Data dashboard illustrating missing values and unusual observations during SPSS data screening
SPSS & Data Analysis

Missing Values and Outliers in SPSS: What Students Should Do

A practical guide to identifying missing data and outliers in SPSS and deciding what to do without deleting cases automatically.

missing data SPSSSPSS outliersmissing valuesdata screeningdissertation analysis

Quick Answer

What should you do with missing values and outliers in SPSS?

First identify whether unusual values are errors or genuine observations. Quantify and inspect missing data, review outliers with plots and standardized measures, consider the planned analysis, then choose a defensible treatment. Never delete values automatically just because SPSS flags them.

Start by Understanding the Missing Values

A blank cell may represent a skipped question, an inapplicable question, a technical failure or a respondent who stopped the survey. These situations are not always statistically equivalent. Begin by checking how the questionnaire and data export recorded missing responses.

If special codes such as 99 or -999 were used, define or recode them as missing before calculating means, scale scores or correlations. Otherwise SPSS may treat them as very large or very small real values.

How to Inspect Missing Data in SPSS

1

Run frequencies or descriptives

Check the valid and missing count for important variables.

2

Review patterns by variable

Identify questions with unusually high non-response.

3

Check patterns by case

Look for respondents with large portions of the questionnaire missing.

4

Consider why values are missing

Use survey design and context rather than assuming all missingness is random.

5

Document the rule used

Record any case exclusion, substitution or analytic method used to handle missing data.

If you need help preparing or checking your statistical analysis, Academia Helper can provide professional academic support tailored to your research questions, dataset and university requirements.

Choosing How to Handle Missing Data

Possible approaches include analysing available cases, excluding a case from a specific analysis, using scale-specific rules, or applying a formal imputation method. The right choice depends on amount and pattern of missingness, sample size, model and research design.

Simple mean substitution is easy but can reduce variability and distort relationships, so it should not be used automatically. More advanced methods can be appropriate in some studies but require methodological justification and correct implementation.

What Is an Outlier?

An outlier is an observation unusually far from the rest of the data or unusual in relation to a statistical model. Some outliers are data-entry errors; others are genuine respondents with extreme but valid experiences.

That distinction matters. An income value with an extra zero may be a clear error if the original survey confirms it. A genuine very high income is different and should not be deleted only because it changes a mean.

For expert SPSS support, data analysis and professionally written academic work, Academia Helper is here to help you turn statistical output into a clearer and better-structured submission.

Ways to Detect Outliers in SPSS

MethodUseful forWhat to remember
BoxplotQuick visual screening of univariate extremesA flagged point is not automatically invalid
Standardized scoresSeeing how far a value is from the variable meanThreshold conventions vary by context
ScatterplotFinding unusual bivariate patternsHelps detect leverage and non-linearity
Regression diagnosticsIdentifying influential cases in a modelUse residual, leverage and influence measures together where appropriate

For multivariate analyses, a case can be unusual because of a combination of variables even when no single value looks extreme. Use diagnostics suited to the planned model.

How to Decide Whether to Keep, Correct or Remove a Case

Correct confirmed data-entry errors using the original response. If the value is impossible but the true value cannot be recovered, treat it according to your missing-data rule. For a genuine extreme observation, investigate its influence rather than deleting it automatically.

You may compare results with and without a highly influential case as a sensitivity check, but report the decision transparently. Exclusion should be based on a defensible rule, not on whether removal makes the hypothesis significant.

If you want an experienced academic writer to review your analysis, tables or results section, Academia Helper offers professional support for dissertations, reports and other university assignments.

How to Report Data Screening

In the methods or results section, state how missing values and outliers were identified and what action was taken. If no cases were removed, you can still report that screening was conducted. If cases were excluded, state how many and the criterion used.

Keep a cleaning log that links each action to a variable or respondent ID. This makes later analysis reproducible and helps you explain your dataset confidently during supervision or viva questions.

Common Missing-Data and Outlier Mistakes

  • Treating 99 or -999 as real values.
  • Deleting all boxplot outliers without investigation.
  • Using mean substitution automatically.
  • Removing cases because they weaken statistical significance.
  • Checking only individual variables when the planned model is multivariate.
  • Failing to document exclusions and recoding decisions.

If the conclusion changes materially, report that limitation and explain the final decision transparently.

When one or two genuine observations strongly influence a model, compare the main result with a reasonable alternative analysis when appropriate. The purpose is not to search for the version that gives the preferred p-value, but to understand whether the conclusion depends heavily on a small number of cases.

Use Sensitivity Checks for Important Decisions

Key Takeaways

  • Define special missing codes before analysis.
  • Inspect both amount and pattern of missing data.
  • Separate data errors from genuine extreme observations.
  • Use outlier diagnostics appropriate to the planned analysis.
  • Never remove cases simply to improve significance.
  • Document every correction, exclusion or imputation decision.

Frequently Asked Questions

Should I delete every missing case?
No. The correct approach depends on how much is missing, why it may be missing and which analysis you plan to run.
What does SPSS mean by an outlier in a boxplot?
It flags a value relatively far from the central distribution using the boxplot rule. The point may still be genuine and should be investigated.
Can I replace missing values with the mean?
You can in some contexts, but simple mean substitution can distort variance and relationships. Do not use it automatically without methodological justification.
Should I remove a respondent with many missing answers?
Possibly, if a pre-defined completeness rule supports exclusion. Apply the rule consistently and document it.
What is an influential case?
It is an observation that has a relatively large effect on a fitted statistical model. Influence is not the same as simply having a high or low value.
How do I report outlier removal?
State the diagnostic or rule used, how many cases were affected, why the decision was made and whether the analysis changed materially if relevant.

Need Expert Help With Your Academic Work?

Expert writers · Plagiarism-free · 0% AI content · On-time delivery · UK & USA academic standards