Quick Answer
What should you do with missing values and outliers in SPSS?
First identify whether unusual values are errors or genuine observations. Quantify and inspect missing data, review outliers with plots and standardized measures, consider the planned analysis, then choose a defensible treatment. Never delete values automatically just because SPSS flags them.
Start by Understanding the Missing Values
A blank cell may represent a skipped question, an inapplicable question, a technical failure or a respondent who stopped the survey. These situations are not always statistically equivalent. Begin by checking how the questionnaire and data export recorded missing responses.
If special codes such as 99 or -999 were used, define or recode them as missing before calculating means, scale scores or correlations. Otherwise SPSS may treat them as very large or very small real values.
How to Inspect Missing Data in SPSS
Run frequencies or descriptives
Check the valid and missing count for important variables.
Review patterns by variable
Identify questions with unusually high non-response.
Check patterns by case
Look for respondents with large portions of the questionnaire missing.
Consider why values are missing
Use survey design and context rather than assuming all missingness is random.
Document the rule used
Record any case exclusion, substitution or analytic method used to handle missing data.
If you need help preparing or checking your statistical analysis, Academia Helper can provide professional academic support tailored to your research questions, dataset and university requirements.
Choosing How to Handle Missing Data
Possible approaches include analysing available cases, excluding a case from a specific analysis, using scale-specific rules, or applying a formal imputation method. The right choice depends on amount and pattern of missingness, sample size, model and research design.
Simple mean substitution is easy but can reduce variability and distort relationships, so it should not be used automatically. More advanced methods can be appropriate in some studies but require methodological justification and correct implementation.
What Is an Outlier?
An outlier is an observation unusually far from the rest of the data or unusual in relation to a statistical model. Some outliers are data-entry errors; others are genuine respondents with extreme but valid experiences.
That distinction matters. An income value with an extra zero may be a clear error if the original survey confirms it. A genuine very high income is different and should not be deleted only because it changes a mean.
For expert SPSS support, data analysis and professionally written academic work, Academia Helper is here to help you turn statistical output into a clearer and better-structured submission.
Ways to Detect Outliers in SPSS
| Method | Useful for | What to remember |
|---|---|---|
| Boxplot | Quick visual screening of univariate extremes | A flagged point is not automatically invalid |
| Standardized scores | Seeing how far a value is from the variable mean | Threshold conventions vary by context |
| Scatterplot | Finding unusual bivariate patterns | Helps detect leverage and non-linearity |
| Regression diagnostics | Identifying influential cases in a model | Use residual, leverage and influence measures together where appropriate |
For multivariate analyses, a case can be unusual because of a combination of variables even when no single value looks extreme. Use diagnostics suited to the planned model.
How to Decide Whether to Keep, Correct or Remove a Case
Correct confirmed data-entry errors using the original response. If the value is impossible but the true value cannot be recovered, treat it according to your missing-data rule. For a genuine extreme observation, investigate its influence rather than deleting it automatically.
You may compare results with and without a highly influential case as a sensitivity check, but report the decision transparently. Exclusion should be based on a defensible rule, not on whether removal makes the hypothesis significant.
If you want an experienced academic writer to review your analysis, tables or results section, Academia Helper offers professional support for dissertations, reports and other university assignments.
How to Report Data Screening
In the methods or results section, state how missing values and outliers were identified and what action was taken. If no cases were removed, you can still report that screening was conducted. If cases were excluded, state how many and the criterion used.
Keep a cleaning log that links each action to a variable or respondent ID. This makes later analysis reproducible and helps you explain your dataset confidently during supervision or viva questions.
Common Missing-Data and Outlier Mistakes
- Treating 99 or -999 as real values.
- Deleting all boxplot outliers without investigation.
- Using mean substitution automatically.
- Removing cases because they weaken statistical significance.
- Checking only individual variables when the planned model is multivariate.
- Failing to document exclusions and recoding decisions.
If the conclusion changes materially, report that limitation and explain the final decision transparently.
When one or two genuine observations strongly influence a model, compare the main result with a reasonable alternative analysis when appropriate. The purpose is not to search for the version that gives the preferred p-value, but to understand whether the conclusion depends heavily on a small number of cases.
Use Sensitivity Checks for Important Decisions
Key Takeaways
- Define special missing codes before analysis.
- Inspect both amount and pattern of missing data.
- Separate data errors from genuine extreme observations.
- Use outlier diagnostics appropriate to the planned analysis.
- Never remove cases simply to improve significance.
- Document every correction, exclusion or imputation decision.