FESK.COMYour global study desk
Email us
Statistics & Data

Handling missing data: why deleting rows and mean imputation can both mislead

A survey of 600 shrinks to 410 once software drops every row with a blank. The three missing-data mechanisms, the cost of deletion and mean imputation, multiple imputation, and reporting that reviewers can trust.

Handling missing data: why deleting rows and mean imputation can both mislead

A survey of 600 students on academic stress. When the regression runs, the software reports N = 410. Almost a third of the sample has vanished because each of those students skipped at least one question, usually family income or grade point average. Worse, the students who skipped the grades question tend to be the ones with low grades. Results based on the remaining 410 describe a rosier group than the one actually surveyed.

Missing data turn up in almost every empirical study. The issue is not whether data are missing but why, and whether your method matches the reason.

Three missing-data mechanisms

MechanismMeaningExample
MCAR (missing completely at random)Missingness is unrelated to any variable, observed or notSome paper questionnaires lost a page in the rain
MAR (missing at random)Missingness depends on observed variables, not on the missing value once those are accounted forFinal-year students skip the income question more often, and year of study is recorded
MNAR (missing not at random)Missingness depends on the missing value itself, even after accounting for observed variablesStudents with low grades decline to report them

The catch: the data cannot prove whether they are MAR or MNAR, because you never see the missing values. MCAR can be partly refuted, if people with missing data differ on observed variables, but never proven. Little’s MCAR test is often reported, yet a non-significant result does not establish MCAR.

The cost of deleting rows

Most software by default drops any observation missing any variable in the model: listwise deletion, or complete-case analysis. Two consequences:

  • Lost power. Going from 600 to 410 throws away nearly a third of the information, including the 190 people’s other, complete answers.
  • Bias if data are not MCAR: the remaining group no longer represents the original one, as in the opening example.

Complete-case analysis can still be acceptable when missingness is minimal and plausibly close to MCAR, or in certain situations where missingness depends only on predictors. It should not be the unexamined default.

Why mean imputation is a bad idea

Filling every blank with the variable’s mean sounds neutral. In reality it:

  • Artificially shrinks variance: hundreds of people with exactly the same value.
  • Weakens correlations between that variable and others.
  • Makes standard errors too small, so results look more precise than they are.

Last observation carried forward in longitudinal studies has similar problems: it assumes people who drop out stay exactly as they were, which is rarely true.

Multiple imputation and full information maximum likelihood

Two methods are widely recommended when data are plausibly MAR.

Multiple imputation creates several complete datasets, filling each gap with a prediction from other variables plus random noise reflecting uncertainty. You run the analysis on each dataset and pool the results using Rubin’s rules. Common mistakes:

  • The imputation model must include every variable in the analysis model, including the outcome. It sounds circular but it is required; leaving the outcome out biases associations towards zero.
  • Add auxiliary variables: variables not in the analysis model that predict the missing values or the probability of missingness. They make the MAR assumption more plausible.
  • Number of imputations: a popular rule of thumb is at least as many as the percentage of incomplete cases, so around 30 for 30% missing. The handful once recommended is often too few.
  • Check imputed values are plausible: no negative ages, distributions not wildly different from observed data.

Full information maximum likelihood (FIML) uses all available information from each person directly in estimation, without creating new datasets. It is common in structural equation modelling and mixed models.

When you suspect MNAR: sensitivity analysis

No standard method fully fixes MNAR. The honest approach is a sensitivity analysis: assume people with missing grades scored 0.5, then 1 point lower than similar people with observed grades, and see whether conclusions change. If they hold under plausible assumptions, you can be more confident; if not, say so clearly in the limitations.

Tools

  • R: the mice package is the most widely used for multiple imputation; lavaan supports FIML.
  • Python: statsmodels offers multiple imputation; imputers in machine learning libraries usually produce a single completed dataset, better suited to prediction than statistical inference.
  • SPSS: has a multiple imputation module with automatic pooling for many analyses.
  • Stata: the mi suite of commands.

In practice: workflow and reporting

  1. Describe: proportion missing for each variable, number of complete cases, pattern of missingness (scattered or in blocks).
  2. Compare people with and without missing data on observed variables to judge whether MCAR is plausible.
  3. Choose the primary method on the basis of an argument about the mechanism, not on which result looks better.
  4. Implement multiple imputation or FIML with suitable auxiliary variables.
  5. Run sensitivity analyses: compare with complete-case results and test MNAR scenarios.
  6. Report the method, software, number of imputations, variables in the imputation model and sensitivity results.

Model paragraph: “Thirty-two per cent of participants were missing at least one variable. Those with missing data were more likely to be final-year students. We used multiple imputation by chained equations (mice package, 40 imputations), with an imputation model including all analysis variables, the outcome and three auxiliary variables. Complete-case results were similar and are shown in Table S3.”

Do this now: before your next analysis, produce a table of missingness for every variable and compare the N in each model with the full sample. If the gap is more than a few per cent, you need a reasoned plan for handling it, rather than letting the software decide quietly.

Câu hỏi thường gặp

What is the difference between MCAR, MAR and MNAR?

MCAR means missingness is completely random; MAR means it depends on observed variables; MNAR means it depends on the missing value itself. The mechanism determines which methods are appropriate.

Should I replace missing values with the mean?

Not for inferential analysis. Mean imputation shrinks variance, weakens correlations and makes standard errors artificially small.

How many imputations do I need?

A common rule of thumb is at least as many imputations as the percentage of incomplete cases, so about 30 for 30% missing. A handful is often too few for stable results.

Should the outcome variable be in the imputation model?

Yes. The imputation model must include every variable in the analysis model, including the outcome, or estimated associations will be biased towards zero.

How much missing data is too much?

There is no universally safe threshold. The mechanism matters most; even a small amount of missing data can cause bias if missingness is strongly related to the missing values.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com