Resource

Deleting Every Incomplete Case Changed the Answer: A Missing-Data Case Study

A small percentage of missing cells can remove a much larger percentage of participants from a regression. This synthetic case study compares complete-case analysis with multiple imputation and sensitivity analysis—and shows why the conclusion can change.

Missing data are not just a data-cleaning nuisance. They can change who contributes to an analysis, how precisely parameters are estimated, and—when the assumptions behind deletion are not credible—what the fitted regression appears to say. Complete-case analysis is especially consequential because every record missing any variable required by the model is excluded, even when most of that record is observed (Dmitrienko & Koch, 2017; Lovric, 2011; Tabachnick & Fidell, 2013).

Synthetic case study: This example is entirely synthetic. Its purpose is methodological: to show why researchers should examine where missingness occurs, which observations become excluded, and which assumptions are required before deciding how incomplete data should be handled.

The research question

A synthetic organizational study enrolled 600 employees. The research question was:

After accounting for baseline workload, age, tenure, and income, is perceived manager support associated with lower six-month burnout scores?

The primary regression was specified as:

Six-month burnout = manager support + baseline burnout + workload + age + tenure + income

All variables and results below were generated solely for this case study. They are not empirical findings.

Interpretation boundary: The coefficient of primary interest is the regression coefficient for manager support. Because this is an observational synthetic study, the coefficient is interpreted as an adjusted association, not a causal effect.

The first analysis: delete every incomplete record

The initial analyst inspected the dataset, saw missing values, and selected complete-case—or listwise—deletion without further investigation.

That decision reduced the regression sample from 600 to 468 employees.

Original sample

600 employees

Complete-case sample

468 employees

Manager-support coefficient

-0.62

Complete-case inference

95% CI: -1.23 to -0.01
p = .046

On this analysis alone, the conclusion would be that higher manager support was associated with slightly lower six-month burnout after adjustment for the other predictors.

But the important question was not simply whether 468 observations were “enough.”

The more important question: Which 468 observations remained, and what assumptions made the exclusion of the other 132 observations defensible?

Complete-case analysis is simple, but it discards observed information from every incomplete case. It can therefore lose information and precision, and its validity depends on conditions concerning the missing-data process rather than merely on the percentage of missing values (Dmitrienko & Koch, 2017; Lovric, 2011).

Step 1: Where was the missingness?

Of the 600 synthetic participants:

  • 468 (78.0%) had complete data on every variable required by the regression.
  • 132 (22.0%) had at least one missing value.
  • No individual analysis variable had anything close to 22% missingness.
Variable-level missingness
Variable Missing n Missing %
Six-month burnout 54 9.0%
Manager support 47 7.8%
Income 31 5.2%
Tenure 24 4.0%
Workload 19 3.2%
Baseline burnout 0 0%
Age 0 0%

Counts overlap because some participants were missing more than one variable.

This distinction matters. Tabachnick and Fidell (2013) recommend examining missing-data patterns during data screening rather than treating missingness as a single undifferentiated property of the dataset. They distinguish missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR), with different implications for how missing observations can be handled.

Why one overall missing percentage is inadequate

Suppose the analyst reported only that approximately 5% of all cells in the regression dataset were missing. That number sounds modest.

It does not reveal that 22% of participants were excluded from the complete-case regression.

Cell-level question

How many values in the rectangular dataset are absent?

Approximately 5% of cells in this example.

Complete-case question

How many observations have any missing value among the variables required for this analysis?

22% of participants in this example.

With several variables, relatively small variable-specific missingness rates can accumulate into substantial case loss. Dmitrienko and Koch (2017) illustrate this general problem: as the number of required measurements grows, complete-case analysis can discard a large proportion of observations even when the missing percentage on each individual measurement is comparatively modest.

For regression, the relevant missing-data description therefore needs at least three levels:

  1. missingness by variable;
  2. missingness patterns across variables; and
  3. the number and characteristics of observations actually entering each analysis.

An overall percentage alone does not provide that information.

Step 2: What patterns were visible?

The 132 incomplete observations did not all have the same configuration.

The most common synthetic patterns were:

Common missing-data patterns among incomplete observations
Pattern n
Six-month burnout only missing 35
Manager support only missing 31
Income only missing 18
Tenure only missing 13
Workload only missing 10
Two analysis variables missing 19
Three or more analysis variables missing 6

This was already more informative than saying “the dataset had about 5% missing data.”

The analyst then compared observed variables between the 468 complete cases and the 132 incomplete cases.

Observed characteristics of complete and incomplete cases
Observed characteristic Complete cases Incomplete cases
Mean baseline burnout 48.1 55.6
Mean workload score 52.4 59.8
Mean age 39.7 38.9

The incomplete cases had higher observed baseline burnout and workload in this synthetic dataset.

What this does not establish: The difference between complete and incomplete cases is an empirically observable feature of the generated dataset. It is not proof of why values were missing.

Associations between missingness indicators and observed variables can show that missingness is patterned, but they do not establish the unobserved missing-data mechanism. In particular, observed data generally cannot establish that missingness is MAR rather than MNAR, because that distinction concerns dependence on values that are not observed (Dmitrienko & Koch, 2017; Lovric, 2011).

Observable facts are not missing-data assumptions

The analysis therefore separated what could be observed from what had to be assumed.

Observable in the synthetic dataset

  • 22% of participants had at least one missing regression variable.
  • Missingness differed substantially across variables.
  • Several missingness patterns occurred.
  • Incomplete observations differed from complete observations on some fully or partially observed characteristics.
  • Complete-case deletion reduced the regression population from 600 to 468.

Not established from those observations

  • That missingness was MCAR.
  • That missingness was MAR.
  • That missingness was MNAR.
  • Why a particular participant had a missing value.
  • That one missing-data method would necessarily recover a uniquely “correct” answer.

The MCAR/MAR/MNAR distinction is an assumption about the missingness process. MCAR requires missingness to be independent of observed and unobserved data; under MAR, conditional on observed data, missingness is independent of the unobserved measurements; MNAR covers processes that do not satisfy MCAR or MAR (Dmitrienko & Koch, 2017; Lovric, 2011).

Because this is a synthetic example, the data-generating procedure is known to the case-study designer. Missingness probabilities were deliberately made dependent on observed baseline burnout and workload. That construction provides a controlled demonstration of an MAR-type scenario. Importantly, an analyst receiving only the incomplete dataset would not be entitled to infer MAR merely from the patterns seen in the data.

Step 3: Complete-case deletion changed the analysis population

The most important effect of listwise deletion was not that the spreadsheet became smaller. It changed the set of people represented by the fitted model.

The original synthetic sample contained 600 participants.

The complete-case analysis contained only participants with observed values for:

  • manager support;
  • six-month burnout;
  • baseline burnout;
  • workload;
  • age;
  • tenure; and
  • income.

Because the excluded and retained groups differed on observed baseline burnout and workload, the complete-case regression was being estimated in a selected subset of the original sample.

Lovric (2011) notes that complete-case analysis loses the information contained in observed variables for incomplete cases and that the resulting complete cases may not be representative of the original sample. Tabachnick and Fidell (2013) similarly warn that deleting cases can distort the sample when cases with missing values are not randomly distributed through the data.

This is why a statement such as “132 cases were deleted because of missing values” is incomplete reporting. Researchers should also ask whether the retained analysis population differs from the population originally measured.

What assumption does complete-case analysis require?

Complete-case analysis is attractive because it is straightforward: retain only observations with all required measurements and fit the intended model using standard software.

That simplicity does not make it assumption-free.

Dmitrienko and Koch (2017) emphasize that complete-case analysis can produce substantial information loss and that, in the setting they discuss, consistency of the corresponding complete-case estimator requires the more restrictive MCAR condition; bias can result when missingness is MAR but not MCAR.

This mattered here because the observed differences between complete and incomplete records provided no reassurance that a strong MCAR interpretation was appropriate.

The point is not that an observed difference mechanically “proves MCAR false.” The safer conclusion is narrower: the automatic complete-case analysis was relying on a missing-data assumption that required justification, while simultaneously discarding 22% of the participants.

Step 4: Reanalyzing the data with multiple imputation

The analysis was repeated using multiple imputation.

Multiple imputation replaces each missing value with multiple plausible values rather than pretending that one filled-in value is known with certainty. The resulting datasets are analyzed separately and their estimates combined so that uncertainty associated with the missing values contributes to the final inference. Dmitrienko and Koch (2017) explicitly contrast this with single imputation and note that multiple imputation acknowledges uncertainty arising from filling in rather than observing values.

For this synthetic analysis, the imputation model included:

  • six-month burnout;
  • manager support;
  • baseline burnout;
  • workload;
  • age;
  • tenure; and
  • income.

Fifty imputations were generated. The same substantive regression was fitted within each completed dataset, and estimates were pooled.

Complete-case and multiple-imputation results
Analysis Records contributing information Manager-support coefficient 95% CI p-value
Complete cases 468 -0.62 -1.23 to -0.01 .046
Multiple imputation 600 -0.34 -0.78 to 0.10 .129

The answer changed.

The complete-case analysis placed the 95% confidence interval just below zero and produced a conventional p-value below .05. The multiply imputed analysis estimated a smaller negative association, with an interval spanning zero.

The appropriate interpretation is not that multiple imputation proved there was no association.

Under the multiple-imputation model and its assumptions, the data did not provide the same evidence for an adjusted association that appeared in the complete-case analysis.

That is a change in inference, not proof that one coefficient represents an underlying causal truth.

What does multiple imputation assume?

Multiple imputation does not “solve” missing data without assumptions.

Dmitrienko and Koch (2017) describe multiple imputation as valid under MAR in the framework they present. Under MAR, missingness may depend on observed information, but conditional on that observed information it does not depend on the missing values themselves. Lovric (2011) likewise describes MAR-based likelihood and Bayesian analyses as ignorable under additional parameter-distinctness conditions.

The imputation model also matters. The variables used to generate imputations need to carry the information required to make the assumed missing-data model plausible and to represent relevant relationships in the analysis.

Thus the comparison was not:

biased deletion versus assumption-free imputation.

It was:

a complete-case analysis relying on restrictive missingness conditions versus an imputation analysis relying on a different, explicitly stated set of assumptions.

That framing keeps the methodological uncertainty visible.

Why not use mean substitution?

A tempting alternative would be to replace each missing value with the observed mean of its variable.

That would retain all 600 rows, but retaining rows is not sufficient.

Simple mean imputation reduces variability and can distort relationships among variables. Lovric (2011) notes that unconditional mean imputation can produce inconsistent estimates for parameters such as variances and regression coefficients, while inferences from the filled-in dataset can have distorted precision. Tabachnick and Fidell (2013) similarly describe the reduction in variance and correlations produced by mean substitution.

Newton and Rudestam (1999) also caution that mean substitution artificially reduces the variance of a variable and may reduce its correlations with other variables.

For a regression question, those are fundamental problems rather than cosmetic disadvantages.

A likelihood-based alternative

Multiple imputation is not the only principled alternative supported by the sources.

For appropriate longitudinal or multivariate models, likelihood-based methods can use the observed portions of incomplete records without first filling every blank with a single value. Under MAR and appropriate model conditions, direct-likelihood analysis can provide valid inference while avoiding automatic deletion of every incomplete case (Dmitrienko & Koch, 2017; Lovric, 2011).

In this case study, multiple imputation was chosen because the substantive analysis was an ordinary multiple regression with missingness distributed across both the outcome and predictors.

If the research question had instead been formulated as a longitudinal mixed model with repeated outcomes, a likelihood-based analysis could have been a particularly natural primary approach. Dmitrienko and Koch (2017) discuss likelihood-based analyses, Bayesian analyses, and multiple imputation as modern approaches to incomplete data rather than defaulting automatically to historical procedures such as complete-case analysis.

The choice should therefore follow the research design, analysis model, missingness pattern, and assumptions—not a universal rule that one missing-data procedure is always best.

Step 5: Sensitivity analysis

An MAR analysis is still an assumption-based analysis.

The observed data cannot generally tell us what the unobserved values would have been under an MNAR process. Dmitrienko and Koch (2017) therefore emphasize sensitivity analysis: a primary model can be supplemented with plausible alternatives, and researchers can examine whether conclusions remain stable across them.

For this synthetic case study, the multiple-imputation analysis under MAR was treated as the primary incomplete-data analysis. We then performed an illustrative sensitivity analysis in which the imputed six-month burnout values were shifted under progressively less favorable scenarios.

Important: These shifts are sensitivity parameters, not estimates obtained from the observed data.

Illustrative sensitivity-analysis results
Sensitivity scenario Manager-support coefficient 95% CI Interpretation
Complete cases -0.62 -1.23 to -0.01 Interval excludes zero narrowly
MI under MAR -0.34 -0.78 to 0.10 Interval includes zero
Mild MNAR departure -0.27 -0.72 to 0.18 Same broad inference as MAR
Moderate MNAR departure -0.18 -0.65 to 0.29 Association attenuates further
Opposite-direction sensitivity scenario -0.45 -0.92 to 0.02 Still close to inferential boundary

The exact numerical shifts are synthetic and illustrative. They should not be treated as empirically estimated missingness effects.

The important result is the pattern: the apparently “significant” complete-case result was not robust to alternative ways of handling the incomplete observations.

Dmitrienko and Koch (2017) describe sensitivity analysis broadly as examining multiple plausible models or variations of a primary model and assessing the extent to which conclusions remain stable. They specifically discuss moving from a primary MAR analysis to MNAR variations, including pattern-mixture approaches.

So which answer should be reported?

Reporting only the complete-case p-value would give an incomplete account of the evidence.

The main finding of this synthetic case is:

Higher manager support had a small negative adjusted association with six-month burnout across the fitted models, but the inferential conclusion depended on how incomplete observations were handled. Complete-case analysis produced an estimate of -0.62 (95% CI [-1.23, -0.01]), whereas multiple imputation under MAR produced -0.34 (95% CI [-0.78, 0.10]). Sensitivity analyses did not support treating the complete-case significance result as robust.

That statement preserves several distinctions.

  • The coefficient is an association, not a causal effect.
  • The confidence interval describes statistical uncertainty under a fitted model; it does not determine practical importance.
  • The disagreement between analyses is itself useful information: the conclusion is sensitive to missing-data assumptions.

Why the complete-case answer changed

Three things happened when every incomplete observation was deleted.

1. Information was discarded

A participant missing income but observed on manager support, baseline burnout, workload, age, tenure, and six-month burnout contributed nothing to the complete-case regression.

Complete-case deletion therefore discarded observed information along with the missing values. This is a recognized disadvantage of listwise deletion (Dmitrienko & Koch, 2017; Lovric, 2011).

2. The analysis population changed

The retained complete cases had lower baseline burnout and workload than the excluded incomplete observations in this synthetic dataset.

The fitted complete-case model therefore described a selected subset rather than all 600 enrolled participants.

3. The coefficient and standard error changed

The issue was not simply “less power.”

The coefficient itself moved from -0.62 to -0.34.

If only the standard error had changed, the story would mainly concern lost precision. Here, both the estimated association and its uncertainty changed, showing that deletion affected more than sample size.

A practical workflow for missing data in regression

This case suggests a more defensible workflow than selecting “listwise deletion” automatically.

  1. Define the intended analysis first.

    Identify the outcome, predictors, adjustment variables, interactions, repeated measurements, and target population before deciding what constitutes an incomplete case.

    A record can be complete for one analysis and incomplete for another.

  2. Describe missingness by variable.

    Report counts and percentages separately for every analysis variable. Do not rely on one dataset-wide percentage.

  3. Examine missingness patterns.

    Determine whether missing values occur mainly in one variable, occur jointly across variables, or create many distinct patterns. Tabachnick and Fidell (2013) treat examination of missing-data patterns as part of data screening.

  4. Quantify the analysis-population consequence.

    State how many observations remain under complete-case analysis.

    Compare observed characteristics of included and excluded observations where appropriate. These comparisons describe the data; they do not by themselves identify the missingness mechanism.

  5. State assumptions explicitly.

    Do not write “the data were MAR” as though MAR were directly observed.

    “The primary multiple-imputation analysis assumed MAR conditional on the variables included in the imputation model.”

    Likewise, if complete-case analysis is used, state the conditions under which it is being treated as valid.

  6. Match the method to the analysis.

    Complete-case analysis, multiple imputation, and likelihood-based methods do not rely on identical assumptions or use incomplete observations in identical ways. Multiple imputation and direct likelihood provide principled ways to retain observed information under suitable assumptions (Dmitrienko & Koch, 2017; Lovric, 2011).

  7. Perform sensitivity analysis when the assumption matters.

    If a substantive conclusion depends on an unverifiable MAR assumption, evaluate plausible departures rather than presenting the primary model as definitive. Sensitivity analysis asks whether reasonable changes to assumptions materially alter the conclusion (Dmitrienko & Koch, 2017).

  8. Report whether the conclusion changes.

    Do not hide disagreement among analyses.

    If complete cases give p = .046 and multiple imputation gives p = .129, that discrepancy is methodologically relevant. Report the coefficients and uncertainty intervals as well as p-values so readers can see whether the change reflects effect estimates, precision, or both.

What should be documented in a research report?

A missing-data section should allow a reader to reconstruct what happened between the original dataset and the final analysis.

At minimum, report:

  • the original number of observations;
  • the number and percentage missing for each analysis variable;
  • the number and percentage of observations with any missing analysis data;
  • important observed missingness patterns;
  • the number of observations entering each primary analysis;
  • whether included and excluded observations differed on relevant observed variables;
  • the missing-data method used;
  • the assumptions required by that method;
  • the variables included in an imputation model, when applicable;
  • the number of imputations and pooling procedure, when applicable;
  • the likelihood-based model used, when applicable;
  • any sensitivity analyses and their prespecified or explicitly chosen sensitivity parameters; and
  • whether substantive conclusions changed across analyses.

Dmitrienko and Koch (2017) emphasize that incomplete-data analysis requires attention not only to the primary method but also to assumptions and sensitivity. More generally, data preparation and missing-value decisions should be documented rather than treated as invisible preprocessing (Newton & Rudestam, 1999; Tabachnick & Fidell, 2013).

Example reporting language

Methods

Of 600 observations, 132 (22.0%) contained at least one missing value among variables required for the primary regression, leaving 468 complete cases. Variable-specific missingness ranged from 0% to 9.0%. Incomplete observations had higher observed baseline burnout and workload than complete observations; these differences were treated as descriptions of the observed missingness pattern rather than evidence establishing a specific missing-data mechanism. The primary incomplete-data analysis used multiple imputation under an MAR assumption conditional on variables included in the imputation model. Fifty imputed datasets were generated and estimates were pooled. Complete-case analysis was retained as a comparison, and sensitivity analyses examined departures from the MAR specification.

Results

Complete-case regression estimated an adjusted manager-support coefficient of -0.62 (95% CI [-1.23, -0.01], p = .046; n = 468). Multiple imputation yielded a smaller estimate of -0.34 (95% CI [-0.78, 0.10], p = .129; N = 600 contributing observed information). Sensitivity analyses under alternative missing-data specifications did not restore a stable conclusion that the adjusted association differed from zero. The inference was therefore considered sensitive to the handling of missing data.

The lesson from the case

The mistake was not simply that the analyst used complete-case analysis.

The mistake was using it automatically.

Twenty-two percent of participants had at least one missing regression value even though no individual variable exceeded 9% missingness. Deleting those participants changed the analysis population and produced a result that was not reproduced under multiple imputation or the sensitivity analyses.

Missing-data methods do not manufacture information that was never observed. Multiple imputation and likelihood approaches still require assumptions, and MAR cannot simply be declared from an inspection of the observed dataset (Dmitrienko & Koch, 2017; Lovric, 2011).

The defensible question is therefore not:

“What percentage of my dataset is missing?”

It is:

Where is information missing, which observations does my proposed method exclude or retain, what assumptions does that method require, and does the scientific conclusion survive reasonable alternatives?

In this synthetic study, asking those questions changed the answer.

References

Dmitrienko, A., & Koch, G. G. (Eds.). (2017). Analysis of clinical trials using SAS: A practical guide (2nd ed.). SAS Institute.

Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2

Newton, R. R., & Rudestam, K. E. (1999). Your statistical consultant: Answers to your data analysis questions. SAGE Publications.

Tabachnick, B. G., & Fidell, L. S. (2013). Using multivariate statistics (6th ed.). Pearson.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry