Resource

Missing Data in Research: MCAR, MAR, MNAR and How to Choose an Analysis Strategy

Missing values can affect who contributes to an analysis, the precision of estimates, and the population represented by the results. This Resource explains how MCAR, MAR, and MNAR assumptions relate to complete-case analysis, multiple imputation, likelihood-based methods, and sensitivity analysis.

Missing values are not merely a data-cleaning inconvenience. They can change who contributes to an analysis, how precisely effects are estimated, and—when missingness selects a systematically different subset—what the resulting estimate represents.

The common default of deleting every incomplete row is therefore not a neutral preprocessing decision. Little and Rubin emphasize that complete-case analysis can lose information and can be biased when complete observations are not representative for the estimand of interest. The consequences depend on the missingness pattern, mechanism, variables affected, and target of the analysis—not simply on the percentage of blank cells (Little & Rubin, 2020).

A practical strategy for missing data in research:

identify the pattern → identify variables with missingness → consider a plausible mechanism → define the analysis target → choose a method → conduct diagnostics and sensitivity analysis.

First: What Are You Trying to Estimate or Predict?

Missing-data decisions should follow the research question rather than precede it.

Before choosing complete case analysis, multiple imputation, or a likelihood-based method, identify:

  • the outcome and predictors required for the analysis;
  • whether the goal is estimation, hypothesis testing, or prediction;
  • the population or sample to which the result is intended to apply;
  • whether observations are independent, clustered, or repeatedly measured;
  • which variables are incomplete; and
  • whether missingness occurs once, intermittently, or through longitudinal dropout.

This matters because an observation can be complete for one analysis but incomplete for another. Complete-case analysis also does not merely remove blanks: it defines a subset of observations on which the model is fitted.

For prediction research, missing predictor values are especially important because excluding people who lack any one required predictor can rapidly reduce the development sample. Steyerberg describes complete-case analysis as inefficient because it discards information from subjects who have some, but not all, predictor measurements (Steyerberg, 2019).

MCAR, MAR and MNAR: What Do They Actually Mean?

MCAR, MAR and MNAR describe assumptions about the relationship between missingness and the data. They should not be treated simply as three labels that can be read directly from a missingness table.

MCAR: Missing Completely at Random

Under missing completely at random (MCAR), missingness does not depend on the observed or missing data relevant to the missingness process. Conceptually, the complete observations behave like a random subset with respect to the variables involved.

MCAR is therefore a strong condition. It is important because many simple deletion or available-data procedures require MCAR for generally valid inference. Little and Rubin note that when complete cases are not effectively a random sample of the full observations, complete-case estimates can lose both precision and validity for relevant estimands (Little & Rubin, 2020).

MAR: Missing at Random

Under missing at random (MAR), missingness may depend on information that has been observed, but conditional on the observed information it does not additionally depend on the values that remain missing.

In a longitudinal setting, for example, dropout may depend on responses recorded before dropout while satisfying MAR if, conditional on the relevant observed history, dropout does not additionally depend on subsequent unobserved responses. Fitzmaurice et al. distinguish this situation from MCAR and discuss its implications for longitudinal analysis (Fitzmaurice et al., 2011).

MNAR: Missing Not at Random

Missing not at random (MNAR) applies when the missingness process retains dependence on missing values after accounting for the relevant observed information. In longitudinal research, for example, dropout can be MNAR when it depends on the unobserved response even after conditioning on observed response history and other relevant measured variables.

MNAR creates a fundamental difficulty: the dependence involves information that, by definition, is not observed. Analyses therefore require additional modeling assumptions, and sensitivity analyses become particularly important (Little & Rubin, 2020; Fitzmaurice et al., 2011).

A practical distinction

Practical distinction among MCAR, MAR, and MNAR
Mechanism Missingness may depend on observed data? May it still depend on the missing value after conditioning on observed information? Practical implication
MCAR No relevant dependence No Some simple deletion/available-data procedures may be valid, although inefficient.
MAR Yes No Likelihood-based or multiple-imputation approaches may be appropriate under suitable models.
MNAR Yes Yes Explicit assumptions about the missingness process and sensitivity analysis become important.

MAR Is an Assumption, Not Something You Prove From the Missingness Table

One of the most important misconceptions about MCAR MAR MNAR is the idea that researchers can inspect their observed dataset and determine conclusively that the data are MAR.

They generally cannot.

Observed information can reveal useful facts. Researchers can determine which variables have missing values, examine patterns of missingness, and compare observed characteristics of complete and incomplete observations. Such comparisons may provide evidence against a simple MCAR explanation. Little and Rubin explicitly note, however, that these comparisons offer no direct evidence establishing the weaker MAR assumption (Little & Rubin, 2020).

The distinction between MAR and MNAR concerns how missingness relates to values that are not observed. Consequently, finding that missingness is associated with observed variables does not prove that there is no additional dependence on unobserved values.

A defensible report usually says:

“The primary analysis assumed MAR conditional on the observed information included in the model.”

It should not simply declare:

“The data were MAR.”

Why Simply Deleting Incomplete Rows Can Change Your Answer

Complete-case analysis—also called listwise deletion—restricts the analysis to observations for which every required variable is observed.

Its attractions are obvious: it is simple, standard complete-data procedures can be applied, and every reported analysis based on those variables uses the same subset. But Little and Rubin identify two major consequences of discarding incomplete observations: loss of precision and potential bias, with the severity depending on the missingness mechanism, missingness pattern, differences between complete and incomplete observations, and the estimand being targeted (Little & Rubin, 2020).

Deleting rows can change an answer through several pathways.

  1. It discards observed information. Someone missing one covariate may nevertheless have a recorded outcome and nearly every other predictor. Complete-case analysis throws all of that information away for that analysis.

  2. It can change the effective analysis population. If becoming a complete case is related to variables relevant to the outcome or model, the retained subset can differ systematically from the original sample.

  3. It usually reduces precision. Fewer observations contribute to estimation.

  4. The effect can accumulate across variables. In a multivariable model, modest missingness in several different predictors can leave substantially fewer complete observations than any one variable-specific missing percentage suggests. Steyerberg highlights this problem in prediction modeling and emphasizes the inefficiency of discarding subjects who retain useful predictor information (Steyerberg, 2019).

The correct question is therefore not merely:

“What percentage of cells are missing?”

It is:

“Which observations will my proposed analysis retain, what information will it discard, and does that retained subset still support the estimand or prediction goal I intended?”

Complete-Case Versus Available-Case Analysis

Complete-case analysis

Uses only observations with all variables required for the analysis.

Available-case analysis

Attempts to use the observations available for each particular calculation. A mean for one variable, for example, may use a different set of observations from the mean for another variable or a covariance involving two variables.

Available-case approaches can preserve more observed information than complete-case deletion, but they do not automatically solve the missing-data problem. Different quantities may be estimated from different subsets, and available-case covariance calculations can even produce problematic covariance structures. Little and Rubin do not generally recommend available-case analysis as a universal alternative to principled likelihood or imputation approaches (Little & Rubin, 2020).

The lesson is not that available-case analysis is always wrong. It is that using every available value is not, by itself, a statistical justification.

Single Imputation Versus Multiple Imputation

Imputation fills in missing values using information in the observed data. But not all imputation strategies treat uncertainty correctly.

Single imputation

Single-imputation procedures create one completed dataset.

Examples discussed in the approved sources include mean imputation, conditional regression imputation, and stochastic regression imputation. These approaches differ in sophistication, but once a single value replaces a missing observation, ordinary analysis of the completed dataset can easily treat that imputed value as though it had actually been observed.

Little and Rubin stress that imputation must not create the illusion that the data are genuinely complete. Missing values can be predicted using observed variables, but the uncertainty associated with those predictions remains part of the statistical problem (Little & Rubin, 2020).

Steyerberg likewise notes that simple imputation can underestimate variability, while regression-based methods can use correlations among predictors more effectively (Steyerberg, 2019).

Multiple imputation

With multiple imputation (MI), missing values are replaced multiple times using draws from an imputation model, producing multiple completed datasets. The substantive analysis is performed in each dataset, and the results are combined.

The crucial point is that multiple imputation does not simply produce several guesses and average them. Its inferential advantage is that variability between imputations contributes to uncertainty alongside the ordinary sampling variability estimated within imputations (Little & Rubin, 2020).

Steyerberg makes the same distinction in prediction modeling: variation across completed datasets represents uncertainty about the missing values, and combined variance includes both within- and between-imputation components. That between-imputation component is a fundamental difference between multiple and single imputation (Steyerberg, 2019).

Multiple imputation does not remove assumptions

MI should not be described as an assumption-free repair for missing data.

Its validity depends on the imputation model, the variables used to predict missing information, the relationships represented by that model, and the missingness assumptions under which the procedure is justified.

The comparison is therefore not:

bad deletion versus assumption-free imputation.

It is:

alternative analyses that use different information and depend on different assumptions.

Missing Predictors and Missing Outcomes Are Not the Same Problem

The location of missingness matters.

Missing predictors

When predictors are incomplete, complete-case regression can discard individuals whose outcomes and most other predictors are known. This reduces the information available for estimating predictor effects or developing prediction models.

For prediction research, Steyerberg discusses imputation models that use available predictors and, where appropriate for model development, outcome information to predict missing predictor values. Multiple imputation can propagate uncertainty about those missing predictors through subsequent model estimation (Steyerberg, 2019).

Harrell likewise treats missing data as part of the regression-modeling strategy rather than as an unrelated preliminary cleaning step: handling incomplete observations has consequences for how regression models are developed and interpreted (Harrell, 2015).

Missing outcomes

A missing outcome removes direct information about the response for that observation. The appropriate response depends strongly on study design.

In repeated-measures research, a participant with a missing later outcome may still have several valid earlier outcomes. Discarding that participant completely can therefore waste substantial longitudinal information. Likelihood-based longitudinal models can, under appropriate assumptions, use observed portions of trajectories rather than requiring every participant to complete every measurement occasion (Fitzmaurice et al., 2011).

Thus “missing data” should not be treated as a single homogeneous problem. Ask what is missing and what role that variable plays in the analysis.

When Likelihood-Based Methods Are Appropriate

Multiple imputation is not the only principled approach.

Likelihood-based methods can often use the observed components of incomplete records directly without first constructing a single filled-in dataset. Under appropriate ignorability conditions and correctly specified response models, likelihood-based inference can accommodate MAR missingness.

This is particularly important in longitudinal analysis. Fitzmaurice et al. explain that likelihood-based models can use partial response trajectories under MAR, whereas ordinary complete-case or available-data procedures generally require stronger conditions for valid inference. The advantage comes with a modeling requirement: the joint response model, including relevant mean and within-subject dependence structures, must be appropriately specified (Fitzmaurice et al., 2011).

Accordingly, a likelihood-based mixed model may be a natural strategy for incomplete repeated outcomes, while multiple imputation may be convenient when missingness affects several predictors and other analysis variables.

There is no universal rule that MI is always preferable to likelihood, or vice versa.

Longitudinal Dropout Needs Special Attention

Dropout creates a structured form of missingness because later measurements become unavailable after a participant leaves follow-up.

Suppose dropout is related to earlier observed outcomes. Such missingness is not naturally MCAR. It may nevertheless be compatible with MAR if, conditional on the observed history and relevant measured variables, dropout has no additional dependence on the later outcomes that become unobserved.

This distinction has direct consequences for analysis.

Fitzmaurice et al. note that complete-case analysis discards all measurements from people who later drop out, even measurements collected legitimately before dropout. Available-data methods preserve more of that information but are not automatically valid merely because they use all observed visits. Under MAR that is not MCAR, appropriately specified likelihood-based longitudinal models can provide valid inference under conditions that do not justify ordinary complete-case or standard available-data procedures (Fitzmaurice et al., 2011).

If dropout is MNAR, analyses that simply ignore the missingness process may be inadequate. The dependence on unobserved outcomes then needs to be addressed through additional modeling assumptions or investigated through sensitivity analysis.

A Practical Workflow for Handling Missing Data in Research

1. Identify the missingness pattern

Begin descriptively. Determine how much information is missing, how many observations are affected, which variables tend to be missing together, whether repeated measurements show intermittent gaps or monotone dropout, and how many observations a complete-case analysis would retain.

Do not reduce this description to one overall percentage.

2. Identify which variables are missing

Separate missing outcomes, primary exposures or predictors, adjustment variables, auxiliary variables, repeated outcomes, and variables required only for secondary analyses.

The statistical consequences depend on the role of the missing variable.

3. Consider a plausible missingness mechanism

Ask what scientific or data-collection processes could have generated the missingness.

Then distinguish between what is observed and what must be assumed.

4. Define the analysis target

Clarify the estimand or prediction goal before choosing the missing-data procedure.

For explanatory research, identify the population quantity the coefficient, mean difference, odds ratio, or trajectory contrast is intended to estimate. For prediction, distinguish model development, validation, and prediction for future individuals.

5. Choose a method

Choose a method that matches the target, design, data structure, and missingness assumptions rather than inheriting a software default.

6. Conduct diagnostics and sensitivity analysis

Examine model adequacy, compare transparent alternatives where useful, and determine whether substantive conclusions depend strongly on the treatment of missingness.

1. Identify the missingness pattern

Begin descriptively.

Determine:

  • how much information is missing for each analysis variable;
  • how many observations have at least one missing analysis value;
  • whether variables tend to be missing together;
  • whether repeated measurements show intermittent gaps or monotone dropout; and
  • how many observations a complete-case analysis would actually retain.

Do not reduce this description to one overall percentage.

2. Identify which variables are missing

Separate missing:

  • outcomes;
  • primary exposures or predictors;
  • adjustment variables;
  • auxiliary variables;
  • repeated outcomes; and
  • variables required only for secondary analyses.

The statistical consequences depend on the role of the missing variable.

3. Consider a plausible missingness mechanism

Ask what scientific or data-collection processes could have generated the missingness.

Then distinguish between what is observed and what must be assumed.

Observed associations between missingness and measured variables may make a simple MCAR explanation less plausible. They do not establish MAR over MNAR because the decisive distinction involves unobserved information (Little & Rubin, 2020).

4. Define the analysis target

Clarify the estimand or prediction goal before choosing the missing-data procedure.

For explanatory research, ask what population quantity the coefficient, mean difference, odds ratio, or trajectory contrast is intended to estimate.

For prediction, ask whether the task is model development, validation, or prediction for future individuals and whether missing predictors will also occur at the intended point of use.

5. Choose a method that matches the target and assumptions

Decision framework for selecting a missing-data strategy
Situation Strategy to consider Main caution
Very limited missing information and defensible conditions for deletion Complete-case analysis may be adequate Quantify information loss and justify assumptions
Simple summaries using different available subsets Available-case analysis may sometimes be useful Different quantities may use different samples; not a general solution
Missing predictors/covariates in multivariable analysis Multiple imputation Imputation model and MAR-type assumptions require justification
Repeated outcomes with incomplete follow-up Likelihood-based longitudinal model under suitable MAR assumptions Correct specification of the response model matters
Single-value imputation Generally requires caution Standard completed-data inference can understate missing-data uncertainty
Material concern about MNAR Explicit MNAR modeling and/or sensitivity analysis Conclusions necessarily depend on unverifiable assumptions
This table is a decision aid, not an algorithm. The appropriate strategy depends on the design, estimand, data structure, and missingness process.

6. Conduct diagnostics and sensitivity analysis

After fitting the primary analysis:

  • examine whether the imputation or likelihood model adequately represents important relationships in the observed data;
  • compare the effective sample and estimates with transparent alternatives where useful;
  • inspect whether conclusions depend strongly on the treatment of missingness; and
  • when MAR is scientifically uncertain and plausible MNAR departures could change the result, perform sensitivity analyses.

The purpose of sensitivity analysis is not to discover the “true” missing values. It is to determine whether the substantive conclusion survives reasonable alternative assumptions.

When Should You Consider Sensitivity Analysis?

Sensitivity analysis deserves particular attention when:

  • substantial outcome data are missing;
  • dropout is plausibly related to unobserved outcomes;
  • the reason for missingness suggests that MAR may be questionable;
  • results are close to a decision or inferential boundary;
  • different reasonable missing-data methods give materially different estimates;
  • a primary scientific conclusion depends strongly on an unverifiable missingness assumption; or
  • the stakes of the inference make robustness to missing-data assumptions important.

Little and Rubin devote explicit attention to sensitivity analyses for MNAR models, including repeated-measures settings and multiple-imputation applications, reflecting the fact that missingness assumptions can be central to inference rather than a minor technical detail (Little & Rubin, 2020).

What Should Researchers Report?

A reproducible missing-data section should make clear what happened between data collection and final inference.

Report, where relevant:

  • missingness by analysis variable;
  • important missingness patterns;
  • the number of observations with any missing analysis data;
  • the number entering complete-case analyses;
  • whether missingness occurred in predictors, outcomes, or both;
  • longitudinal dropout patterns;
  • the missing-data method used;
  • the assumptions under which that method is interpreted;
  • variables and structure used in an imputation model;
  • how multiple-imputation estimates were combined;
  • the likelihood model used for incomplete longitudinal outcomes;
  • sensitivity analyses to alternative assumptions; and
  • whether substantive conclusions changed across approaches.

Avoid writing simply that “missing values were handled using multiple imputation.” The reader needs enough information to understand the assumptions and how uncertainty from missingness entered the analysis.

Common Mistakes

Automatically deleting every incomplete row

Simplicity does not establish validity, and deletion can simultaneously reduce precision and change the population represented by the analysis.

Calling data MAR because missingness is associated with observed variables

Such associations may contradict MCAR, but they do not establish the absence of additional dependence on unobserved values.

Treating imputed values as observed measurements

Imputation predicts missing information; it does not recreate observations with certainty.

Using single imputation with ordinary standard errors

Using single imputation and ordinary standard errors without accounting for imputation uncertainty can fail to represent the additional uncertainty arising from missing values. Multiple imputation was developed specifically to incorporate that uncertainty (Little & Rubin, 2020).

Assuming every available longitudinal measurement solves dropout

The validity of available-data methods still depends on the missingness process and model.

Treating MI as assumption-free

Multiple imputation exchanges one set of assumptions for an explicit model-based strategy; it does not eliminate assumptions.

Do not ignore disagreement between analyses. If complete cases, MI, likelihood analyses, or sensitivity analyses give materially different conclusions, that disagreement is itself important evidence about robustness.

Bottom Line: How to Handle Missing Data

There is no single best method for all missing data in research.

The defensible strategy is to begin with the research target and data structure, not with a software default.

Use the following sequence:

identify pattern → identify variables with missingness → consider plausible mechanism → define analysis target → choose method → conduct diagnostics/sensitivity analysis.

Complete case analysis may sometimes be adequate, but it should be chosen rather than inherited automatically. Multiple imputation can preserve information and account for uncertainty from imputed values when its modeling assumptions are appropriate. Likelihood-based methods can be especially natural for incomplete longitudinal outcomes under suitable MAR assumptions. When conclusions depend materially on assumptions that distinguish MAR from MNAR, sensitivity analysis should make that dependence visible.

The central principle is simple: missing values are part of the statistical model and the scientific interpretation—not merely cells to delete before the “real” analysis begins.

FAQs

What is the difference between MCAR, MAR and MNAR?

MCAR means missingness does not depend on the relevant observed or missing values. MAR allows missingness to depend on observed information but assumes no additional dependence on the missing values after conditioning on that information. MNAR allows residual dependence on missing values themselves. These distinctions concern the missingness process, not simply the visible pattern of blank cells (Little & Rubin, 2020).

Can I test whether my data are MAR?

Observed-data comparisons can provide evidence against some MCAR explanations, but they cannot generally prove MAR because distinguishing MAR from MNAR requires assumptions about values that are not observed. MAR should therefore be stated as an assumption conditional on specified observed information, not as an empirical fact established from the missingness table (Little & Rubin, 2020).

Is complete case analysis always wrong?

No. Its adequacy depends on the estimand, amount and pattern of missingness, missingness mechanism, and information contained in incomplete observations. The problem is using complete-case deletion automatically without examining those conditions (Little & Rubin, 2020).

Why is multiple imputation usually preferable to simple single imputation?

Multiple imputation creates multiple completed datasets and uses variation across them to represent uncertainty about the missing values. Combined inference therefore incorporates both within-imputation and between-imputation uncertainty, whereas ordinary analysis after single imputation can fail to represent that additional uncertainty adequately (Little & Rubin, 2020; Steyerberg, 2019).

Should missing predictors and missing outcomes be handled the same way?

Not necessarily. Missing predictors can remove otherwise informative observations from multivariable or prediction models, while missing outcomes remove direct response information. Repeated-outcome studies add another complication because participants may retain informative measurements before dropout. The method should reflect which variable is missing and its role in the intended analysis (Fitzmaurice et al., 2011; Steyerberg, 2019).

When are likelihood-based methods useful for missing data?

They are particularly useful when the substantive model itself provides a likelihood for incomplete observations—for example, appropriately specified longitudinal repeated-measures or mixed models. Under suitable MAR and model assumptions, these methods can use observed portions of incomplete trajectories without requiring complete follow-up for every participant (Fitzmaurice et al., 2011).

When should I perform an MNAR sensitivity analysis?

Sensitivity analysis should be considered when MAR is scientifically uncertain, missingness or dropout may plausibly depend on unobserved outcomes, missing information is consequential, or the substantive conclusion could change under reasonable departures from the primary missingness assumption (Little & Rubin, 2020).

References

Fitzmaurice, G. M., Laird, N. M., & Ware, J. H. (2011). Applied longitudinal analysis (2nd ed.). Wiley.

Harrell, F. E., Jr. (2015). Regression modeling strategies: With applications to linear models, logistic and ordinal regression, and survival analysis (2nd ed.). Springer.

Little, R. J. A., & Rubin, D. B. (2020). Statistical analysis with missing data (3rd ed.). Wiley.

Steyerberg, E. W. (2019). Clinical prediction models: A practical approach to development, validation, and updating (2nd ed.). Springer.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry