Propensity Scores in Observational Research: What They Can and Cannot Fix
Propensity scores can improve comparability on measured pretreatment covariates and reveal poor overlap in observational studies. This Resource explains what matching, stratification, and weighting can accomplish—and why they cannot establish randomization or eliminate unmeasured confounding.
Propensity scores are widely used to adjust treatment or exposure comparisons in observational research. The attraction is understandable: reduce a potentially large collection of pretreatment covariates to a treatment-assignment probability, then use that score for matching, stratification, weighting, or related adjustment.
The dangerous interpretation is that a successful propensity score matching procedure has somehow converted the observational study into a randomized trial.
It has not.
Propensity-score methods can reorganize or reweight observed data so that treatment groups become more comparable with respect to measured covariates represented in the treatment model. They cannot establish that treatment was randomized, demonstrate that all confounding has disappeared, recover information where treatment groups do not overlap, or correct confounding caused by variables that were never adequately measured.
Causal interpretation still depends on the study design, target estimand, causal assumptions, measurement quality, overlap, and implementation of the analysis (Hernán & Robins, 2020; Lash et al., 2021).
This resource provides a practical framework for deciding what propensity scores can accomplish—and what claims remain unjustified after propensity score analysis.
The Core Idea: Model Treatment Assignment, Not the Outcome
For a binary treatment (A) and pretreatment covariates (L), the propensity score is the conditional probability of receiving treatment given those covariates:
e(L) = P(A = 1 | L)
Gelman et al. describe the propensity score as the probability, as a function of the covariates, that a unit receives treatment. They emphasize its usefulness in observational studies because it provides a scalar summary that can help reveal poor overlap between treatment groups even when the covariate vector is multidimensional (Gelman et al., 2014).
This immediately distinguishes a propensity-score model from an outcome regression.
Propensity-score model
Models treatment or exposure assignment given covariates.
Outcome regression
Models the outcome given treatment and covariates.
Both can be used as components of observational study adjustment, but they attack the adjustment problem from different directions. Hernán and Robins treat outcome regression and propensity-score approaches as related strategies for causal effect estimation rather than as evidence that one approach is universally superior (Hernán & Robins, 2020).
The first question should therefore not be:
“Should I match or weight?”
It should be:
“What causal effect am I trying to estimate, what measured variables are required for exchangeability, and do the data contain adequate treatment-group overlap for that effect?”
What the Propensity Score Is Trying to Balance
The practical appeal of a propensity score is dimensional reduction. Instead of trying to form groups that are identical on every combination of several measured pretreatment variables, investigators can use the estimated treatment probability as an adjustment device.
Under the relevant treatment-assignment assumptions, propensity-score methods can be used to create comparisons in which the distribution of the measured covariates used to construct the score is made more comparable between treatment groups. This is the balancing role that motivates matching, stratification and weighting.
Gelman et al. connect propensity scores directly to the design problem in observational studies. When treated and untreated groups occupy very different regions of covariate space, estimated treatment effects can become highly dependent on outcome-model extrapolation. Estimated propensity scores can expose this lack of overlap and can be used to restrict comparisons to treated and control observations with more similar covariate distributions (Gelman et al., 2014).
The important qualifier is measured.
Balancing observed covariates does not demonstrate balance in variables that were not measured, were omitted from the relevant adjustment strategy, or were measured too poorly to represent the confounding structure adequately.
Propensity Scores Do Not Recreate Randomization
Randomization and propensity-score adjustment are fundamentally different.
In a randomized experiment, the treatment-assignment mechanism itself supplies the basis for treatment exchangeability. In an observational study, treatment assignment generally depends on characteristics of the participants, clinicians, institutions, environments, or other processes.
Causal adjustment therefore requires investigators to identify a sufficient set of measured variables such that treatment can plausibly be regarded as conditionally exchangeable with respect to the relevant counterfactual outcomes. Conditional exchangeability is an identification assumption about the causal data-generating process; it is not something statistical software can prove from the fitted propensity-score model (Hernán & Robins, 2020).
Propensity-score methods can implement an adjustment strategy under conditional exchangeability. They cannot establish that conditional exchangeability is true.
An impressive matched dataset, a well-behaved propensity-score model, or excellent measured-covariate balance does not prove that an important unmeasured common cause of treatment and outcome is absent.
What Propensity Scores Can Do
1. Support matching on measured pretreatment characteristics
Propensity score matching uses the treatment-assignment score to identify treated and untreated observations with similar estimated treatment probabilities.
This can reduce reliance on comparisons between observations that are very different with respect to the measured covariates represented by the score. Gelman et al. describe matching, subclassification, blocking and stratification as organizational strategies for limiting dependence on outcome-model extrapolation by restricting the range of covariate space across which comparisons are made (Gelman et al., 2014).
Matching also changes the analyzed population. Modern Epidemiology notes that matching treated observations to controls according to the propensity score can alter the covariate distribution of the resulting cohort. When controls are selected to resemble the exposed group, for example, the resulting standard can correspond to the covariate distribution of the exposed population (Lash et al., 2021).
Implication: The estimand and target population should be specified before matching rules are chosen.
2. Support stratification or subclassification
Observations can also be grouped into strata defined by similar propensity scores, with treatment comparisons performed or standardized across those strata.
This provides another route for reducing multidimensional covariate differences to comparisons among observations with similar treatment probabilities. But categorizing a continuous propensity score also involves analytical choices; propensity-score stratification should not be treated as an automatic guarantee of adequate adjustment.
3. Support propensity score weighting
Instead of discarding or selecting observations, investigators can weight them according to their estimated treatment probabilities.
In inverse probability weighting, observations whose actual treatment was relatively unlikely given their measured covariates receive greater weight. Under the required assumptions and appropriate implementation, the resulting weighted population can remove the measured dependence of treatment assignment on the covariates used for confounding adjustment (Hernán & Robins, 2020).
For a simple binary treatment, inverse-probability treatment weights are built from probabilities of the treatment actually received. Conceptually, weighting uses observed participants to represent individuals with similar measured covariates whose treatment assignment differed.
Modern Epidemiology notes that weighting avoids having to categorize the propensity score or specify a score effect directly in an outcome model, while also warning that fitted treatment probabilities near zero can make the analysis unstable (Lash et al., 2021).
4. Reveal lack of overlap
One of the most useful functions of propensity score analysis may occur before the treatment effect is estimated.
If treated and untreated observations have little overlap in their estimated treatment probabilities, the study is trying to compare groups for whom the alternative treatment was rarely observed.
Gelman et al. explicitly use propensity-score distributions as a diagnostic for this problem: little or no overlap reveals potential sensitivity of causal-effect estimates to modeling assumptions and extrapolation (Gelman et al., 2014).
A propensity score can therefore reveal that the available data are poorly suited to answering the intended causal question. That is useful information—not a statistical inconvenience to hide.
Positivity and Overlap: Can Both Treatments Actually Be Compared?
For causal identification, positivity requires that the treatment alternatives relevant to the causal contrast have positive probability within the covariate strata represented in the target population (Hernán & Robins, 2020).
This has a direct propensity-score interpretation.
If some participants have treatment probabilities extremely close to 0 or 1, the groups contain little empirical information about what would happen under the alternative treatment for people with those characteristics.
Structural lack of positivity
For some covariate patterns, one treatment may simply not occur. The desired counterfactual comparison is then unsupported by observed treatment variation in that region.
No propensity-score algorithm can manufacture the missing treatment experience.
Near-positivity and extreme estimated probabilities
Even when treatment is theoretically possible, estimated probabilities close to 0 or 1 can create unstable comparisons.
This is particularly visible with inverse probability weighting. A very small probability appears in the denominator of a weight, potentially producing a very large contribution from a small number of observations.
Modern Epidemiology identifies fitted probabilities near zero as a major source of instability in weighting analyses (Lash et al., 2021).
Extreme weights are not merely a cosmetic problem in a table of diagnostics. They can signal weak empirical support for the intended comparison.
Why Measured and Unmeasured Confounding Must Be Separated
Suppose treatment choice depends on disease severity, age and prior treatment history, and all three are adequately measured and appropriately incorporated into the causal adjustment strategy.
A propensity-score method may improve treatment-group comparability with respect to those measured variables.
Now suppose an important clinical characteristic affects both treatment choice and outcome but was never recorded.
The propensity score contains no information about that characteristic.
The same problem arises when a required confounder is recorded but measured badly. Hernán and Robins show that measurement error in a confounder can prevent adjustment from fully removing confounding. Conditioning on an imperfect measured version is not necessarily equivalent to conditioning on the underlying variable needed for exchangeability. Modern Epidemiology likewise treats confounder measurement error or misclassification as a source of residual confounding (Hernán & Robins, 2020; Lash et al., 2021).
A propensity score cannot correct confounding by information that the analysis does not adequately possess.
Increasing the sophistication of the matching algorithm or weighting model does not solve that identification problem.
What Propensity Scores Do Not Prove
A completed propensity-score analysis does not prove any of the following:
“The observational study is now equivalent to a randomized trial.”
No. Treatment was still assigned through the observational treatment process. Propensity-score adjustment operates on measured data after that process; it does not replace the assignment mechanism with randomization.
“There is no remaining confounding.”
No. Balance on measured variables cannot establish balance on unmeasured or inadequately measured common causes of treatment and outcome.
“Conditional exchangeability has been verified.”
No. Conditional exchangeability is a causal identification assumption. Statistical adjustment can be constructed under that assumption but cannot demonstrate that every relevant confounding pathway has been controlled (Hernán & Robins, 2020).
“A good propensity-score model means the treatment groups are adequately comparable.”
Not necessarily. The relevant question is what happened to the distributions of the covariates and whether adequate overlap exists for the intended comparison. Gelman et al. emphasize the usefulness of propensity scores for revealing lack of overlap and model sensitivity rather than treating the fitted treatment model itself as the endpoint of the analysis (Gelman et al., 2014).
“Matching automatically identifies the same causal effect as weighting.”
No. Different implementations can imply different target populations and different weighting of covariate strata. Modern Epidemiology explicitly notes that propensity-score matching can change the cohort covariate distribution and that different approaches may estimate different parameters when effects vary across covariate levels (Lash et al., 2021).
“Extreme weights can be ignored if the sample is large.”
No. Very small fitted treatment probabilities can generate unstable weights and give a small number of observations substantial influence (Lash et al., 2021).
“Good measured-covariate balance proves the causal effect.”
No. Balance assessment is evidence about the success of a particular observed-data adjustment. It is not proof of the untestable causal assumptions required to convert an adjusted association into a causal effect.
Matching, Stratification and Weighting Answer Related—but Not Identical—Questions
There is no universally preferred propensity-score implementation.
The appropriate choice depends on the target estimand, population, treatment structure, overlap and intended analysis.
| Approach | Main role | Key caution |
|---|---|---|
| Propensity score matching | Construct a comparison sample with similar estimated treatment probabilities | Matching can change the population represented by the analysis |
| Propensity-score stratification | Compare or standardize treatment groups within score-defined strata | Categorization may leave residual differences within strata |
| Propensity score weighting | Reweight observations according to treatment-assignment probabilities | Extreme probabilities can create unstable weights |
| Outcome regression | Model the outcome conditional on treatment and covariates | Depends on adequate outcome-model specification |
| Propensity-score + outcome modeling | Combine treatment-assignment and outcome-model information in an analysis | Does not remove the underlying identification requirements |
The important decision is therefore not whether matching is “better” than weighting.
It is whether the chosen procedure corresponds to the causal estimand and population the investigator actually intends to describe.
Outcome Regression Versus Propensity-Score Approaches
Propensity scores do not make outcome modeling obsolete.
An outcome-regression approach specifies the conditional relationship between outcome, treatment and covariates. A propensity-score approach instead models the treatment-assignment process and then uses that model to organize or weight comparisons.
The two approaches place modeling effort in different parts of the observed-data structure.
Gelman et al. highlight an important reason to examine treatment assignment even when outcome regression is planned: when treatment groups occupy substantially different covariate regions, the estimated treatment effect may become highly sensitive to the form of the outcome model. Propensity-score overlap can reveal this vulnerability (Gelman et al., 2014).
Consequently, the choice should not be framed as:
“Propensity scores or regression—which one removes confounding?”
Neither method independently guarantees that.
A more defensible question is:
“Given the target causal effect, measured confounders, treatment-assignment structure and available overlap, what estimation strategy implements the required adjustment with assumptions we can defend?”
Diagnostics: Assess the Comparison the Propensity Score Created
Constructing the score is not the end of the analysis.
After matching, weighting or stratification, investigators should examine whether the resulting treatment groups are actually more comparable with respect to the measured pretreatment variables the adjustment was intended to address.
At minimum, assessment should focus on:
- distributions of important measured pretreatment covariates across treatment groups after adjustment;
- remaining regions of poor treatment-group overlap;
- observations that could not be matched or that contribute disproportionately to the weighted analysis;
- the distribution and extremity of estimated treatment probabilities and, for weighting, the resulting weights;
- whether the resulting analyzed or weighted population still corresponds to the intended target population.
Gelman et al. specifically support inspecting propensity-score overlap as a way to diagnose sensitivity caused by imbalanced observational treatment assignment (Gelman et al., 2014). Modern Epidemiology likewise warns that probabilities near zero can destabilize weighting and that matching and weighting can imply different target distributions (Lash et al., 2021).
What diagnostics can answer
“Did the procedure produce the measured-data comparison we intended?”
What diagnostics cannot answer
“Have we proved that no unmeasured confounding remains?”
Propensity Scores and Missing-Data “Response Propensity” Are Not the Same Problem
The term propensity score also appears in missing-data and survey-nonresponse methods, which can create confusion.
In treatment-effect analysis, the propensity score concerns a treatment or exposure assignment probability:
P(A = 1 | L)
Its role is tied to treatment-group comparability and confounding adjustment.
In missing-data analysis, a response propensity instead describes the probability of response or observation conditional on available variables. Little and Rubin describe response-propensity methods in which the missing-data indicator is modeled using observed covariates; respondents may then be stratified by the estimated response propensity or weighted by its inverse (Little & Rubin, 2020).
The mathematical resemblance does not make the scientific problems identical.
Treatment propensity
Who received which treatment, given measured pretreatment characteristics?
Response propensity
Who remained observed or responded, given variables relevant to the missingness process?
Treatment weighting addresses treatment-assignment comparability. Response or censoring weighting addresses selection created by observation or response. A study may require attention to both, and solving one does not automatically solve the other.
Practical Checklist for Propensity Score Analysis
Before PS analysis
- Define the causal question, treatment/exposure contrast, outcome and target population.
- State the target estimand before selecting propensity score matching, weighting or another adjustment method.
- Identify the measured pretreatment variables required for the intended confounding adjustment using substantive and causal knowledge.
- Check whether those variables were measured with sufficient quality for the intended adjustment.
- Distinguish confounders from mediators, colliders and variables created after treatment.
- Ask whether important common causes of treatment and outcome may be unmeasured.
- Examine whether both treatment alternatives actually occur across the covariate patterns relevant to the target population.
- Decide whether the scientific problem is treatment confounding, missingness/selection, or both; do not substitute response propensity for treatment propensity.
After PS construction
- Examine the treatment-probability distributions in both groups.
- Identify regions with weak or absent overlap.
- Assess treatment-group comparability on the measured pretreatment covariates the procedure was intended to balance.
- After matching, determine which observations and which target population remain represented.
- After weighting, inspect the distribution of weights and identify influential or extreme values.
- Investigate fitted treatment probabilities near 0 or 1.
- Do not treat successful model fitting as evidence that confounding has been eliminated.
- Determine whether the implemented matching, stratification or weighting procedure still corresponds to the intended estimand.
Before causal interpretation
- Reassess whether conditional exchangeability is scientifically plausible given the variables actually measured.
- Confirm that positivity/overlap is adequate for the causal contrast being reported.
- Consider whether confounder measurement error could leave residual confounding.
- Distinguish measured-covariate balance from absence of unmeasured confounding.
- Confirm that the effect estimate refers to the population implied by the actual matching or weighting procedure.
- Report remaining limitations from unmeasured confounding, measurement problems, weak overlap and modeling assumptions.
- Do not describe the adjusted observational comparison as equivalent to randomization.
- Reserve causal wording for analyses whose design, estimand and identification assumptions justify it.
The Interpretation to Aim For
A defensible propensity-score conclusion is not:
“After propensity score matching, the groups were randomized and confounding was eliminated.”
It is closer to:
“The propensity-score procedure improved comparability between treatment groups with respect to the measured pretreatment variables used for adjustment. Causal interpretation additionally depends on adequate control of relevant confounding, measurement quality, positivity/overlap, consistency, and the assumptions of the implemented models.”
That wording preserves what the method accomplished without claiming what the data cannot establish.
Bottom Line
Propensity score matching is an observational study adjustment method, not a substitute for randomization.
Propensity scores summarize treatment assignment conditional on measured covariates. They can support matching, stratification and weighting; help improve measured treatment-group comparability; expose lack of overlap; and reduce reliance on comparisons across very different regions of measured covariate space (Gelman et al., 2014; Lash et al., 2021).
But propensity-score methods cannot recover an unmeasured confounder, repair a badly measured confounder simply by including its recorded proxy, create treatment alternatives where positivity fails, or prove conditional exchangeability. Extreme treatment probabilities can also make inverse probability weighting unstable, and different matching or weighting schemes can correspond to different target populations (Hernán & Robins, 2020; Lash et al., 2021).
Recommended sequence
- Causal question
- Target estimand
- Causal structure
- Measured adjustment set
- Overlap
- Propensity-score method
- Diagnostics
- Effect estimation
- Assumptions
- Causal interpretation
Not: fit propensity score → match → declare the study randomized.
FAQs
What is propensity score matching?
Propensity score matching pairs or otherwise selects treated and untreated observations according to similar estimated probabilities of treatment given measured pretreatment covariates. Its purpose is to construct a more comparable observational treatment contrast on those measured characteristics. It does not reproduce randomized treatment assignment.
Does propensity score matching remove all confounding?
No. It can address confounding represented by adequately measured variables incorporated into an appropriate adjustment strategy. It cannot demonstrate that unmeasured confounding is absent, and poorly measured confounders can leave residual confounding (Hernán & Robins, 2020; Lash et al., 2021).
What is propensity score weighting?
Propensity score weighting uses estimated treatment probabilities to give observations different analytical weights. In inverse probability weighting, observations receiving treatments that were relatively unlikely given their measured covariates receive greater weight. Causal interpretation still requires the relevant exchangeability, positivity and other identification assumptions.
What is the difference between propensity score matching and inverse probability weighting?
Matching constructs a selected comparison sample based on similar treatment probabilities, whereas inverse probability weighting reweights observed individuals according to their treatment probabilities. These procedures can represent different target populations and should not be assumed to estimate identical quantities in every setting (Lash et al., 2021).
Why are extreme propensity scores a problem?
Treatment probabilities close to 0 or 1 indicate weak overlap. For inverse probability weighting, a probability close to zero can generate a very large weight, making the estimate unstable and potentially highly influenced by relatively few observations (Lash et al., 2021).
How should propensity-score balance be assessed?
The relevant assessment is whether treatment groups became more comparable with respect to the measured pretreatment covariates the procedure was intended to address, together with examination of propensity-score overlap and, for weighting, the behavior of the weights. Good measured balance supports the implementation of the adjustment but does not prove absence of unmeasured confounding.
Is propensity score analysis better than outcome regression?
The approved sources do not support a universal ranking. Outcome regression models the outcome conditional on treatment and covariates, whereas propensity-score approaches model treatment assignment and use that information for adjustment. The appropriate strategy depends on the estimand, causal structure, data, overlap and modeling assumptions.
Does a propensity score make an observational study equivalent to a randomized controlled trial?
No. Randomization concerns how treatment was actually assigned. A propensity-score procedure adjusts an observational comparison using measured information. It cannot retroactively randomize treatment or establish that unmeasured treatment-group differences are absent.
Is a response propensity for missing data the same as a treatment propensity score?
No. A treatment propensity concerns the probability of receiving treatment conditional on measured covariates. A missing-data response propensity concerns the probability of response or observation conditional on relevant variables. The two use related probability-weighting ideas but target different selection processes (Little & Rubin, 2020).
References
Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., & Rubin, D. B. (2014). Bayesian data analysis (3rd ed.). CRC Press.
Hernán, M. A., & Robins, J. M. (2020). Causal inference: What if. Chapman & Hall/CRC.
Lash, T. L., VanderWeele, T. J., Haneuse, S., & Rothman, K. J. (2021). Modern epidemiology (4th ed.). Wolters Kluwer.
Little, R. J. A., & Rubin, D. B. (2020). Statistical analysis with missing data (3rd ed.). Wiley.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.