Not Everyone Had the Event: Kaplan–Meier and Cox Regression With Censored Data
This synthetic case study shows why time-to-event outcomes with right censoring require survival methods rather than simple comparisons of mean follow-up or event status. It demonstrates Kaplan–Meier estimation, the log-rank test, Cox proportional hazards regression, and careful interpretation of hazard ratios.
A clinical team wants to compare two synthetic treatment strategies after hospital discharge. The outcome is not simply whether a participant experiences a cardiovascular readmission. The research question is how long participants remain free of readmission, while recognizing that some participants complete follow-up without the event and others are observed for shorter periods.
That distinction changes the analysis. Survival data require a clearly defined time origin and event, and right censoring occurs when the event time is known only to exceed the participant’s observed follow-up time (Lovric, 2011). Kaplan–Meier estimation can describe survival while retaining censored observations, whereas Cox proportional hazards regression can examine associations between survival time and treatment or other covariates while accommodating censoring (Tabachnick & Fidell, 2013; Dmitrienko & Koch, 2017).
Synthetic-data disclosure: This case study uses entirely synthetic data generated for methodological illustration. It contains no real patients, institutions, treatments, or clinical findings.
The research question
Suppose 640 synthetic adults begin follow-up at hospital discharge. The event is the first cardiovascular readmission after discharge.
Standard care
316 participants
Enhanced follow-up programme
324 participants
Age, sex, and a standardized baseline clinical risk score are also recorded. Recruitment is staggered, so participants enter the synthetic study at different calendar times. Consequently, their maximum possible follow-up differs even before additional censoring is considered.
Is time to first cardiovascular readmission different between the two groups, and does that association persist after adjustment for baseline characteristics?
For survival analysis, the time origin and event must be defined precisely because the outcome is the elapsed time between that origin and the event of interest (Lovric, 2011).
Why the first proposed analyses are inadequate
The initial analyst suggests two simpler approaches.
Compare average observed follow-up
Mean observed follow-up is 9.70 months under standard care and 10.88 months under enhanced follow-up.
Problem: Observed follow-up ends either because the participant experiences the event or because observation stops. The measure therefore mixes event timing with censoring.
Reduce the outcome to event versus no event
By study closure, 52.5% of the standard-care group and 44.1% of the enhanced-follow-up group have experienced the event.
Problem: This discards when events occurred and treats participants with very different amounts of event-free follow-up as equivalent non-events.
A participant followed for 14 months without readmission has not experienced a 14-month event time; all that is known is that the true event time, if it eventually occurs, exceeds 14 months. Right-censored observations contain exactly this kind of partial information (Lovric, 2011).
Likewise, a participant censored after 5 months and another followed event-free for 20 months do not provide the same time-to-event information. Survival methods are designed specifically to retain the timing of events and the follow-up contributed by censored participants (Tabachnick & Fidell, 2013; Lovric, 2011).
Core distinction: The simpler summaries are not necessarily numerically wrong. They answer a less informative question than the time-to-event question posed by the study.
What the time-to-event dataset contains
Each participant contributes two essential pieces of outcome information:
- Observed time: months from discharge until readmission or censoring.
- Event indicator: whether readmission occurred at that observed time.
Total sample
640 synthetic participants
Readmissions observed
309 participants
Censored
331 participants
Thus, slightly more than half the sample has no observed event during available follow-up.
Right censoring does not mean that these participants should be classified as permanent non-events. It means only that their event time exceeds the time for which they were observed.
Censoring assumption: The standard right-censoring framework relies importantly on the censoring process not conveying additional information about the unobserved event time in a way that invalidates the analysis. The basic random-censoring model depends crucially on independence between survival and censoring times (Lovric, 2011).
Kaplan–Meier estimation: retaining censored follow-up
The Kaplan–Meier estimator provides a nonparametric estimate of the survival function: here, the probability of remaining free of cardiovascular readmission beyond a given time. The estimate changes when events occur; censored observations reduce the later risk set rather than being counted as events (Lovric, 2011).
| Time | Standard care event-free survival | Enhanced follow-up event-free survival |
|---|---|---|
| 6 months | 0.695 | 0.764 |
| 12 months | 0.510 | 0.600 |
| 18 months | 0.365 | 0.510 |
At 12 months, estimated event-free survival is 0.510 for standard care and 0.600 for enhanced follow-up. Approximate 95% confidence intervals are 0.450–0.567 and 0.541–0.654, respectively.
At 18 months, the corresponding estimates are 0.365 (95% CI 0.301–0.430) and 0.510 (95% CI 0.445–0.571).
What Kaplan–Meier adds: The estimates show information that an event/no-event table cannot: the separation between the groups develops over follow-up, and participants contribute information for as long as they remain observed and at risk.
Kaplan–Meier methods are commonly used to display survival curves for two or more groups, with group differences assessed using procedures such as the log-rank test (Tabachnick & Fidell, 2013). Clinical-trial analyses likewise use Kaplan–Meier estimates for time-to-event outcomes and log-rank procedures for treatment comparisons (Dmitrienko & Koch, 2017).
Comparing the Kaplan–Meier curves
For these synthetic data, the log-rank comparison gives:
Log-rank result
χ²(1) = 5.79, p = .016.
The result indicates evidence that the two survival distributions differ over the observed follow-up period.
But the log-rank test is still an unadjusted group comparison. It does not estimate an adjusted treatment association after accounting for age or baseline clinical risk. For that purpose, a survival regression model is more useful.
Cox proportional hazards regression
Cox regression relates covariates to the hazard while leaving the baseline hazard unspecified. For a binary predictor, exponentiating the corresponding Cox coefficient gives the hazard ratio comparing the two covariate levels while holding the other included covariates constant (Lovric, 2011).
Unadjusted model
Treatment group alone
Hazard ratio = 0.76
95% CI 0.61–0.95
p = .017
Adjusted model
The adjusted model includes treatment group, age, standardized baseline risk score, and sex.
| Predictor | Hazard ratio | 95% CI | p-value |
|---|---|---|---|
| Enhanced follow-up vs standard care | 0.73 | 0.58–0.91 | .006 |
| Age, per additional year | 1.03 | 1.02–1.04 | <.001 |
| Baseline risk score, per 1-SD increase | 1.44 | 1.29–1.61 | <.001 |
| Sex indicator | 1.27 | 1.01–1.59 | .039 |
Adjustment can matter in survival models when important prognostic covariates are related to the endpoint, and clinical-trial methods commonly use Cox proportional hazards models for adjusted time-to-event analyses (Dmitrienko & Koch, 2017).
Here, adjustment moves the treatment estimate from 0.76 to 0.73, while the confidence interval remains below 1.
How to interpret the hazard ratio
The adjusted treatment hazard ratio is 0.73.
Under the fitted proportional-hazards model, this means that, at a given follow-up time, among participants who remain event-free and under observation, the estimated hazard of first readmission in the enhanced-follow-up group is approximately 27% lower than in the standard-care group, conditional on the included covariates.
Do not interpret this as a risk ratio. The value 0.73 is a hazard ratio. It should not be translated into a claim that the enhanced programme reduces a participant’s cumulative probability of readmission by exactly 27%.
The distinction matters because the hazard describes an instantaneous event rate conditional on surviving event-free to that time, whereas the survival probability describes the chance of remaining event-free beyond a specified time (Lovric, 2011).
This is also why reporting the 12- and 18-month Kaplan–Meier estimates alongside the Cox hazard ratio improves interpretation. The hazard ratio supplies a relative model-based comparison, while the survival estimates provide an absolute time-specific description of the synthetic groups.
What proportional hazards requires
The Cox model assumes that the modeled hazard ratio is proportional over time. The model represents covariate effects through hazard ratios while allowing an unspecified baseline hazard, but the proportional-hazards structure remains an assumption that should be examined rather than taken for granted (Lovric, 2011; Dmitrienko & Koch, 2017).
Tabachnick and Fidell (2013) explicitly include proportionality of hazards among the checks preceding interpretation of Cox regression and illustrate assessing whether covariates interact with time. Lovric (2011) likewise describes model checking through time-by-covariate interactions and residual-based methods.
Synthetic-data limitation: In this demonstration, proportional effects were imposed when the event times were generated. That construction is not a substitute for diagnostics in real data. In an applied study, the analyst should examine whether the treatment or other covariate effects change materially with time.
If the treatment hazard ratio were not reasonably stable over follow-up, reporting one constant hazard ratio could hide that pattern. The model specification and interpretation would then need to reflect the time dependence rather than treating the single coefficient as universally applicable.
What the better analysis changes
The three candidate analyses answer different questions.
| Approach | What it uses | What it loses or requires |
|---|---|---|
| Mean observed follow-up | One observed duration per participant | Confounds event timing with censoring |
| Event/no-event analysis | Whether an event was observed | Discards when the event occurred and how long censored participants were observed |
| Kaplan–Meier / Cox analysis | Event indicator and observed time | Retains censoring and timing, subject to survival-model assumptions |
Central methodological principle: Censoring is not a nuisance to be deleted or converted automatically into “no event.” It is part of the data structure.
A practical analysis sequence
1. Define the question
Specify the time origin, the event of interest, and the time-to-event outcome.
2. Represent censoring correctly
Retain both observed time and the event indicator rather than treating censored observations as permanent non-events.
3. Describe survival
Use Kaplan–Meier estimation to describe event-free survival over follow-up while retaining censored information.
4. Compare groups
Use a procedure such as the log-rank test for an unadjusted comparison of survival distributions.
5. Model covariates
Use Cox proportional hazards regression when an adjusted hazard comparison is required.
6. Examine assumptions
Assess whether the proportional-hazards structure is reasonable before treating a single hazard ratio as stable over follow-up.
Uncertainty matters
The point estimates should not be read in isolation.
The adjusted treatment hazard ratio of 0.73 has a 95% confidence interval from 0.58 to 0.91. The interval expresses substantial uncertainty about the size of the association even though it excludes 1 in this synthetic sample.
The Kaplan–Meier estimates also become less precise later in follow-up as events and censoring reduce the number still at risk. Consequently, late portions of survival curves generally deserve greater caution than early portions supported by more participants.
A p-value from a log-rank or Cox analysis does not tell researchers whether an effect is clinically important. Interpretation should combine relative effects, absolute survival probabilities at meaningful times, interval estimates, study design, and the clinical consequences of the event.
Limitations of the case study
Synthetic evidence only
Every participant and outcome in this example is synthetic. The numerical results demonstrate analysis and interpretation; they are not clinical evidence.
Proportional hazards were imposed
The proportional-hazards structure was built into the synthetic event-generation process. Real clinical data can depart from that assumption, so proportional-hazards diagnostics remain necessary (Tabachnick & Fidell, 2013; Lovric, 2011).
Censoring assumptions matter
Standard right-censoring methods depend on assumptions about how censoring relates to the underlying event process. Informative withdrawal or loss to follow-up can threaten interpretation because a censored participant may then differ systematically in unobserved event risk (Lovric, 2011).
Adjustment does not establish causation
Even an adjusted Cox model does not automatically establish causation. Interpretation depends on study design, treatment assignment, measurement quality, covariate specification, and whether important confounding remains (Dmitrienko & Koch, 2017).
Finally, a hazard ratio is only one summary of a survival experience. Time-specific survival probabilities provide important absolute context and should not be replaced by a single relative effect estimate.
Conclusion
The defining feature of this study is not merely that some participants had a cardiovascular readmission. It is that participants were followed for different lengths of time and many had not experienced the event when observation ended.
Reducing the data to mean follow-up ignores why follow-up stopped. Reducing the outcome to event/no-event ignores when the event occurred.
Kaplan–Meier estimation preserves the evolving event-free experience of each group while accommodating right censoring. The log-rank test can compare the resulting survival distributions. Cox proportional hazards regression then extends the analysis to covariates and expresses group differences through hazard ratios, provided the model’s proportional-hazards structure is appropriate (Tabachnick & Fidell, 2013; Dmitrienko & Koch, 2017; Lovric, 2011).
In the synthetic example, enhanced follow-up is associated with longer event-free survival and an adjusted hazard ratio of 0.73 (95% CI 0.58–0.91). That estimate should be read as a hazard comparison conditional on remaining event-free, not as a risk ratio and not as a claim of a 27% reduction in cumulative risk.
References
Dmitrienko, A., & Koch, G. G. (Eds.). (2017). Analysis of clinical trials using SAS: A practical guide (2nd ed.). SAS Institute.
Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2
Tabachnick, B. G., & Fidell, L. S. (2013). Using multivariate statistics (6th ed.). Pearson.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.