Resource

Prediction Model Recalibration and Updating: What to Do When Performance Changes in a New Population

External validation can reveal changed calibration without implying that the original prediction model should be discarded. This Resource explains how to choose the least extensive update that adequately addresses the observed transport problem, from baseline-risk recalibration through coefficient revision, model extension, and redevelopment.

External validation has identified a problem. Perhaps discrimination remains acceptable, but predicted risks are systematically too high. Perhaps the calibration slope suggests that predictions are too extreme. Or perhaps both calibration and discrimination have deteriorated in the new population.

The next decision is not automatically “discard the model and start again.”

An existing clinical prediction model contains information estimated from its development data. When performance changes in a new population, the appropriate response can range from leaving the model unchanged, through relatively parsimonious prediction model recalibration, to revising coefficients, extending the model with additional predictors, or eventually redeveloping it. Steyerberg describes updating as a hierarchy in which increasingly extensive changes require increasingly strong information from the new data (Steyerberg, 2019).

Practical principle: Use the least extensive update that adequately addresses the observed transport problem, and require stronger evidence before estimating more parameters from the new population.

This resource begins after external validation has been performed. Its purpose is not to explain general prediction-model validation, but to decide what to do once validation shows that the original model does not transport perfectly.

Start by Diagnosing What Changed

Before updating a model, separate three questions:

  1. Has baseline outcome risk changed?
  2. Have the average predictor effects changed?
  3. Has the model lost useful discrimination or omitted important predictive information in the new population?

These are not interchangeable problems.

External validation should distinguish discrimination from calibration. Discrimination concerns how well the model separates patients with different outcomes; calibration concerns agreement between predicted and observed outcomes. A model can retain useful discrimination while its absolute risk estimates become inaccurate (Harrell, 2015; Steyerberg, 2019).

That distinction matters because a model that still ranks patients effectively but systematically overpredicts risk presents a very different updating problem from a model whose predictor effects no longer transport.

Calibration-in-the-large: did baseline risk change?

Calibration-in-the-large assesses systematic overprediction or underprediction at the population level. For binary outcomes, it addresses whether predicted risks are systematically too high or too low relative to the observed outcome frequency.

A difference in baseline outcome risk between development and validation populations can produce this pattern without implying that every predictor coefficient needs to be replaced (Steyerberg, 2019).

This is precisely the situation in which a baseline or intercept update may be more defensible than rebuilding the entire model.

Calibration slope: did the overall strength of the predictor effects change?

The calibration slope assesses the average strength of the model's predictor effects in the validation population. An ideal slope is 1.

A slope below 1 indicates that predictions are too extreme on average: low predicted risks tend to be too low and high predicted risks too high. It therefore indicates that the predictor effects need to be reduced on average (Steyerberg, 2019). Harrell similarly describes the characteristic calibration pattern of overly extreme predictions produced by overfitting (Harrell, 2015).

External-validation qualification: A calibration slope below 1 in a new population should not automatically be attributed entirely to overfitting during original development. It can reflect both original overfitting and genuine differences in predictor effects between the development and validation settings (Steyerberg, 2019).

Discrimination: is the model still separating patients?

Discrimination asks whether the model still assigns relatively higher predictions to patients with worse outcomes. It does not establish that the absolute predicted probabilities are correct.

Preserved discrimination + poor calibration often supports considering recalibration before abandoning the model.

By contrast, deterioration in discrimination suggests that the problem may extend beyond baseline risk alone. The predictors or their effects may not separate outcomes as effectively in the new population, and more substantial revision may need consideration (Steyerberg, 2019).

Validation, Recalibration, Updating, and Redevelopment Are Different Tasks

These terms should not be treated as synonyms.

How common prediction-model adaptation tasks differ
Task What happens to the existing model? Main purpose
External validation Nothing is changed Measure how the original model performs in new data
Recalibration Limited parameters governing absolute risk and/or overall predictor-effect strength are adjusted Correct identifiable calibration problems while retaining most of the original model
Model updating/revision Some existing model parameters are modified using new data Adapt predictor effects when simple recalibration is insufficient
Model extension New predictors or model components are added Capture additional predictive information supported in the new setting
Redevelopment A substantially new model is estimated Replace the original model when its structure or predictive information is no longer adequate

The distinction is important for reporting. Performance of the unchanged original model is the external validation result. Performance after fitting recalibration or updating parameters in those data is performance of an updated model, not evidence that the original model transported unchanged.

Steyerberg's updating framework moves from relatively parsimonious adjustments of baseline risk and calibration slope toward more extensive parameter re-estimation. The choice should depend on both the type of performance problem and the amount of information available for updating (Steyerberg, 2019).

The Prediction Model Updating Ladder

A useful strategy is to move upward only when the evidence shows that the preceding, less extensive step is inadequate.

1. Validate unchanged

Establish how the original prediction rule performs before modifying it.

2. Adjust baseline/intercept

Correct systematic differences in the overall level of predicted risk while preserving the original relative predictor effects.

3. Adjust calibration slope

Apply a common adjustment when the overall strength of the original predictor effects is mismatched to the new population.

4. Revise coefficients

Allow individual predictor effects to change when simpler recalibration is insufficient and the new data support more extensive estimation.

5. Extend the model

Add supported predictors when the original model retains useful structure but important predictive information is missing.

6. Consider redevelopment

Move toward a substantially new prediction model when the original structure is no longer an adequate starting point and sufficient new data are available.

Validate unchanged → adjust baseline/intercept → adjust calibration slope → revise coefficients → extend model → consider redevelopment

Each step spends more information from the new dataset and moves farther away from the original model.

Step 1: Validate the Original Model Unchanged

Always establish the performance of the original prediction rule before modifying it.

Apply the published model as specified and examine, at minimum, whether discrimination and calibration transport. For calibration, calibration-in-the-large, the calibration slope, and a graphical calibration assessment provide complementary information (Steyerberg, 2019).

Stay at this step when

No clinically or statistically important deterioration requiring correction is apparent, and uncertainty around the performance estimates does not suggest a clear need for adaptation.

An imperfect-looking calibration curve should not automatically trigger model reconstruction. Validation estimates themselves are uncertain, particularly in small samples.

Move upward when

There is credible evidence of systematic miscalibration or other deterioration that matters for the intended predictions.

The pattern of failure determines the next step.

Step 2: Adjust the Baseline Risk or Intercept

Suppose discrimination is preserved and the calibration slope is reasonably compatible with the original predictor effects, but predicted risks are systematically too high or too low.

This pattern points first toward a baseline-risk problem.

For a logistic prediction model, an intercept update can adjust the overall level of predicted risk while retaining the original relative predictor effects. Analogous baseline-risk updating principles apply to other model structures.

Steyerberg presents baseline-risk adjustment as one of the more parsimonious updating approaches when a model's general predictive structure remains useful but outcome incidence differs in the new population (Steyerberg, 2019).

Evidence that supports this step

Move from unchanged validation to baseline recalibration when the evidence suggests:

  • systematic overprediction or underprediction;
  • a difference in baseline outcome risk between populations;
  • relatively preserved discrimination; and
  • no compelling evidence that individual predictor effects require extensive revision.

This approach deliberately preserves information from the original model instead of treating a shift in average outcome risk as evidence that every coefficient has failed.

Do not stop here when

Adjusting baseline risk does not resolve important calibration problems, particularly when the calibration slope substantially differs from 1 or the calibration curve shows that the problem involves the spread of predictions rather than merely their overall level.

Step 3: Adjust the Calibration Slope

A more extensive recalibration allows the overall strength of the original linear predictor to change.

This is relevant when the external calibration slope indicates that the original predictor effects are collectively too strong or too weak for the new population.

When the slope is below 1, predictions are too extreme on average. Multiplying the original predictor effects by an estimated calibration factor provides a form of uniform shrinkage of those effects, while an intercept or baseline component can simultaneously correct the overall risk level (Steyerberg, 2019).

Evidence that supports this step

Move beyond an intercept-only update when:

  • calibration-in-the-large alone does not explain the miscalibration;
  • the calibration slope provides evidence that the overall predictor-effect strength differs from that required in the new population; and
  • there remains reason to preserve the relative structure of the original predictor effects rather than estimate each coefficient separately.

This distinction is important. A slope update assumes that the original effects can be adjusted collectively. It does not claim that each predictor has been individually re-estimated correctly for the new population.

Shrinkage is particularly relevant here

A slope below 1 is associated with predictions that are too extreme and indicates a need to reduce predictor effects on average. Steyerberg connects calibration slopes with shrinkage, while Harrell emphasizes shrinkage and validation as mechanisms for controlling the instability and optimism associated with overfitting (Harrell, 2015; Steyerberg, 2019).

In external validation, however, the estimated slope can reflect both original overfitting and differences in predictor effects between settings. The reason for the slope change should therefore not be oversimplified (Steyerberg, 2019).

Step 4: Revise Individual Predictor Coefficients

The next step is qualitatively different.

Rather than modifying the overall baseline risk or applying one common adjustment to predictor effects, model revision permits individual predictor effects to change.

This may be reasonable when the external evidence suggests that the original model's predictor relationships themselves do not adequately transport.

Evidence that should justify coefficient revision

Moving to individual coefficient revision requires stronger evidence than moving to intercept or slope recalibration.

Look for a combination of:

  • residual miscalibration after simpler updating;
  • evidence that predictor effects differ meaningfully between the original and target populations;
  • sufficient information in the new data to estimate those differences with acceptable stability; and
  • a substantive reason to believe that heterogeneous coefficient changes are real rather than sample noise.

The calibration slope alone does not identify which individual coefficients should change. It summarizes average predictor-effect strength.

Why small validation samples are dangerous here

The more parameters that are re-estimated, the more heavily the updated model depends on the new sample.

Steyerberg cautions that re-estimating all regression coefficients in a small validation dataset can replace comparatively stable estimates from the original development dataset with highly variable new estimates. Parsimonious updating is therefore attractive when the new data contain limited information, whereas extensive revision requires stronger data support (Steyerberg, 2019).

Decision rule: Do not confuse evidence that the original coefficients are imperfect with evidence that a small external sample can estimate better coefficients.

Consider shrinkage when revising coefficients

Once multiple parameters are estimated from new data, overfitting becomes a renewed concern. Updating is itself a modeling exercise.

Shrinkage can therefore be relevant when revised effects are estimated, particularly when the amount of information is modest relative to the complexity of the update. Harrell emphasizes that predictive modeling procedures must account for overfitting and instability rather than treating fitted coefficients as though they were known without error (Harrell, 2015).

Step 5: Extend the Model With New Predictors

Sometimes the original predictors retain useful information, but the target population contains additional information that could materially improve prediction.

In that situation, model extension can preserve the existing prediction model while adding one or more new predictors rather than replacing the entire structure.

Evidence that should justify extension

Adding predictors should require more than observing that an additional variable has a statistically significant association with the outcome.

The relevant question is predictive: Does the new predictor provide reproducible information beyond the existing model that improves prediction for the target setting?

Extension is most defensible when:

  • the original model retains useful predictive structure;
  • a plausible new predictor contains additional predictive information;
  • the new dataset provides enough information to estimate its contribution;
  • the expanded model improves the performance dimension that actually needs improvement; and
  • the added complexity can be validated rather than judged only from apparent performance.

Adding predictors simply because the external model is miscalibrated is not a substitute for diagnosing whether the problem is baseline risk, overall coefficient strength, specific predictor effects, or genuinely missing predictive information.

Step 6: Consider Redevelopment

Redevelopment is the most extensive option.

It means moving away from incremental adaptation of the existing model toward estimation of a substantially new prediction model.

That may eventually be justified—but it should not be the reflex response to every disappointing external calibration result.

Evidence that should justify redevelopment

Redevelopment becomes more defensible when evidence indicates that:

  • simpler recalibration cannot adequately restore performance;
  • important predictor effects differ in ways that cannot reasonably be handled by limited revision;
  • discrimination or broader predictive performance has deteriorated substantially;
  • important predictive information is absent from the original model;
  • the intended population or measurement environment differs enough that the original model structure is no longer a useful starting point; and
  • sufficient new data are available to support development and validation of a new model.

The critical last condition is easily overlooked.

If the external dataset is too small to support extensive coefficient updating, it is unlikely to become adequate merely because the analysis is relabelled “redevelopment.”

How Much Updating Does the Evidence Support?

The updating ladder can be summarized as an evidence-escalation framework.

Evidence-escalation framework for prediction model updating
External finding First strategy to consider What would justify moving farther?
Acceptable discrimination and calibration Keep model unchanged Reproducible evidence of important performance deterioration
Systematic over-/underprediction with otherwise preserved structure Update baseline/intercept Calibration problems remain after correcting overall risk
Predictor effects collectively too strong/weak Update intercept + calibration slope Evidence that a common effect adjustment is insufficient
Specific predictor effects appear nontransportable Revise selected coefficients Adequate data and evidence supporting heterogeneous effect changes
Existing model omits important target-population information Extend with supported predictors Demonstrated incremental predictive information and adequate validation
Original structure no longer provides adequate prediction Redevelop Strong evidence simpler updates are inadequate plus sufficient development data

The progression is intentionally conservative.

Every step upward estimates more information from the new data. That can increase flexibility, but it also increases the opportunity to fit sample-specific noise.

Do Not Let a Small Updating Dataset Erase Useful Prior Information

The original model was typically estimated from a development dataset containing information about predictor–outcome relationships. Updating does not make that information disappear.

Replacing all original coefficients with unrestricted estimates from a small external sample can therefore be statistically inefficient and unstable. Steyerberg specifically emphasizes the attraction of parsimonious updating when validation data are limited and cautions against extensive re-estimation in small samples (Steyerberg, 2019).

This is one of the main reasons for using an updating ladder rather than a binary choice between:

“Use the original model unchanged”

“Build an entirely new model.”

Intermediate strategies allow the data to correct the parts of the model for which transport problems are supported while preserving information that does not need to be discarded.

Differences in Baseline Risk Are Not the Same as Differences in Predictor Effects

This distinction should drive the updating decision.

A new hospital, country, calendar period, referral system, or patient population may have a different overall outcome frequency. That difference can cause systematic overprediction or underprediction even if relative predictor effects remain broadly transportable.

That is principally a baseline-risk calibration problem.

A calibration slope that differs from 1 instead suggests that the original predictor effects are collectively mismatched to the new data. In external validation, this can arise from development overfitting, genuine population differences in predictor effects, or both (Steyerberg, 2019).

Still more complex discrepancies may indicate that some effects differ more than others.

The updating ladder therefore maps progressively different hypotheses:

Baseline risk changed → overall effect strength changed → individual effects changed → relevant predictive information is missing → original model structure is inadequate.

Each hypothesis requires more evidence before more parameters are estimated.

Reassess Performance After Every Update

Updating a model does not end the evaluation process.

Once parameters have been estimated using the new data, the resulting predictions are partly adapted to those observations. Performance calculated naively in the same data can consequently be optimistic.

Harrell and Steyerberg emphasize the broader principle that predictive performance should be assessed in a way that accounts for the modeling process and its potential optimism. Data-dependent modeling steps need to be reflected in validation rather than treating the final fitted equation as if it had been prespecified (Harrell, 2015; Steyerberg, 2019).

After updating, reassess:

  • calibration-in-the-large;
  • calibration slope;
  • the calibration curve;
  • discrimination; and
  • other performance measures relevant to the intended prediction task.

More extensive updates deserve particularly careful validation because more information has been learned from the updating sample.

Where possible, subsequent external evaluation in further data provides evidence about whether the updated model, rather than merely the original model, transports.

Report the Original and Updated Models Separately

A clean report should preserve the chronology of the analysis.

First report the performance of the unchanged external validation model.

Then report:

  1. what performance problem was identified;
  2. why that pattern motivated the chosen level of updating;
  3. exactly which parameters were changed or added;
  4. whether shrinkage or other safeguards against overfitting were used;
  5. performance after updating; and
  6. how the updated performance was validated.

Do not retrospectively replace the original external-validation results with the improved fit after recalibration.

If recalibration was necessary, the correct conclusion is not:

“The original model was well calibrated externally.”

It is:

“The original model showed external miscalibration and was subsequently updated for the target population.”

That distinction separates validation from adaptation.

Common Mistakes After External Validation

Rebuilding immediately because calibration is imperfect

Miscalibration does not establish that the entire model is useless. First determine whether the problem can be explained by baseline risk or overall predictor-effect strength.

Updating only because the AUC changed

Discrimination and calibration are separate performance dimensions. Diagnose what changed before deciding what to modify (Harrell, 2015; Steyerberg, 2019).

Using intercept recalibration to solve a slope problem

An intercept adjustment changes the overall risk level. It does not, by itself, correct predictions whose spread is systematically too extreme or too narrow.

Treating a calibration slope below 1 as proof of original overfitting

In external data, the slope can reflect both development overfitting and genuine differences in predictor effects between populations (Steyerberg, 2019).

Re-estimating every coefficient in a small validation dataset

Greater flexibility does not guarantee better transportability. Extensive re-estimation can replace stable original information with noisy new estimates (Steyerberg, 2019).

Adding predictors before identifying the source of miscalibration

A new predictor is not automatically the solution to a shifted baseline event rate or a calibration-slope problem.

Reporting updated performance as external validation of the original model

Once the external data have been used to estimate updating parameters, they have become part of model adaptation. Keep unchanged-model validation and updated-model performance conceptually and numerically separate.

Practical Prediction Model Recalibration Checklist

  • Apply and evaluate the original model unchanged before fitting updates.
  • Separate discrimination from calibration.
  • Examine calibration-in-the-large for systematic overprediction or underprediction.
  • Examine the calibration slope for mismatch in overall predictor-effect strength.
  • Use a calibration plot to identify patterns not captured by summary measures.
  • Ask whether differences are primarily in baseline risk, predictor effects, or both.
  • Prefer an intercept/baseline-risk adjustment when the evidence supports a baseline shift alone.
  • Consider calibration-slope adjustment when predictor effects need a common correction.
  • Require stronger evidence and adequate data before revising individual coefficients.
  • Consider shrinkage when estimating multiple updating parameters.
  • Add new predictors only when their additional predictive contribution is supported.
  • Avoid extensive re-estimation merely because a small validation dataset is available.
  • Reserve redevelopment for situations in which simpler updating is inadequate and sufficient new data support a new development process.
  • Reassess discrimination and calibration after updating.
  • Clearly distinguish original external-validation performance from updated-model performance.

Bottom Line

Prediction model recalibration should be treated as an evidence-based escalation problem, not an all-or-nothing decision.

External miscalibration does not automatically invalidate everything learned by the original model. First identify whether the problem concerns baseline risk, the average strength of predictor effects, specific predictor effects, missing predictive information, or the broader model structure.

Then climb the updating ladder only as far as the evidence supports:

Validate unchanged → adjust baseline/intercept → adjust calibration slope → revise coefficients → extend model → consider redevelopment.

The farther you move along that sequence, the more parameters you estimate from the new population and the stronger the data requirements become. In limited updating samples, parsimonious recalibration can preserve useful information from the original model while correcting identifiable transport problems; unrestricted re-estimation can instead introduce new instability (Steyerberg, 2019).

Finally, updating creates a new prediction rule. Reassess its performance. A model that has been recalibrated, revised, or extended should not inherit the validation status of either the original model or the apparent performance obtained during updating.

FAQs

What is prediction model recalibration?

Prediction model recalibration adapts an existing model when its absolute predictions no longer agree adequately with outcomes in a target population. Relatively parsimonious approaches can adjust baseline risk or the overall strength of the original predictor effects while preserving most of the original model (Steyerberg, 2019).

When should I update only the intercept of a clinical prediction model?

An intercept or baseline-risk update is most appropriate to consider when the model systematically overpredicts or underpredicts risk in the new population while its relative predictor structure remains reasonably transportable. A difference in baseline outcome risk is an important reason to investigate this type of recalibration (Steyerberg, 2019).

What does the calibration slope tell me about model updating?

The calibration slope summarizes whether the model's predictor effects have the appropriate overall strength in the validation population. A slope below 1 indicates predictions that are too extreme on average and supports considering a common reduction in predictor-effect strength. In external validation, however, such a slope can reflect both original overfitting and differences in predictor effects between populations (Steyerberg, 2019).

Is recalibration the same as rebuilding a prediction model?

No. Recalibration preserves most of the existing prediction structure while modifying limited components such as baseline risk or overall predictor-effect strength. Revision estimates more model parameters, extension adds predictive information, and redevelopment constructs a substantially new model.

Should I re-estimate all coefficients when an external model is miscalibrated?

Not automatically. Extensive coefficient re-estimation requires considerably more information than intercept or slope recalibration. In a small updating sample, replacing the original coefficients with unrestricted new estimates can increase variability and instability, making parsimonious updating preferable when it adequately addresses the problem (Steyerberg, 2019).

When should new predictors be added during model updating?

New predictors should be considered when there is reason to believe that the existing model omits predictive information important in the target population and the new data adequately support estimating and validating that additional contribution. Adding predictors should not substitute for correcting a simpler baseline-risk or calibration-slope problem.

Does a recalibrated model need to be validated again?

Yes. Once the new data are used to estimate recalibration or updating parameters, the resulting model has been adapted to those observations. Its discrimination and calibration should be reassessed, with appropriate attention to optimism from the updating procedure, and further external evaluation is valuable for establishing transportability of the updated prediction rule (Harrell, 2015; Steyerberg, 2019).

References

Harrell, F. E., Jr. (2015). Regression modeling strategies: With applications to linear models, logistic and ordinal regression, and survival analysis (2nd ed.). Springer. https://doi.org/10.1007/978-3-319-19425-7

Steyerberg, E. W. (2019). Clinical prediction models: A practical approach to development, validation, and updating (2nd ed.). Springer. https://doi.org/10.1007/978-3-030-16399-0

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry