Resource

Interaction vs Effect Modification: How to Interpret Subgroup Differences Correctly

Interaction, effect modification, and subgroup significance are not interchangeable. Learn how effect scale, regression product terms, causal framing, and direct comparisons between subgroup effects determine the correct interpretation.

Researchers often use interaction, effect modification, subgroup analysis, and heterogeneity of effects as though they were interchangeable. They are not. The distinction matters because the same data can show effect modification on one effect-measure scale but not another, a regression interaction term is defined on the scale of the model’s linear predictor, and a subgroup-specific statistically significant result does not by itself show that effects differ between subgroups.

The practical question is not simply:

“Was the treatment significant in subgroup A but not subgroup B?”

It is: What effect is being estimated, on what scale, in which population or subgroup, and is the analysis asking about statistical heterogeneity or a causal interaction between interventions?

This resource provides a decision framework for interpreting interaction vs effect modification without confusing subgroup-specific effects, statistical interaction, or causal interaction.

Quick Decision Guide

Question What to examine
Does an exposure or treatment effect differ across levels of another variable? Subgroup-specific effect estimates on a clearly specified scale
Does the difference persist on the risk-difference scale? Additive effect modification/interaction
Does it persist on a ratio scale? Multiplicative effect modification/interaction
Does a regression model contain a product term? Statistical interaction on the model’s specified scale
Is the claim about joint intervention on two causal factors? Causal interaction, requiring an explicitly causal framework
Is one subgroup significant and another nonsignificant? Do not infer interaction from this alone; evaluate the difference between effects
Is the modifier continuous? Preserve meaningful continuous information where possible; do not create arbitrary subgroups merely to simplify interpretation
Are many possible interactions being searched? Consider prespecification, scientific rationale, model complexity, and overfitting

1. Interaction and Effect Modification Are Not Interchangeable

Effect modification describes heterogeneity of an effect measure

In epidemiology, effect-measure modification occurs when the magnitude of an effect measure varies across levels of another variable. The effect measure must be named because modification depends on the measure being examined. A risk difference may vary across strata even when a risk ratio does not, and vice versa (Lash et al., 2021; Hernán & Robins, 2020).

This is why “effect modification” without specifying the scale is incomplete. Researchers should ask whether there is modification of the risk difference, risk ratio, odds ratio, or another explicitly defined measure.

Effect-measure modification is also not a bias that should automatically be removed. Modern Epidemiology distinguishes it from confounding: confounding is a source of bias investigators generally seek to prevent or control, whereas effect-measure modification is a feature of the effect under study that may itself be scientifically important and worth reporting (Lash et al., 2021).

Causal interaction asks a different question

Hernán and Robins distinguish effect modification from interaction as a causal concept. Their causal definition of interaction concerns two treatments or interventions and therefore requires considering the outcomes under their joint interventions. They explicitly note that the terms effect modification and interaction are sometimes used synonymously in scientific writing even though the concepts are related but distinct (Hernán & Robins, 2020).

Effect modification

Does the effect of treatment or exposure (A) vary across strata defined by (V)?

Causal interaction

What happens under joint interventions on (A) and another treatment or exposure (E)?

When the second variable is itself a randomized treatment, Hernán and Robins show that the concepts can coincide on the specified effect scale. Outside that setting, moving from observed heterogeneity to a causal-interaction claim requires the assumptions needed to identify the relevant causal effects (Hernán & Robins, 2020).

Modern Epidemiology makes a similar distinction between statistical interaction, effect heterogeneity, and causal interaction. For example, a multiplicative interaction parameter from a regression model requires appropriate absence-of-confounding assumptions before it can be interpreted as a causal interaction parameter (Lash et al., 2021).

Practical rule: Do not upgrade a regression product term or subgroup pattern into a causal interaction merely because the word interaction appears in the statistical output.

2. The Scale Matters

The most important technical point in interaction vs effect modification is that heterogeneity is scale-dependent.

Additive and multiplicative effect modification can disagree

Suppose an intervention has the same causal risk difference in two strata but different causal risk ratios. There is then no additive effect modification but there is multiplicative effect modification. Hernán and Robins demonstrate precisely this possibility and emphasize that heterogeneity on one scale does not imply heterogeneity on another (Hernán & Robins, 2020).

Modern Epidemiology similarly emphasizes that epidemiologists commonly work with both differences and ratios and that their heterogeneity can behave differently across strata. Ratio and difference measures can even vary in opposite directions (Lash et al., 2021).

A statement such as “There was effect modification” should usually specify the scale.

For example: “The estimated treatment effect varied across subgroups on the risk-difference scale.”

Or: “There was evidence of multiplicative interaction on the risk-ratio scale.”

The scale is part of the finding.

Which scale should be reported?

There is no justification for selecting a scale merely because software makes it convenient.

Modern Epidemiology recommends presenting both additive and multiplicative measures of interaction in general. The text notes that multiplicative interaction is often reported because common regression software produces it conveniently, whereas obtaining additive interaction measures may require additional calculations. That convenience is not itself a scientific reason to prefer the multiplicative scale (Lash et al., 2021).

The additive scale is especially informative when the question concerns absolute public-health impact. If treatment produces a larger risk reduction in one subgroup on the difference scale, treating the same number of individuals in that subgroup would prevent or cure more outcomes, assuming the causal interpretation is warranted (Lash et al., 2021; Hernán & Robins, 2020).

A ratio-scale interaction can answer a different question. It should not automatically be translated into a claim about larger absolute benefit.

Report the underlying risks when possible

Hernán and Robins argue that reporting absolute counterfactual risks within levels of the potential effect modifier can be more informative than reporting only their ratios or differences (Hernán & Robins, 2020).

Likewise, Modern Epidemiology recommends presentations that make the joint exposure pattern visible, including stratum-specific effects, additive and multiplicative interaction measures, uncertainty, and—where available in cohort settings—the actual risks (Lash et al., 2021).

Practical reporting sequence:

subgroup risks → subgroup-specific effects → effect scale → contrast between effects → uncertainty → substantive interpretation

This is more informative than reporting only an interaction-term p-value.

3. How Interaction Is Represented in Regression

A product term represents departure from additivity on the model’s linear-predictor scale

Regression models commonly represent statistical interaction by including product terms.

For two predictors X and Z, a simple model can be written as:

g{E(Y | X,Z)} = β0 + β1X + β2Z + β3XZ.

The product term XZ allows the modeled relationship between X and the outcome to depend on Z.

Harrell describes standard generalized regression models as additive on their linear-predictor scale unless interaction terms are included. Thus, the meaning of “no interaction” depends on the model and link being used (Harrell, 2015).

This point is crucial in generalized models. In logistic regression, for example, the linear predictor is the log-odds scale. A product term naturally tests departure from additivity on that scale, corresponding to a multiplicative odds interpretation after exponentiation. It does not automatically test additive interaction in risks.

Modern Epidemiology notes that standard logistic-regression software conveniently produces multiplicative interaction estimates, which helps explain why multiplicative interaction is reported more often than additive interaction (Lash et al., 2021).

Lower-order coefficients become conditional

Once a product term is included, the lower-order coefficient for X no longer represents one universal effect of X. It represents the modeled effect when the interacting variable Z is at its reference or zero value.

For a linear specification:

Y = β0 + β1X + β2Z + β3XZ,

the modeled X effect is:

∂Y/∂X = β1 + β3Z.

Thus, β3 describes how the X effect changes with Z, while β1 is the X effect specifically at Z = 0.

The interpretation of zero and of reference categories therefore matters. Interaction should be interpreted through meaningful contrasts or predicted effects rather than by reading the product-term coefficient in isolation.

Continuous variables should not be turned into subgroups automatically

A continuous candidate effect modifier does not have to be divided into “low” and “high” groups.

Harrell emphasizes the disadvantages of categorizing continuous predictors: categorization discards information and imposes artificial discontinuities and flat relationships within categories. His regression strategy instead supports flexible continuous modeling, including spline functions when simple linearity is inadequate (Harrell, 2015).

This matters for interaction analysis because a continuous modifier can be represented continuously. If either predictor has a nonlinear relationship with the outcome, the interaction structure may also require greater flexibility than one simple linear-by-linear product.

Do not create arbitrary subgroups merely because subgroup-specific coefficients look easier to report.

Preserve continuous information where scientifically and statistically appropriate, specify the functional forms deliberately, and interpret effects over meaningful values or contrasts.

Interaction terms increase model complexity

Interactions are not cost-free additions to a model.

Harrell treats interactions and nonlinear terms as important components of model specification while also emphasizing the need for parsimony. Increasing model flexibility consumes information and increases the opportunity for overfitting, particularly when model structure is chosen through data-dependent searching (Harrell, 2015).

Consequently, an analysis that tries every plausible subgroup, every pairwise interaction, multiple cutpoints, and several functional forms can generate an unstable collection of apparent differences.

Interactions with a clear scientific rationale should therefore be distinguished from exploratory interaction searches. Prespecification is especially valuable when the interaction is intended to support a confirmatory scientific conclusion rather than generate a hypothesis.

4. How to Interpret Subgroup Effects

Start with the estimand, not the subgroup p-value

Suppose a treatment effect is estimated separately for two groups.

The first question is not: “Which subgroup has p < .05?”

The first questions are:

  1. What effect measure is being estimated?
  2. What is the estimated effect in each subgroup?
  3. How uncertain is each estimate?
  4. How different are the subgroup-specific effects on the chosen scale?
  5. Is that heterogeneity scientifically important?
  6. How compatible is the observed difference with random variability?

Modern Epidemiology explicitly frames stratified analysis in this way: after calculating stratum-specific estimates, investigators must consider both whether the variation is scientifically or publicly important and the extent to which it is compatible with random statistical fluctuation (Lash et al., 2021).

Hernán and Robins likewise note that observed heterogeneity between stratum-specific estimates may reflect either true effect-measure modification or sampling variability, with finer stratification generally increasing uncertainty (Hernán & Robins, 2020).

“Significant here, nonsignificant there” is not a test of interaction

A subgroup-specific test asks whether the effect in that subgroup is compatible with its null value under the specified testing procedure.

Interaction asks a different question:

Do the effects differ from each other on the specified scale?

Those are not the same hypothesis.

Therefore, the pattern:

  • subgroup A: statistically significant;
  • subgroup B: not statistically significant

does not by itself establish interaction or effect modification.

The two subgroup estimates may be similar while having different standard errors, sample sizes, or precision. The correct inferential target for interaction is the contrast between the subgroup-specific effects—for example, a difference in risk differences, a ratio of risk ratios, or the corresponding interaction parameter from an appropriately specified model.

This follows directly from the distinction in the approved sources between variation in stratum-specific effect estimates and the need to determine whether that variation exceeds what can reasonably be attributed to random fluctuation (Lash et al., 2021; Hernán & Robins, 2020).

Report effect estimates and uncertainty, not just subgroup labels

A useful subgroup presentation should make the pattern inspectable.

Modern Epidemiology recommends reporting effect estimates across the joint exposure strata, subgroup-specific effects, interaction measures on additive and multiplicative scales, and confidence intervals and p-values for interaction measures where relevant (Lash et al., 2021).

Avoid

“Treatment worked in younger patients but not older patients.”

Prefer

“The estimated treatment effect was [effect, CI] in the younger subgroup and [effect, CI] in the older subgroup. The estimated between-subgroup contrast on the prespecified [risk-difference/risk-ratio/etc.] scale was [interaction estimate, CI].”

Only after presenting those quantities should the substantive importance of the difference be discussed.

Distinguish heterogeneity from causality

Even convincing subgroup heterogeneity is not automatically causal.

If the analysis is observational, a subgroup-specific adjusted association remains an association unless the study design, causal estimand, adjustment strategy, and identification assumptions support causal interpretation. Modern Epidemiology specifically notes that interpreting statistical interaction parameters causally requires assumptions about confounding of the relevant exposure effects (Lash et al., 2021).

Inferential framework Appropriate wording
Statistical “The modeled association differed by sex on the odds-ratio scale.”
Effect-measure “The estimated risk difference varied across baseline-risk strata.”
Causal “The causal treatment effect differed across strata,” only when the design and assumptions justify that interpretation.
Causal interaction Requires a clearly defined joint-intervention question rather than merely observing that one variable modifies an association.

5. Common Reporting Mistakes

Mistake 1: Treating interaction and effect modification as synonyms

Effect modification concerns variation in an effect measure across strata. Causal interaction concerns joint interventions. Statistical interaction concerns departure from additivity on a specified model scale. These concepts overlap in some settings but should not be treated as universally interchangeable (Hernán & Robins, 2020; Lash et al., 2021).

Mistake 2: Failing to name the effect scale

Effect-measure modification is scale-dependent. A risk difference and risk ratio can show different heterogeneity patterns across the same strata (Lash et al., 2021; Hernán & Robins, 2020).

Better: “There was heterogeneity on the risk-difference scale.”

Mistake 3: Treating a logistic-regression product term as a universal interaction test

A product term in logistic regression naturally concerns the model’s log-odds scale and therefore multiplicative odds. It does not automatically answer whether interaction exists on an additive risk scale (Lash et al., 2021; Harrell, 2015).

Mistake 4: Comparing subgroup significance decisions

Different subgroup significance decisions do not directly test whether the subgroup effects differ. Evaluate the between-effect contrast on the prespecified scale and its uncertainty rather than comparing two binary significance labels (Lash et al., 2021; Hernán & Robins, 2020).

Mistake 5: Reporting only the interaction p-value

The interaction test does not communicate the subgroup risks, effect magnitudes, direction of heterogeneity, or practical importance. Report subgroup-specific estimates and uncertainty together with the interaction measure (Lash et al., 2021).

Mistake 6: Automatically dichotomizing a continuous modifier

Turning a continuous variable into arbitrary categories discards information and imposes artificial discontinuities. Continuous predictors can be retained and, where necessary, represented flexibly rather than forced into “low” and “high” subgroups (Harrell, 2015).

Mistake 7: Searching every possible interaction

Interaction and nonlinear terms add model complexity. Data-driven model searching can create unstable and overfit findings. Scientifically motivated interactions should be distinguished from exploratory searches, and model complexity should be kept compatible with the information available (Harrell, 2015).

Mistake 8: Calling statistical interaction a causal mechanism

A statistical interaction parameter does not automatically establish biological, mechanistic, or causal interaction. Causal interpretation requires the appropriate causal question and assumptions (Lash et al., 2021; Hernán & Robins, 2020).

A Practical Interaction vs Effect Modification Workflow

  1. Define the research question.
    Is the goal descriptive heterogeneity, effect-measure modification, a regression interaction, or causal interaction?

  2. Define the target effect.
    Specify the treatment/exposure contrast, outcome, population, and whether the target is associational or causal.

  3. Choose the effect-measure scale deliberately.
    State whether the analysis concerns risk differences, risk ratios, odds ratios, or another measure. Consider both additive and multiplicative scales when scientifically relevant (Lash et al., 2021).

  4. Estimate subgroup-specific effects.
    Report effect estimates and uncertainty rather than only subgroup-specific p-values.

  5. Estimate the difference between effects.
    Use an interaction parameter or another explicit contrast appropriate to the selected scale.

  6. Inspect absolute risks where possible.
    They can clarify whether apparently different relative effects correspond to meaningful differences in absolute benefit (Hernán & Robins, 2020).

  7. Preserve continuous modifiers where appropriate.
    Avoid arbitrary categorization; consider flexible continuous functional forms when needed (Harrell, 2015).

  8. Separate statistical from causal interpretation.
    A fitted product term is not sufficient evidence of causal interaction.

  9. Consider model complexity and prespecification.
    Interactions should be motivated by the scientific question or explicitly labeled exploratory. Avoid allowing an unrestricted search across many interactions to masquerade as a prespecified analysis (Harrell, 2015).

  10. Report the pattern, not merely the threshold decision.
    Explain which effects differ, on which scale, by how much, with what uncertainty, and what the difference means substantively.

Bottom Line

The central distinction in interaction vs effect modification is not terminology alone. It determines what quantity is being estimated and what conclusion the data can support.

Effect modification describes variation in an effect measure across levels of another variable, and that variation depends on the effect-measure scale. Statistical interaction usually refers to departure from additivity on the scale specified by a statistical model. Causal interaction concerns joint interventions and requires a causal framework and the assumptions needed to identify the relevant causal effects (Hernán & Robins, 2020; Lash et al., 2021).

A significant effect in one subgroup and a nonsignificant effect in another does not by itself demonstrate that the effects differ. The interaction question concerns the contrast between effects, not the separate significance labels attached to each one.

A defensible interpretation therefore follows this sequence:

research question → causal or statistical target → subgroup-specific effects → effect-measure scale → contrast between effects → uncertainty → substantive interpretation → limitations

That sequence keeps subgroup analysis focused on the quantity researchers actually want to compare.

FAQs

What is the difference between interaction and effect modification?

Effect modification describes variation in an effect measure across levels of another variable. Interaction can refer to a statistical model property or, in a causal framework, to the joint effects of two interventions. Hernán and Robins explicitly distinguish causal interaction from effect modification even though the terms are sometimes used synonymously (Hernán & Robins, 2020).

Why does the effect-measure scale matter for interaction?

Because heterogeneity can exist on one scale but not another. An effect may be constant across subgroups as a risk difference while varying as a risk ratio, or vice versa. Interaction or effect modification should therefore always be interpreted with the scale stated explicitly (Lash et al., 2021; Hernán & Robins, 2020).

Does a significant interaction term prove effect modification?

It establishes evidence against the no-interaction restriction represented by that particular statistical model and scale. It does not automatically establish heterogeneity on every other effect scale, nor does it automatically establish causal interaction (Lash et al., 2021; Harrell, 2015).

If an effect is significant in one subgroup but not another, is the interaction significant?

Not necessarily. The subgroup tests ask whether each subgroup effect differs from its own null value. Interaction asks whether the subgroup effects differ from each other on a specified scale. The latter requires an explicit comparison of effects rather than comparison of significance labels (Lash et al., 2021; Hernán & Robins, 2020).

How is interaction represented in regression?

A common approach adds a product term such as XZ to the regression model. Without such terms, standard regression specifications assume additivity on their linear-predictor scale. The meaning of the interaction therefore depends on the model and link function (Harrell, 2015).

Should continuous effect modifiers be divided into high and low groups?

Not routinely. Categorizing continuous predictors discards information and imposes artificial discontinuities. Continuous predictors can instead remain continuous, with flexible functions used when simple linearity is inadequate (Harrell, 2015).

Should interaction analyses be prespecified?

Prespecification is particularly important for confirmatory interaction questions because interactions increase model complexity and searching many possible model terms creates opportunities for unstable, data-dependent findings. Harrell’s broader modeling strategy favors subject-matter-guided specification and cautions against treating data-driven model searches as though the final model had been fixed in advance (Harrell, 2015).

References

Harrell, F. E., Jr. (2015). Regression modeling strategies: With applications to linear models, logistic and ordinal regression, and survival analysis (2nd ed.). Springer.

Hernán, M. A., & Robins, J. M. (2020). Causal inference: What if. Chapman & Hall/CRC.

Lash, T. L., VanderWeele, T. J., Haneuse, S., & Rothman, K. J. (2021). Modern epidemiology (4th ed.). Wolters Kluwer.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry