Resource

Statistical Significance vs Clinical Importance: Why p < 0.05 Is Not the Final Research Question

Statistical significance and clinical importance answer different questions. This Resource explains how to interpret effect magnitude, confidence intervals, clinically meaningful thresholds, sample size, and p values together before drawing a practical conclusion.

A result crosses the conventional threshold:

p < 0.05.

What has the study established?

Core principle: It may have provided statistical evidence against a specified null hypothesis. It has not, from that fact alone, established that the effect is large, clinically important, useful, or worth acting on.

This distinction is central to understanding statistical significance vs clinical significance. Altman explicitly warns that statistical and clinical significance are not equivalent. A statistically significant difference can be too small to matter clinically, while a nonsignificant result can remain compatible with an important effect when the study is imprecise (Altman, 1991).

Borenstein et al. make the corresponding point from an effect-size perspective: the question that often matters to clinicians, patients, and researchers is the magnitude of the effect. A p value can address evidence relative to a null value, but it does not by itself tell us how large the effect is or whether that magnitude matters (Borenstein et al., 2021).

A better interpretation therefore follows this sequence:

1. Magnitude

What effect was actually estimated?

2. Uncertainty

How precisely was the effect estimated?

3. Clinical threshold and context

How large would an effect need to be to matter?

4. Statistical evidence

What does the p value add?

5. Practical conclusion

What does the total evidence justify saying?

Magnitude → uncertainty → clinical threshold/context → statistical evidence → practical conclusion

That sequence keeps the scientific question in front of the significance threshold.

Statistical significance and clinical importance answer different questions

Statistical significance concerns an inferential comparison with a null hypothesis. In conventional hypothesis testing, a p value is evaluated against a prespecified significance level such as 0.05.

Clinical importance asks something different:

Is the magnitude of the effect large enough to matter in the clinical or scientific context?

Those questions can produce different answers.

Altman notes that the word “significant” itself encourages confusion between statistical and clinical significance. He gives an example in which a very small difference was highly statistically significant but clinically unimportant. He also cautions against reversing the mistake: a nonsignificant result should not automatically be interpreted as evidence of no effect (Altman, 1991).

The practical implication is that four combinations are possible:

Possible combinations of statistical evidence and clinical importance
Statistical evidence Clinical importance Interpretation
Statistically significant Clinically important Evidence against the null accompanies an effect that may matter.
Statistically significant Clinically unimportant The effect is detectable but too small to matter for the relevant decision.
Not statistically significant Potentially clinically important Important effects may remain compatible with the data because uncertainty is substantial.
Not statistically significant Clinically unimportant effects reasonably excluded The estimate may be sufficiently precise to make important effects difficult to reconcile with the data.

A binary significant/nonsignificant label cannot distinguish these possibilities.

The StatsAlly interpretation framework

1. Magnitude: What effect was actually estimated?

Start with the effect estimate, not the significance label.

Depending on the research question and outcome, the effect might be expressed as a mean difference, risk difference, risk ratio, odds ratio, correlation, standardized mean difference, or another appropriate measure.

The important question is:

How large is the estimated effect, in what direction, and on what scale?

Borenstein et al. emphasize that effect size carries information about magnitude. In their discussion of clinical importance, they contrast this directly with the p value: for a decision about an intervention, knowing the size of the effect is often the information of substantive interest (Borenstein et al., 2021).

Effect-size interpretation must therefore remain tied to the scale being reported. A difference in the original clinical units may have a different practical interpretation from a standardized mean difference, and a relative effect may communicate something different from an absolute effect.

Do not translate p < 0.001 into “The treatment had a large effect.”

The p value does not supply that information.

2. Uncertainty: How precisely was the effect estimated?

A point estimate is incomplete without information about its uncertainty.

Confidence intervals are especially useful because they display the precision of an estimate. Altman argues that confidence intervals convey substantially more useful information than reducing results to whether they are statistically significant. They show the uncertainty around the estimated effect and can also indicate how the estimate relates to the null value (Altman, 1991).

This changes the interpretation from:

“Was the result significant?”

to:

“What range of effects is compatible with the estimate and its sampling uncertainty?”

A narrow confidence interval indicates greater precision than a wide interval. That distinction can radically alter the meaning of otherwise similar results.

Consider two nonsignificant results.

One confidence interval might be tightly concentrated around effects too small to matter clinically. Another might span substantial benefit, little effect, and clinically important harm.

Calling both results simply “not significant” discards the difference that matters most.

Altman describes this problem directly in his discussion of trials with nonsignificant findings: some studies were too imprecise to rule out clinically valuable treatment effects. Their lack of statistical significance was therefore not evidence that an important effect had been excluded (Altman, 1991).

3. Clinical threshold and context: How large would an effect need to be to matter?

After estimating magnitude and uncertainty, compare them with the substantive question.

The relevant benchmark should come from the clinical or scientific context rather than being created by the p value.

For study planning, Altman frames the problem in terms of specifying the smallest treatment difference that would be clinically worthwhile. Piantadosi similarly identifies the smallest treatment effect of interest based on clinical considerations as an important quantitative design parameter. Chow et al. explicitly incorporate clinically or scientifically meaningful differences into power-based sample-size planning (Altman, 1991; Chow et al., 2018; Piantadosi, 2005).

This principle matters both before and after data collection.

Before the study

What difference would be important enough for this study to detect or estimate adequately?

After the study

Where do the estimated effect and its confidence interval lie relative to values that would matter?

The clinically meaningful threshold is not necessarily the null value.

Zero difference may be statistically important because it defines the conventional null for a difference measure. But the clinical question may concern whether the benefit is large enough to influence treatment, policy, resource allocation, or another substantive decision.

Clinical importance cannot be inferred from statistical significance alone.

4. Statistical evidence: What does the p value add?

Only after considering magnitude, uncertainty, and context should the p value be used to complete the interpretation.

The p value contributes information about the statistical evidence relative to the tested null hypothesis. It should not be treated as a measurement of effect magnitude.

Altman specifically rejects sharp interpretive discontinuities around 0.05. Results such as p = 0.045 and p = 0.055 should not lead to radically different scientific interpretations simply because they fall on opposite sides of an arbitrary threshold. He argues that forcing results into “significant” and “nonsignificant” categories obscures the uncertainty inherent in statistical inference (Altman, 1991).

Borenstein et al. reinforce the point in meta-analysis. Individual studies with larger observed effects may fail to achieve statistical significance when they contain relatively little information, while studies with smaller effects may achieve significance when they are more precise. They emphasize that the p value reflects both the effect magnitude and the amount of information available (Borenstein et al., 2021).

Statistical evidence is therefore one part of the conclusion—not the conclusion itself.

5. Practical conclusion: What does the total evidence justify saying?

The final conclusion should integrate all four preceding components.

Instead of:

“The treatment was statistically significant (p = 0.03).”

a research interpretation should answer:

What was estimated? How large was it? How uncertain was it? How does that range compare with effects that matter? What statistical evidence was observed?

The resulting conclusion may be that the data support a clinically important effect.

But it could also be that:

  • the effect is statistically detectable but clinically small;
  • the point estimate appears clinically promising but remains too imprecise for a firm conclusion;
  • the result is nonsignificant but important effects remain plausible;
  • or the estimate is precise enough to weigh against effects of a clinically important magnitude.

That is why binary reporting is insufficient.

Why large samples can make small effects statistically significant

Statistical significance depends partly on sample size because sample size affects precision.

Altman notes that uncertainty decreases as sample size increases. He also gives examples showing why large datasets require particular care: with a very large sample, statistical significance can be achieved for effects that are extremely small in magnitude (Altman, 1991).

Borenstein et al. similarly explain that a p value reflects both effect magnitude and the volume of information available. A small p value can therefore arise from a large effect, but it can also arise from a modest or small effect estimated with substantial information (Borenstein et al., 2021).

This creates the familiar situation of a result that is:

statistically significant but not clinically significant.

The mistake is to read stronger statistical evidence as evidence of a larger effect.

It is not.

With increasing information, researchers become better able to distinguish smaller departures from the null. Whether those departures are important remains a separate substantive question.

Why small studies can leave important effects unresolved

The reverse problem is equally important.

Small studies usually estimate effects less precisely than otherwise comparable larger studies. A clinically important true effect can therefore fail to achieve conventional statistical significance because the available data do not estimate it precisely enough.

Altman describes published trials with nonsignificant results whose confidence intervals nevertheless remained compatible with substantial therapeutic improvement. He notes that such findings can also be understood as consequences of low power and insufficient sample size (Altman, 1991).

Borenstein et al. provide the same lesson from collections of studies. In their discussion of meta-analysis, many individual studies failed to reach statistical significance not because their observed treatment effects were smaller, but because their sample sizes and statistical power were limited (Borenstein et al., 2021).

p > 0.05 does not establish that an important effect is absent.

The confidence interval is essential for distinguishing a reasonably precise near-null result from an inconclusive result that leaves important effects unresolved.

Effect size interpretation is not optional

A useful research report should identify what the effect measure actually means.

This matters because different effect measures answer different quantitative questions. Borenstein et al. distinguish multiple effect-size indices and emphasize that effect size is central to understanding magnitude (Borenstein et al., 2021).

For clinical interpretation, researchers should therefore ask:

What does this effect measure represent on the scale relevant to the research question?

A standardized effect may be useful for comparison or synthesis, but clinical decisions may also require interpretation on an outcome scale that has direct substantive meaning.

Likewise, relative and absolute effects should not be treated as interchangeable descriptions merely because both can achieve the same statistical significance decision.

The aim is not simply to attach an “effect size” somewhere in the results section. It is to report a magnitude whose meaning is clear for the scientific question.

Statistical significance vs clinical significance begins at study planning

The distinction should not first appear after the analysis has been run.

It belongs in the design.

Altman describes sample-size planning as choosing a study size with a high probability of detecting a worthwhile effect if such an effect exists. Piantadosi identifies clinically based treatment effects among the quantitative parameters required for trial planning. Chow et al. describe prestudy power analysis as choosing the sample size required to detect a clinically or scientifically meaningful difference at a specified Type I error rate and desired power (Altman, 1991; Chow et al., 2018; Piantadosi, 2005).

Chow et al. also distinguish power-based planning from precision-based planning. A study can be designed around detecting a meaningful difference, or sample size can be selected to achieve a desired precision at a specified confidence level (Chow et al., 2018).

Detection question

How much information is needed to have adequate power to detect an effect that would matter?

Estimation question

How much information is needed to estimate the effect with sufficiently useful precision?

Neither question is answered by deciding in advance that p must be below 0.05.

The meaningful effect or precision target must also be defined.

A five-step reporting checklist

When interpreting a treatment effect, association, group difference, or other inferential result, use this sequence:

  1. Magnitude: Report the effect estimate and identify its scale and direction.
  2. Uncertainty: Report the confidence interval and evaluate the precision of the estimate.
  3. Clinical threshold/context: Compare the estimate and interval with differences that would matter clinically or scientifically.
  4. Statistical evidence: Report and correctly interpret the p value or other planned inferential evidence.
  5. Practical conclusion: State what the combination of magnitude, uncertainty, context, and statistical evidence actually supports.

The key is the order.

Starting with the effect and its uncertainty prevents the significance threshold from becoming the entire interpretation.

Common mistakes to avoid

“The result was significant, so the effect is clinically important.”

Incorrect. Statistical significance does not establish that the estimated magnitude matters clinically (Altman, 1991; Borenstein et al., 2021).

“The p value was extremely small, so the effect must be large.”

Incorrect. Statistical evidence depends on both effect magnitude and information or precision. Large samples can make small effects statistically detectable (Altman, 1991; Borenstein et al., 2021).

“The result was nonsignificant, so there is no effect.”

Incorrect. An imprecise study may remain compatible with clinically important effects (Altman, 1991).

“The confidence interval crosses the null, so there is nothing useful to interpret.”

Incorrect. The width and location of the interval show which effect magnitudes remain compatible with the estimate and its uncertainty. A wide interval and a narrow interval around the null do not carry the same information (Altman, 1991).

“Clinical importance can be decided after seeing whether the result is significant.”

Poor practice. The effect considered scientifically or clinically meaningful is part of study planning and can directly affect sample-size requirements (Altman, 1991; Chow et al., 2018; Piantadosi, 2005).

Bottom line

The debate over statistical significance vs clinical significance is not an argument that p values contain no information.

It is an argument against asking them to answer a question they were not designed to answer.

A significance test addresses statistical evidence relative to a specified null hypothesis. An effect estimate describes magnitude. A confidence interval communicates uncertainty and precision. Clinical or practical importance requires comparing those magnitudes with the consequences and thresholds relevant to the research problem (Altman, 1991; Borenstein et al., 2021).

Sample size connects these pieces. Very large studies can make small effects statistically detectable, while small studies may leave important effects unresolved. Study planning should therefore begin with the effect or precision that matters, not merely with the desire to achieve p < 0.05 (Altman, 1991; Chow et al., 2018; Piantadosi, 2005).

The final research question is not:

“Was p < 0.05?”

It is:

“What is the estimated effect, how uncertain is it, how does it compare with effects that matter, what statistical evidence supports it, and what practical conclusion does the total evidence justify?”

Frequently Asked Questions

What is the difference between statistical significance and clinical significance?

Statistical significance concerns evidence relative to a statistical null hypothesis. Clinical significance or clinical importance concerns whether the magnitude of an effect is meaningful in the clinical context. A result can be statistically significant without being clinically important, and a nonsignificant study can remain compatible with clinically important effects when uncertainty is substantial (Altman, 1991).

Does p < 0.05 mean an effect is clinically important?

No. A p value below 0.05 does not establish that the effect is large enough to matter. Effect magnitude, uncertainty, and clinical context must be examined separately (Altman, 1991; Borenstein et al., 2021).

Can a result be statistically significant but not clinically significant?

Yes. Large amounts of information can make very small effects statistically detectable. Altman specifically describes situations in which tiny effects achieve statistical significance in large samples (Altman, 1991).

Can a clinically important effect be statistically nonsignificant?

Yes. When a study is small or otherwise imprecise, its confidence interval may include both the null value and clinically important effects. Statistical nonsignificance alone does not distinguish that situation from a precise estimate showing little clinically meaningful effect (Altman, 1991; Borenstein et al., 2021).

Why should I report a confidence interval with a p value?

The confidence interval communicates the estimated effect's precision and helps show the range of effect values compatible with the estimate and sampling uncertainty. Altman argues that this provides information that a significance label alone cannot convey (Altman, 1991).

How does sample size affect statistical significance?

Increasing sample size generally increases information and improves precision, making smaller departures from the null easier to detect statistically. Conversely, limited sample size can leave even potentially important effects too imprecisely estimated to reach conventional significance (Altman, 1991; Borenstein et al., 2021).

When should a clinically meaningful difference be defined?

Ideally, the difference that matters should be considered during study planning. It can serve as the effect targeted in power calculations and helps connect the statistical design to the clinical or scientific objective (Altman, 1991; Chow et al., 2018; Piantadosi, 2005).

What should researchers report instead of only “significant” or “not significant”?

Report the effect magnitude, confidence interval, clinically or scientifically relevant context, and statistical evidence, then state a practical conclusion that reflects all of them. Binary significance reporting alone suppresses important information about both magnitude and uncertainty (Altman, 1991; Borenstein et al., 2021).

References

Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.

Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2021). Introduction to meta-analysis (2nd ed.). Wiley.

Chow, S.-C., Shao, J., Wang, H., & Lokhnygina, Y. (2018). Sample size calculations in clinical research (3rd ed.). Chapman & Hall/CRC.

Piantadosi, S. (2005). Clinical trials: A methodologic perspective (2nd ed.). John Wiley & Sons.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry