Resource

Interpreting Nonsignificant Results: Why P > .05 Does Not Mean No Effect

Learn how to interpret a nonsignificant result using effect estimates, confidence intervals, precision, and sample size—and why a p value greater than .05 does not by itself demonstrate that there is no effect.

A nonsignificant result interpretation should not begin and end with the statement that p > .05.

Yet researchers frequently translate a nonsignificant comparison into language such as:

“There was no difference.”

That conclusion is often stronger than the analysis supports.

A conventional hypothesis test evaluates evidence against a specified null hypothesis. If the result does not cross the chosen significance threshold, the test has not supplied sufficient evidence for rejection at that threshold. The result does not, merely by being nonsignificant, demonstrate that the true effect is zero or that scientifically important differences are absent. The confidence interval is critical because it shows the estimated effect together with the uncertainty surrounding it (Altman, 1991; Rosner, 2016).

The practical question is not simply: Is the p value greater than 0.05?

It is: What effect was estimated, how precisely was it estimated, and which scientifically or clinically important effects remain compatible with the confidence interval?

That distinction separates an inconclusive study from a study whose data are sufficiently precise to exclude effects large enough to matter.

Hypothesis Testing and Estimation Answer Different Questions

Hypothesis testing starts with a null hypothesis, commonly representing no difference or no association. The analysis then evaluates the observed data relative to what would be expected under that hypothesis. The resulting p value is used to determine whether the evidence is sufficiently inconsistent with the null hypothesis according to the prespecified testing rule (Altman, 1991; Rosner, 2016).

Estimation asks a different question:

What is the magnitude of the effect, and how uncertain is that estimate?

Altman emphasizes that an isolated p value does not communicate the magnitude of the effect. A confidence interval does. It places an interval around the estimate and thereby conveys uncertainty or lack of precision in the estimated quantity (Altman, 1991).

This distinction matters because the scientific question is rarely just whether an effect is exactly zero. Researchers usually also need to know whether an effect could be large enough to matter.

What Does P > .05 Actually Tell You?

When a conventional test produces p > .05, the data have not met the conventional 5% criterion for rejecting the specified null hypothesis.

That is a statement about the hypothesis test. It is not automatically a statement that the true effect equals zero.

The possibility of failing to detect a real effect is built into hypothesis-testing theory. Power concerns the probability of rejecting the null hypothesis for a specified alternative, and the probability of failing to reject when that alternative is true is a Type II error probability. Power depends on the alternative effect under consideration and is affected by the amount of information available in the study (Rosner, 2016).

Consequently, a nonsignificant result may occur because the underlying effect is small, but it may also occur because the estimate is too imprecise to distinguish an important effect from the null value.

“No significant difference” and “no difference” are not interchangeable conclusions.

Why the Confidence Interval Changes the Interpretation

A confidence interval adds information that a binary significance classification cannot provide.

Altman describes the confidence interval as showing the uncertainty, or lack of precision, in the effect estimate. He also emphasizes that a nonsignificant finding is not necessarily ignorable: interpretation requires estimates and confidence intervals rather than a result being reduced to “significant” or “not significant” (Altman, 1991).

For a conventional 95% confidence interval paired with the corresponding two-sided 5% significance test, there is also a direct relationship between the two procedures. If the 95% confidence interval excludes the null value, the corresponding result is statistically significant at the 5% level; if the interval includes the null value, it is not statistically significant at that level (Altman, 1991).

For a difference measure, the conventional no-effect value is usually zero. For an appropriate ratio measure, the no-effect value is typically one.

But whether the interval crosses the null value is only the beginning of confidence interval interpretation.

The next question is:

How far does the interval extend from the null, and does it include effects that would be important?

Wide and Narrow Confidence Intervals Tell Different Stories

Consider two studies that both report p > .05.

Their statistical significance classification is the same. Their scientific interpretation may be very different.

Wide confidence interval

A wide interval may contain the null value while also extending into effects that would represent important benefit, important harm, or both.

Interpretation priority: unresolved uncertainty and limited precision.

Narrow confidence interval

A narrow interval around the null can be substantially more informative if it excludes effects large enough to matter scientifically or clinically.

Interpretation priority: determine which substantively important effects have been excluded.

A wide confidence interval signals unresolved uncertainty

A wide interval may contain the null value while also extending into effects that would represent important benefit, important harm, or both.

Such a result is not strong evidence of “no effect.” It is evidence that the study has not estimated the effect precisely enough to distinguish among materially different possibilities.

Altman gives examples in which comparisons are nonsignificant despite sizeable observed differences. A wide confidence interval in that setting indicates that a larger study may be needed to draw more precise conclusions. He explicitly notes that a wide interval from a small study is not meaningless: it indicates that the study was too small for precise conclusions (Altman, 1991).

This is the setting in which absence of statistically significant evidence should not be converted into evidence that an important effect is absent.

A narrow confidence interval can be much more informative

Now consider a nonsignificant estimate close to the null whose confidence interval is narrow.

If that interval also excludes effects that would be large enough to matter scientifically or clinically, the study has provided considerably more information than a wide-interval nonsignificant result.

The conclusion should still be framed around the interval rather than claiming that an exact zero effect has been proved. But the researcher may be able to state that effects beyond the limits represented by the confidence interval are not supported by the interval.

Wide CI crossing the null → important effects may remain unresolved.

Narrow CI around the null → effects of important magnitude may be excluded, depending on the substantive threshold.

A p value alone does not reveal this distinction.

Precision Is Not the Same as Statistical Significance

Precision concerns how tightly an effect has been estimated.

Statistical significance concerns whether the data meet the rejection criterion for a specified hypothesis test.

These are related but different ideas.

Confidence-interval width provides a direct indication of uncertainty around an estimate. Altman illustrates how small samples can produce wide confidence intervals and how larger studies can produce narrower intervals for comparable kinds of estimates (Altman, 1991).

Sample-size planning can therefore be approached from two different directions.

Hypothesis-testing design

A design can be planned to achieve adequate power for a specified alternative effect.

Estimation-oriented design

A design can instead be planned so that the expected confidence interval has acceptable precision.

Piantadosi discusses confidence-interval-based planning as choosing sufficient sample size to obtain the required precision, while Rosner develops power and sample-size calculations in relation to specified alternatives (Piantadosi, 2005; Rosner, 2016).

A study can be nonsignificant because it has little information, not because it has demonstrated that the effect is negligible.

Sample Size Affects Both Precision and the Chance of Statistical Significance

Sample size affects the information available about an effect.

With other relevant features held comparable, more information generally permits greater precision. Small studies can therefore produce wide confidence intervals that encompass many substantially different effects. Altman's examples explicitly connect small sample size with wide confidence intervals and limited ability to draw precise conclusions (Altman, 1991).

Power is also linked to sample size and the alternative effect being considered. Rosner treats power as the probability of rejecting the null under a specified alternative and uses power calculations as a study-planning tool (Rosner, 2016).

This explains why two common interpretations are both wrong:

“It was nonsignificant, therefore there is no effect.”

“It was significant, therefore the effect is important.”

Neither follows from statistical significance alone.

Effect magnitude, uncertainty, and substantive importance must also be examined.

The Relationship Between Confidence Intervals and Significance Tests

For corresponding conventional two-sided procedures, confidence intervals and hypothesis tests are closely related.

Altman states that a p value is below .05 precisely when the corresponding 95% confidence interval excludes the value specified by the null hypothesis. The same relationship extends to matching confidence levels and significance levels more generally (Altman, 1991).

This means that a confidence interval already communicates whether the corresponding test crosses the conventional significance threshold while providing additional information about magnitude and precision.

Significance-only reporting

p = .08; not significant.”

This gives the significance classification but does not show the magnitude or precision of the estimate.

Estimate-and-interval reporting

“The estimated effect was X, with a 95% confidence interval from L to U.”

This also reveals whether the data remain compatible with effects that would matter.

That is why confidence intervals are central to responsible nonsignificant result interpretation.

A Decision Guide for Interpreting P Values and Confidence Intervals

Interpretation priorities for combinations of statistical significance and confidence-interval precision
Statistical pattern What it tells you Interpretation priority What not to conclude
p > .05 + wide CI Null not rejected; estimate is imprecise Identify the important benefit and harm still included in the interval “There is no effect”
p > .05 + narrow CI Null not rejected; estimate is comparatively precise Ask whether effects large enough to matter have been excluded “Zero effect has been proved”
p < .05 + small effect Null rejected; estimated effect may nevertheless be small Interpret magnitude and practical importance “The effect is important because it is significant”
p < .05 + wide uncertainty Null rejected, but magnitude remains uncertain Report the estimate and full interval; determine whether materially different effect sizes remain compatible “The effect size is known precisely”

1. P > .05 + wide confidence interval

Primary interpretation: insufficient precision.

This pattern should trigger caution about any statement of “no difference.”

Ask:

Does the interval still contain effects that would matter?

If it contains meaningful benefit, meaningful harm, or both, the study has not resolved those possibilities. The correct emphasis is uncertainty, not absence of an effect.

Altman's worked examples illustrate exactly this problem: nonsignificant comparisons can coexist with large observed differences and wide intervals, with the width indicating that a larger study would be required for greater precision (Altman, 1991).

Prefer

“The analysis did not demonstrate a statistically significant difference, but the confidence interval is wide and remains compatible with effects of potentially important magnitude.”

Avoid

“There was no difference.”

2. P > .05 + narrow confidence interval

Primary interpretation: determine what important effects have been excluded.

A narrow interval concentrated around the null is fundamentally different from a wide interval spanning large effects.

If the interval excludes effects exceeding a scientifically or clinically meaningful threshold, the study provides evidence against effects of that magnitude—even though the conventional test did not reject an exact no-effect null.

This is the situation closest to evidence of absence of an important effect, but the claim must be defined relative to an effect magnitude that matters. A nonsignificant test by itself does not establish that conclusion.

Prefer

“The estimate was close to the null and comparatively precise; the confidence interval excludes effects larger than the prespecified clinically important range.”

Avoid

“The true effect is zero.”

3. P < .05 + small effect

Primary interpretation: separate statistical significance from practical importance.

Rejecting a null hypothesis does not establish that the effect is large enough to matter.

Altman explicitly distinguishes statistical significance from clinical significance and argues that effect estimates and confidence intervals are required to determine what was actually observed (Altman, 1991).

A small effect may be estimated precisely enough to exclude zero and therefore produce statistical significance.

The appropriate question is:

Is the magnitude important for the scientific or clinical decision?

Prefer

“The estimated effect differed statistically from the null, but its magnitude was small; practical importance should be judged from the effect estimate and confidence interval.”

Avoid

“The result is important because p < .05.”

4. P < .05 + wide uncertainty

Primary interpretation: evidence against the null does not imply precise knowledge of magnitude.

An interval can exclude the null value yet still span substantially different effect sizes.

In that situation, the data support a nonzero effect under the specified testing procedure, but important uncertainty about its magnitude remains.

Prefer

“The result was statistically significant, but the confidence interval indicates substantial uncertainty about the magnitude of the effect.”

Avoid

Treating the point estimate as though it were known without meaningful uncertainty.

Absence of Evidence Versus Evidence of Absence

The phrase absence of evidence is useful only when its meaning is made precise.

A nonsignificant result with a wide confidence interval is a clear example of why failing to demonstrate an effect cannot automatically establish its absence. Piantadosi warns that low power or low precision can make treatments appear “not significantly different,” which is precisely why ordinary failed superiority testing cannot establish equivalence (Piantadosi, 2005).

Failure to detect a difference can result from inadequate ability to distinguish the relevant alternatives.

Evidence that an important effect is absent requires a stronger argument: the study must have enough precision to exclude effects of the magnitude regarded as important.

This principle is particularly visible in equivalence and noninferiority research. Piantadosi's treatment makes clear that conventional “not significantly different” reasoning is inadequate because low precision can favor an apparent equivalence conclusion. Meaningful similarity requires an analysis designed around appropriate boundaries and sufficient precision (Piantadosi, 2005).

Interpretive rule: Do not ask only whether the confidence interval contains the null. Ask which important effects it excludes.

Why “No Effect” Is Especially Risky as a Post Hoc Conclusion

A study designed to test for a difference cannot automatically be reinterpreted after a nonsignificant result as having demonstrated equivalence or absence of a meaningful effect.

That reversal is particularly problematic when the study is imprecise. Piantadosi notes that conventional comparative testing is poorly suited to demonstrating equivalence because low power can actually favor a conclusion based on “no significant difference” (Piantadosi, 2005).

The scientific criterion therefore needs to precede the result.

If the research question is whether an effect is negligibly small, the investigator must define what “negligibly small” means in the substantive context and use a design and inferential approach capable of addressing that question.

A post hoc statement such as “Because p > .05, we conclude there is no clinically meaningful effect” is not justified merely by the nonsignificant p value.

The confidence interval must first show whether clinically meaningful effects have actually been excluded.

A Practical Confidence Interval Interpretation Workflow

When interpreting any estimated treatment effect, association, difference, or regression parameter, use this sequence:

  1. Define the target effect. What quantity answers the research question?
  2. Identify the no-effect value. For the chosen effect measure, what value corresponds to the null hypothesis?
  3. Read the point estimate. What magnitude and direction were actually observed?
  4. Read the full confidence interval. Do not reduce it to whether it crosses the null.
  5. Assess precision. Is the interval narrow enough to distinguish scientifically different possibilities?
  6. Compare the interval with effects that matter. Does it include important benefit, important harm, or only effects small enough to be considered unimportant?
  7. Interpret the p value in its proper role. Does the corresponding hypothesis test reject the specified null at the prespecified level?
  8. Match the conclusion to the design. Do not turn a failed superiority test into a post hoc equivalence claim.
  9. Report magnitude and uncertainty together. Statistical significance should supplement, not replace, interpretation of the estimated effect.

Better Language for Reporting Nonsignificant Results

Instead of concluding that nonsignificance means no effect

Instead of

“There was no significant difference between groups, indicating that the intervention had no effect.”

Write

“The analysis did not provide sufficient evidence to reject the no-effect null hypothesis. Interpretation should therefore focus on the estimated effect and its confidence interval.”

If the interval is wide

“The confidence interval includes the null value but also includes effects that could be substantively important; the estimate is therefore too imprecise to support a conclusion of no meaningful effect.”

If the interval is narrow

“The estimate is close to the null and comparatively precise. The confidence interval excludes effects larger than the prespecified magnitude considered important, although it should not be interpreted as proving that the effect is exactly zero.”

The distinction is not stylistic. It changes the scientific claim.

Common Interpretation Mistakes

Mistake 1: “P > .05 means no effect”

No. It means that the specified test did not reject the null hypothesis at the chosen threshold. The confidence interval is needed to determine what effect magnitudes remain compatible with the estimate and its uncertainty (Altman, 1991).

Mistake 2: “The confidence interval includes zero, so the result tells us nothing”

Also no. A wide interval may reveal substantial unresolved uncertainty. A narrow interval may exclude effects large enough to matter. Both provide information about precision (Altman, 1991).

Mistake 3: “P < .05 means the effect is important”

Statistical significance and clinical or practical importance are different questions. Effect magnitude and uncertainty must be considered separately (Altman, 1991).

Mistake 4: “A small study with no significant difference shows the treatments are similar”

An imprecise study can produce nonsignificance precisely because it cannot distinguish important differences from the null. Piantadosi warns against using this logic to infer equivalence (Piantadosi, 2005).

Mistake 5: “The point estimate is the effect”

The point estimate is the observed estimate of the unknown parameter. The confidence interval communicates the uncertainty surrounding that estimate.

Reporting Checklist

Before writing “no difference” or “no effect,” check the following:

  • What effect was actually estimated?
  • What is the null value for that effect measure?
  • What is the confidence interval?
  • Is the interval wide or narrow relative to effects that matter?
  • Does the interval still include clinically or scientifically important benefit?
  • Does it include important harm?
  • Does it exclude effects large enough to matter?
  • Was the study designed to demonstrate superiority, equivalence, noninferiority, or another claim?
  • Is the conclusion about effect magnitude, or only about statistical significance?
  • Would the interpretation change if the p value were hidden and only the estimate and confidence interval were shown?

The final question is particularly useful. It forces interpretation back onto the estimated effect and its uncertainty.

Bottom Line

A p value greater than 0.05 does not mean that there is no effect.

It means that the specified hypothesis test did not reject its null hypothesis at the chosen significance level. What the study says scientifically depends on the estimated effect and the precision with which that effect was estimated.

A wide confidence interval crossing the null may indicate that important benefit or harm remains unresolved.

A narrow confidence interval near the null may provide much stronger evidence against effects large enough to matter—provided “large enough to matter” has been defined substantively rather than invented after seeing the data.

Likewise, statistical significance does not establish practical importance. A statistically significant effect can be small, and a statistically significant result can still leave substantial uncertainty about magnitude.

The most useful question after any hypothesis test is not: “Was p below .05?”

It is: “What effect was estimated, how uncertain is that estimate, and which scientifically important effects does the confidence interval actually exclude?”

Frequently Asked Questions

Does p > .05 mean there is no effect?

No. A p value above .05 indicates that the specified hypothesis test did not reject the null hypothesis at the 5% level. It does not by itself demonstrate that the true effect equals zero. The effect estimate and confidence interval are needed to assess magnitude and uncertainty (Altman, 1991; Rosner, 2016).

What does a wide confidence interval mean after a nonsignificant result?

A wide confidence interval indicates limited precision. If it includes the null value together with effects large enough to matter, the study has not distinguished adequately among those possibilities. Altman illustrates that small studies can produce wide intervals that prevent precise conclusions (Altman, 1991).

Can a nonsignificant result provide evidence that an important effect is absent?

Potentially, but not from the p value alone. A comparatively narrow confidence interval may exclude effects exceeding a prespecified scientifically or clinically important magnitude. The conclusion should be framed around what the interval excludes rather than claiming that an exact zero effect has been proved.

Why can two studies with p > .05 have different interpretations?

Because their effect estimates and confidence intervals may differ substantially. One study may be highly imprecise and remain compatible with large benefit or harm. Another may estimate an effect close to the null with enough precision to exclude effects large enough to matter. The significance label alone cannot distinguish those situations (Altman, 1991).

Does a 95% confidence interval that includes zero always correspond to p > .05?

For corresponding conventional two-sided procedures, yes: Altman describes the direct relationship whereby a 95% confidence interval includes the null value when the corresponding test is not significant at the 5% level and excludes it when the test is significant (Altman, 1991).

Does statistical significance mean an effect is clinically important?

No. Statistical significance concerns evidence relative to a null hypothesis. Clinical or practical importance concerns the magnitude and consequences of the effect. Altman explicitly emphasizes that statistical and clinical significance should not be treated as equivalent (Altman, 1991).

Can “no significant difference” establish equivalence?

No. Piantadosi warns that low power or low precision can make treatments appear not significantly different. Equivalence requires an inferential framework designed to exclude differences beyond clinically acceptable limits; it is not the fallback conclusion from a failed superiority test (Piantadosi, 2005).

Should I report the confidence interval if the result is nonsignificant?

Yes. Altman emphasizes estimates and confidence intervals because they communicate magnitude and uncertainty that a significance classification alone cannot provide (Altman, 1991).

References

Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.

Piantadosi, S. (2005). Clinical trials: A methodologic perspective (2nd ed.). John Wiley & Sons.

Rosner, B. (2016). Fundamentals of biostatistics (8th ed.). Cengage Learning.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry