P-Values, Effect Sizes, and Confidence Intervals: How to Read Them Together
Learn how to interpret p-values, effect estimates, and confidence intervals together rather than relying on statistical significance alone. This Resource shows how evidence, magnitude, precision, and practical importance contribute to a defensible statistical conclusion.
A statistical result is rarely well interpreted by asking only whether p < .05. A p-value addresses evidence against a specified null hypothesis; an effect estimate describes the observed direction and magnitude; and a confidence interval shows how precisely that effect has been estimated and what range of parameter values is compatible with the interval procedure. These quantities answer related but different questions, so interpretation is strongest when they are read together rather than treated as competing summaries (Moore et al., 2014, 2021; Devore & Berk, 2012; Newton & Rudestam, 1999).
The practical question is not simply “Is the result statistically significant?”
What does the result tell us about the evidence, direction, magnitude, precision, and practical importance of the effect?
The integrated interpretation framework
Read a result in the following sequence:
| Quantity or decision | Question it helps answer |
|---|---|
| p-value | How incompatible is the observed test statistic with the null hypothesis, under the assumptions of the test? |
| Effect estimate | What direction and magnitude of effect was observed? |
| Confidence interval | How precisely is the effect estimated, and what range of parameter values is represented by the interval? |
| Practical interpretation | Are effects of the estimated magnitude—or values still compatible with the interval—important in the research context? |
A p-value is calculated under the assumption that the null hypothesis is true and reflects how extreme the observed test statistic is relative to that null distribution. Smaller p-values provide stronger evidence against the null hypothesis, but they do not by themselves describe how large an effect is (Moore et al., 2021; Devore & Berk, 2012).
An effect estimate addresses that missing question of magnitude. An estimated mean difference, correlation, regression coefficient, risk difference, odds ratio, or other effect measure tells the researcher what was observed and in which direction. Statistical significance and effect magnitude are not interchangeable: a test can establish strong evidence that an effect differs from its null value while the effect itself remains small (Newton & Rudestam, 1999; Lovric, 2011).
A confidence interval adds information about uncertainty. Narrow intervals indicate greater precision, whereas wide intervals communicate greater uncertainty about the parameter being estimated. Devore and Berk (2012) explicitly connect confidence-interval width with precision: a narrow interval can indicate relatively precise knowledge, while a wide interval represents a broad range of plausible parameter values. Confidence intervals therefore help researchers see information that an isolated significance decision hides (Devore & Berk, 2012; Newton & Rudestam, 1999).
1. Start with the p-value: what evidence does the test provide?
A significance test begins with a null hypothesis and asks whether the data provide evidence against it. The p-value is the probability, calculated assuming the null hypothesis is true, of obtaining a test statistic at least as extreme as the one observed. Smaller p-values indicate stronger evidence against the null hypothesis (Moore et al., 2021).
Evidence question: How strongly do these data challenge the specified null hypothesis?
It does not answer several other questions researchers commonly want answered. In particular, it does not tell you the size of the effect or whether that size matters in practice. Lovric (2011) distinguishes these roles directly: a hypothesis test can inform statistical significance without conveying the magnitude of an effect size (Lovric, 2011).
It is also preferable to retain the actual p-value rather than reduce the result immediately to a binary label. Moore et al. (2021) note that the p-value provides more information about the strength of evidence than merely saying whether a result crosses a fixed significance level.
2. Then examine the effect estimate: what happened, and in which direction?
Once the evidence question has been considered, move to the estimated effect.
Suppose a treatment-control comparison produces an estimated mean difference of +3.2 points. That estimate provides information that “p = .02” cannot: the estimated treatment effect is positive, and its observed magnitude is 3.2 points.
Direction
What is the estimated direction of the effect?
Magnitude
How large is the estimated effect on the chosen scale?
The substantive meaning of that magnitude cannot be determined mechanically from statistical significance. Newton and Rudestam (1999) emphasize that statistical and substantive significance are different and that the meaning assigned to a difference depends on the research problem. Similarly, the International Encyclopedia of Statistical Science notes that statistically significant results can have effect sizes too small to achieve practical importance (Lovric, 2011).
An effect-size number should not automatically be classified as important or unimportant by applying an unsupported universal threshold. Context matters. The same numerical effect can have different implications depending on the outcome, decision, population, costs, risks, and research objective (Newton & Rudestam, 1999; Lovric, 2011).
3. Add the confidence interval: how uncertain is the magnitude?
A point estimate is only one estimate produced by one sample. The confidence interval makes sampling uncertainty visible.
Confidence intervals are constructed by procedures with a stated long-run coverage property. For example, a 95% confidence procedure is designed so that, over repeated sampling under its assumptions, 95% of the resulting intervals contain the true parameter. It should not be interpreted as saying that there is a 95% probability that the fixed parameter lies inside the particular interval already calculated (Moore et al., 2014, 2021).
Interpretation boundary: A 95% confidence interval should not be described as assigning a 95% probability to the fixed parameter being inside the particular interval already calculated.
For interpretation, interval width is especially useful. A narrow confidence interval indicates greater precision, whereas a wide interval indicates substantial uncertainty about the parameter value (Devore & Berk, 2012).
Point estimate
What effect did we estimate?
Confidence interval
How tightly have we estimated it?
That distinction matters because two studies can report the same point estimate while supporting very different conclusions if one estimate is precise and the other is highly uncertain.
4. Compare the interval with values that matter
A confidence interval becomes especially informative when it is interpreted against substantively meaningful values rather than only against the null value.
Consider a synthetic study in which researchers decide, based on the substantive context, that an improvement of 5 points would be large enough to influence a practical decision. The 5-point value is a synthetic study-specific decision threshold, not a universal statistical rule.
Now suppose the estimated improvement is 6 points.
| Result | 95% confidence interval | Interpretation |
|---|---|---|
| Same point estimate: +6 points | 5.2 to 6.8 | A comparatively precise estimate whose entire interval lies above the study-specific 5-point benchmark. |
| Same point estimate: +6 points | −1 to 13 | The interval includes deterioration, negligible effects, and effects substantially exceeding the benchmark. |
The second study therefore does not justify the simple conclusion that a practically important effect has been established. Its interval remains compatible with meaningfully different substantive conclusions. This follows from treating the interval as information about effect magnitude and precision rather than merely asking whether it contains the null value (Devore & Berk, 2012; Newton & Rudestam, 1999).
5. Sample size connects all three quantities
Sample size is one reason p-values, effect estimates, and confidence intervals must be interpreted together.
As sample size increases, sampling variability generally decreases, confidence intervals become narrower under otherwise comparable conditions, and statistical tests become better able to detect departures from the null hypothesis (Moore et al., 2021; Newton & Rudestam, 1999).
That increased sensitivity has an important consequence: with sufficiently large samples, very small effects can become statistically significant. Moore et al. (2021) explicitly warn that large samples can make tiny deviations from a null hypothesis statistically significant, while Newton and Rudestam (1999) similarly emphasize the need to examine practical significance and effect size as sample size grows. Devore and Berk (2012) likewise caution that a small p-value can result from a large sample combined with a departure from the null that has little practical importance.
A very small p-value does not necessarily imply a large effect.
Conversely, a potentially important effect can fail to achieve conventional statistical significance when it is estimated imprecisely. Failure to reject a null hypothesis does not demonstrate that the effect is absent; it indicates that the available evidence was insufficient for rejection under the specified test. Confidence intervals can reveal whether such a result is compatible only with small effects or with a much wider range that includes practically important effects (Moore et al., 2021; Newton & Rudestam, 1999).
Four synthetic result scenarios
The following scenarios use a fictional outcome for which positive values indicate improvement. For illustration, suppose the research team has independently determined that an improvement of 5 points would be practically important for its decision. That value is a synthetic contextual assumption, not a statistical convention.
Scenario 1: Statistically significant and potentially important, with good precision
Synthetic result: p = .002; estimated effect = +7.0 points; 95% CI [4.8, 9.2].
p-value: The result provides strong evidence against a zero-effect null hypothesis under the assumptions of the test.
Effect estimate: The estimated effect is positive and exceeds the fictional 5-point practical benchmark.
Confidence interval: The interval is reasonably concentrated around the positive estimate. It excludes zero, although its lower bound falls slightly below the fictional 5-point benchmark.
Defensible conclusion: The data provide evidence of a positive effect, estimated at 7 points. The interval suggests that the effect could plausibly range from somewhat below the study's practical benchmark to substantially above it. The evidence for a nonzero effect is therefore stronger than the evidence that the effect necessarily exceeds the practical benchmark throughout the interval (Moore et al., 2021; Devore & Berk, 2012).
Overclaim to avoid: “The intervention has definitively produced a practically important 7-point improvement.”
The estimate is 7 points; the parameter is not known to equal exactly 7, and part of the interval lies below the fictional practical benchmark.
Scenario 2: Statistically significant but practically trivial
Synthetic result: p < .001; estimated effect = +0.6 points; 95% CI [0.4, 0.8].
p-value: There is strong evidence against the zero-effect null hypothesis.
Effect estimate: The estimated effect is positive but very small relative to the fictional 5-point practical benchmark.
Confidence interval: The estimate is precise. The interval excludes zero but remains far below the study-specific benchmark throughout.
Defensible conclusion: The study provides precise evidence that the parameter differs positively from zero, but the estimated magnitude and its confidence interval indicate an effect that is trivial for the decision criterion specified in this synthetic study. This is precisely why statistical significance should not be equated with practical importance (Moore et al., 2021; Devore & Berk, 2012; Newton & Rudestam, 1999; Lovric, 2011).
Overclaim to avoid: “The very small p-value shows that the effect is large or important.”
It does not. Here, the p-value supplies strong hypothesis-test evidence while the effect estimate supplies evidence of small magnitude.
Scenario 3: Potentially important effect, but too imprecise for a firm conclusion
Synthetic result: p = .11; estimated effect = +6.0 points; 95% CI [−1.5, 13.5].
p-value: The test does not provide sufficient evidence to reject a zero-effect null hypothesis at a .05 significance level.
Effect estimate: The observed estimate is positive and exceeds the fictional practical benchmark.
Confidence interval: The interval is wide. It includes a small negative effect, zero, effects below the practical benchmark, and effects well above it.
Defensible conclusion: The study has produced a potentially important positive point estimate, but it is too imprecisely estimated to distinguish among substantively different possibilities. The data are compatible with little or no benefit as well as a large benefit. More information or greater precision would be needed for a firmer conclusion (Devore & Berk, 2012; Newton & Rudestam, 1999).
Overclaim to avoid: “There is no effect because p > .05.”
Failure to reject the null is not evidence that the true effect equals zero. The wide interval makes the unresolved uncertainty visible.
Scenario 4: Statistically significant, but the practical decision remains uncertain
Synthetic result: p = .03; estimated effect = +3.8 points; 95% CI [0.4, 7.2].
p-value: The data provide evidence against the zero-effect null hypothesis at the .05 level.
Effect estimate: The estimated effect is positive but below the fictional 5-point practical benchmark.
Confidence interval: The interval excludes zero, yet it spans effects ranging from very small to larger than the practical benchmark.
Defensible conclusion: There is evidence that the effect is positive, but the study does not precisely establish whether its magnitude is practically important. Statistical significance resolves the null-hypothesis question more clearly than it resolves the substantive decision (Newton & Rudestam, 1999; Devore & Berk, 2012).
Overclaim to avoid: “Because p < .05, the intervention has an important effect.”
The significance test and the practical decision are answering different questions.
A fifth scenario: precise evidence against a practically important effect
Consider one more synthetic pattern:
Synthetic result
p = .18; estimated effect = +0.8 points; 95% CI [−0.4, 2.0].
The result is not statistically significant, but that is not the most informative feature. The interval is relatively narrow and, under the fictional 5-point benchmark, does not include effects approaching the size required for practical importance.
The defensible conclusion is therefore stronger than simply saying “not significant.” The data do not provide evidence of a nonzero effect, and the interval also suggests that effects as large as the study-specific practical benchmark are not represented by this interval. That differs fundamentally from Scenario 3, where the nonsignificant result came with an interval extending far into practically important territory (Devore & Berk, 2012; Newton & Rudestam, 1999).
The overclaim to avoid is still “we proved there is no effect.” The more appropriate statement is that the estimate is small and comparatively precise, with the reported interval excluding effects as large as the fictional practical benchmark.
Why two identical p-values can tell very different stories
Suppose two studies both report p = .03.
| Study | Effect estimate | Precision |
|---|---|---|
| Study A | +0.4 | A narrow confidence interval close to zero. |
| Study B | +8 | A much wider interval. |
The identical p-values do not make these results substantively equivalent. Their estimated magnitudes and uncertainties differ. The p-value primarily addresses evidence relative to the null hypothesis; it is not a standardized measure of practical importance (Devore & Berk, 2012; Lovric, 2011).
The same principle works in reverse. Two studies can report the same effect estimate but have very different confidence intervals because their estimates differ in precision. An effect estimate without its uncertainty can therefore be misleading just as a p-value without an effect estimate can be incomplete (Devore & Berk, 2012; Newton & Rudestam, 1999).
A practical reading sequence for statistical results
When reviewing a table, software output, manuscript, or analysis report, use this sequence:
- Identify the effect being estimated. What population parameter or comparison does the estimate represent, and what does its sign mean?
- Read the effect estimate. What magnitude and direction were observed?
- Read the confidence interval. Is the estimate precise or uncertain? What substantively different parameter values are represented within the interval?
- Read the p-value in relation to the hypothesis actually tested. How much evidence does the test provide against that null hypothesis?
- Compare the estimate and interval with meaningful values. Does the interval contain only effects that would lead to similar practical conclusions, or does it span negligible, important, or even opposite-direction effects?
- Consider sample size. Could a very large sample have made a trivial effect statistically detectable? Could limited precision leave a potentially important effect unresolved?
- State the conclusion at the level the data support. Distinguish evidence of an effect from evidence about its magnitude and from a judgment about whether that magnitude matters (Moore et al., 2014, 2021; Devore & Berk, 2012; Newton & Rudestam, 1999; Lovric, 2011).
Common interpretation failures
“p < .05, therefore the effect is important.”
Statistical significance does not establish practical significance. Large samples can make very small effects statistically significant (Moore et al., 2021; Devore & Berk, 2012; Newton & Rudestam, 1999).
“p > .05, therefore there is no effect.”
Failure to reject the null hypothesis does not demonstrate that the null hypothesis is true. The confidence interval helps show which effect magnitudes remain compatible with the estimate and its uncertainty (Moore et al., 2021; Newton & Rudestam, 1999).
“The effect is 6 points.”
Six points is the point estimate, not a known population value. The confidence interval is needed to communicate the estimate's precision (Devore & Berk, 2012).
“The confidence interval includes zero, so there is nothing to interpret.”
An interval containing the null value can still be highly informative. A narrow interval concentrated around negligible effects tells a different story from a wide interval spanning substantial effects in both directions (Devore & Berk, 2012; Newton & Rudestam, 1999).
“The effect size is small, so it cannot matter.”
Practical importance is contextual. Effect-size measures quantify magnitude, but the substantive importance assigned to that magnitude depends on the research problem rather than being dictated automatically by statistical significance or an effect-size label (Newton & Rudestam, 1999; Lovric, 2011).
Why interpretation should not stop at p < .05
A threshold-based significance decision compresses a multidimensional result into one binary classification. That classification may be useful for a prespecified testing decision, but it cannot tell the researcher whether the estimated effect is positive or negative, large or small in practical terms, precisely or imprecisely estimated, or compatible with several substantively different conclusions.
The p-value, effect estimate, and confidence interval therefore work best as complementary pieces of evidence.
p-value
The p-value asks about evidence against the null.
Effect estimate
The effect estimate asks about direction and magnitude.
Confidence interval
The confidence interval asks about precision and the range of parameter values represented by the interval.
Practical interpretation
Practical interpretation asks what those magnitudes mean for the actual research problem.
A defensible statistical conclusion combines all four questions.
A result can be statistically convincing but practically trivial; practically promising but too imprecise to establish; or statistically nonsignificant yet precise enough to make a large practically important effect difficult to reconcile with the reported interval. Those distinctions disappear when interpretation ends at p < .05 (Moore et al., 2021; Devore & Berk, 2012; Newton & Rudestam, 1999; Lovric, 2011).
References
Devore, J. L., & Berk, K. N. (2012). Modern mathematical statistics with applications (2nd ed.). Springer.
Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2
Moore, D. S., McCabe, G. P., & Craig, B. A. (2014). Introduction to the practice of statistics (8th ed.). W. H. Freeman.
Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). Macmillan Learning.
Newton, R. R., & Rudestam, K. E. (1999). Your statistical consultant: Answers to your data analysis questions. SAGE Publications.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.