Resource

Multiple Comparisons in Research: Why Repeated Testing Increases False-Positive Findings

Learn why repeated testing creates more opportunities for false-positive findings and how to define the inferential family, distinguish confirmatory from exploratory analyses, choose an appropriate error-control strategy, and report multiple comparisons transparently.

Testing one prespecified hypothesis and testing 20 outcomes, subgroups, predictors, time points, or pairwise differences are not the same inferential problem.

The multiple comparisons problem arises when researchers make several statistical inferences but interpret each result as though it were the only test performed. As the number of opportunities to reject null hypotheses increases, so does the opportunity to obtain apparently significant findings by chance. Multiple-comparison methods therefore consider a collection of inferences jointly rather than assigning the usual error criterion independently to every comparison (Agresti, 2013; Altman, 1991; Hollander et al., 2014).

The practical question is not simply “Which multiple testing correction should I use?”

How many questions am I testing? → Are they primary or exploratory? → Were the comparisons prespecified? → What error rate needs to be controlled? → What adjustment and interpretation match those objectives?

This Resource provides a decision framework for answering those questions without turning multiplicity analysis into a catalog of correction procedures.

What Is the Multiple Comparisons Problem?

A conventional significance level applies to a specified hypothesis test. Once a study generates many inferential decisions, researchers also need to consider the error behavior of the collection of decisions.

Hollander et al. describe this directly for all-pair treatment comparisons. With k treatments, an all-treatment procedure makes k(k−1)/2 pairwise decisions. Their experimentwise approach controls the probability of at least one incorrect decision under the global null at the specified α level (Hollander et al., 2014).

Agresti makes the same distinction through simultaneous confidence intervals: multiple-comparison methods apply the confidence level to the simultaneous set of comparisons, rather than separately to every individual interval (Agresti, 2013).

Per-comparison inference

Asks about the error criterion for one test or interval.

Family-level or experimentwise inference

Asks about errors across a specified collection of inferential decisions.

If the scientific conclusion depends on searching across many tests and selecting whichever results cross a significance threshold, interpreting each raw p-value as an isolated test ignores that broader search.

Why More Testing Creates More Opportunities for False Positives

The problem is not that performing many analyses mechanically makes every individual p-value invalid. The problem is that the researcher now has more opportunities to obtain at least one apparently significant result by chance.

Altman warns against “data-dredging”: performing large numbers of analyses in the hope that something interesting will emerge. He notes that apparent relationships can arise by chance when no corresponding relationship exists in the population and argues that principal objectives and analyses should therefore be identified in advance. Exploratory analyses can be useful for generating hypotheses, but should not automatically be treated as confirmatory evidence (Altman, 1991).

The same issue occurs in more structured settings. Multiple outcomes, subgroup analyses, several predictors, repeated time-point comparisons, and many pairwise group comparisons can all create multiple inferential opportunities. The statistical strategy should therefore follow the number and organization of conclusions the study intends to support.

The Error Rate Must Match the Question

Before selecting a multiple testing correction, decide what error criterion matters.

For many multiple-comparison problems, the relevant target is an overall or family-level probability of making at least one false rejection. Hollander et al. use the term experimentwise error rate for procedures designed to control the probability of one or more incorrect decisions across the collection under the null hypothesis (Hollander et al., 2014).

Agresti similarly describes controlling a family of inferences. His Bonferroni presentation constructs a collection of confidence intervals so that the probability of at least one coverage error across the set has an upper bound of α (Agresti, 2013).

Core distinction: Does .05 apply to each individual comparison, or to the inferential family as a whole?

Those are different operating characteristics.

Step 1: Count the Questions You Are Actually Testing

Multiplicity often becomes visible only after the research question is translated into statistical decisions.

Ask how many distinct claims may emerge from the analysis:

Research structures that can create multiple inferential claims
Research structure Multiplicity question
Several outcomes Will each outcome generate a separate inferential claim?
Several treatment groups Are all pairwise differences being tested?
Several time points Will treatment effects be tested separately at each time?
Several subgroups Will each subgroup produce its own treatment-effect claim?
Several predictors Are many individual associations being searched for and highlighted?
Several contrasts Were these specific comparisons defined before examining results?

Pairwise comparisons can expand especially quickly. Hollander et al. explicitly formulate all-treatment procedures around k(k−1)/2 pairwise decisions (Hollander et al., 2014).

Important: The solution is not automatically to adjust every p-value produced anywhere in a study as one enormous family. The first task is to define which hypotheses belong to the inferential question whose error rate needs protection.

Step 2: Separate Primary Questions From Exploratory Questions

A study with one clearly defined primary question and several exploratory analyses has a different inferential structure from a study claiming confirmatory evidence for every measured outcome.

Altman argues that main research objectives and principal analyses should be identified in advance. Analyses generated through broad searches of the data should instead be recognized as exploratory and, where appropriate, used to generate hypotheses for future investigation (Altman, 1991).

This does not mean that exploratory analyses are statistically worthless. It means their interpretation should reflect how the questions arose.

Confirmatory question

The study was designed to provide a formal inferential conclusion about this hypothesis.

Exploratory question

The analysis is being used to identify patterns or hypotheses that may deserve further study.

That distinction should be established from the research objectives—not reconstructed after seeing which p-values are smallest.

Step 3: Were the Comparisons Prespecified?

Prespecification reduces the temptation to let the observed results determine which questions become important.

Altman specifically discourages large numbers of comparisons because they can indicate poorly specified research objectives. In his discussion of comparisons following analysis of variance, he demonstrates adjustment of pairwise p-values and recommends against indiscriminate large-scale comparison (Altman, 1991).

Prespecification does not imply that multiple planned hypotheses can always be treated as one isolated test each. If several prespecified comparisons will each support separate confirmatory claims, the relevant multiplicity structure still needs to be considered.

What prespecification does provide is a defensible answer to a different question:

Were these comparisons chosen because they represented the research objectives, or because they happened to look interesting after the data were examined?

That distinction matters for both analysis and reporting.

Step 4: Decide Whether You Need Pairwise Comparisons at All

Researchers often create unnecessary multiplicity by automatically comparing every group with every other group.

If there are k groups, all-pair comparisons require k(k−1)/2 decisions. But not every research question requires all of them.

Hollander et al. distinguish all-treatment multiple comparisons from treatment-versus-control procedures. When the scientific purpose is to compare several treatments with a control, there may be no initial reason to perform comparisons among the noncontrol treatments themselves (Hollander et al., 2014).

Design principle: Do not enlarge the multiple-testing family by asking comparisons that the scientific objective does not require.

A study interested in whether three interventions outperform standard care does not automatically need every intervention-versus-intervention comparison.

Step 5: Do Not Treat an Omnibus Test as the Answer to Every Pairwise Question

An omnibus procedure and a multiple-comparison procedure answer different questions.

An overall test can establish evidence against a common null hypothesis without identifying which particular comparisons account for that result. Hollander et al. explicitly describe multiple-comparison procedures as going beyond deciding whether treatments are equivalent to determining which treatments differ from one another (Hollander et al., 2014).

A common analysis workflow

Overall question → omnibus analysis → scientifically justified follow-up comparisons → multiplicity-aware inference

But a significant omnibus result does not remove the multiple-comparison problem. Hollander et al. note that experimentwise procedures may remain conservative even when an earlier omnibus test has already supplied evidence against treatment equivalence (Hollander et al., 2014).

The follow-up analysis still needs an inferential strategy appropriate to the comparisons being made.

Step 6: When Is a Bonferroni Correction Reasonable?

The Bonferroni correction is valuable because its rationale is simple and general.

Suppose g inferences form a family for which the desired overall error probability is α. Agresti describes the Bonferroni method as assigning each inference an error probability:

α* = α / g

For confidence intervals, using the corresponding more stringent confidence level for each interval produces simultaneous coverage of at least 1−α. Equivalently, the probability of at least one error across the collection is bounded above by α (Agresti, 2013).

Altman demonstrates the same basic logic for pairwise comparisons after analysis of variance by multiplying the individual p-value by the number of comparisons being made (Altman, 1991).

Bonferroni logic

More inferential opportunities → stricter criterion for each opportunity → protection of the error rate across the collection.

Bonferroni is not automatically the optimal procedure for every multiple-testing problem. Agresti explicitly describes it as a simple, multipurpose but somewhat conservative approach (Agresti, 2013).

Use it because its family-level error protection matches the inferential problem—not merely because it is the correction available in the software menu.

Step 7: Consider Simultaneous Confidence Intervals, Not Only Adjusted P-Values

Multiplicity is not solely a p-value problem.

Agresti describes simultaneous confidence intervals as applying the confidence level to a collection of comparisons jointly. With Bonferroni intervals, the individual intervals use a stricter confidence level so that simultaneous coverage across the family is at least the desired level (Agresti, 2013).

This is useful because research interpretation should not stop at whether a threshold was crossed.

Where meaningful effect estimates are available, simultaneous intervals can communicate:

  • the estimated direction and magnitude of each comparison;
  • its uncertainty; and
  • the stronger uncertainty requirement created by making several inferential statements jointly.

That is often more informative than reporting a column of adjusted significance labels alone.

Step 8: Recognize the Error-Control–Power Trade-Off

Stronger protection against false-positive findings has a cost.

Hollander et al. explicitly describe experimentwise error control as conservative. Protecting the probability of making only correct decisions under the null makes it more difficult to declare individual treatment differences statistically significant when differences really exist, and this difficulty becomes more severe as the number of treatments increases (Hollander et al., 2014).

Central trade-off: Stronger family-level false-positive protection ↔ harder individual rejection decisions.

Agresti likewise characterizes Bonferroni as somewhat conservative and notes that less conservative multiple-comparison approaches exist in particular settings (Agresti, 2013).

This does not justify ignoring multiplicity to preserve significance. Instead, it reinforces the importance of defining the research questions efficiently.

  • Do not test 20 questions when only three matter scientifically.
  • Do not create all pairwise comparisons when only comparisons with a control answer the research objective.
  • Do not redefine exploratory findings as primary after seeing the results.

Reducing unnecessary hypotheses is a design strategy, not a statistical trick.

Step 9: Build Multiplicity Into Clinical-Research Planning

Multiplicity should be considered before the results table exists.

Clinical-study planning connects the prespecified significance criterion to statistical power and sample-size requirements. Chow et al. describe prestudy power analysis as determining the sample size needed to achieve desired power for detecting a clinically meaningful difference at a prespecified significance level (Chow et al., 2018).

This matters for multiple testing because a multiplicity strategy that changes the effective significance criterion can also affect the amount of information needed to achieve the desired inferential performance.

Planning sequence

Define the primary questions → identify the hypothesis family → specify the error criterion → choose the multiplicity strategy → evaluate power/sample-size implications.

Do not calculate sample size around one inferential structure and then expand the confirmatory analysis into many additional claims after data collection.

Step 10: Do Not Selectively Report the Significant Comparisons

Selective reporting compounds the multiple-comparisons problem because readers can no longer see how many opportunities existed to obtain the highlighted result.

Altman identifies both multiple comparisons and selective reporting of only significant results as statistical problems that may be overlooked when published research is assessed (Altman, 1991).

Suppose a manuscript reports:

“Outcome 7 differed significantly between groups, p=.03.”

That statement means something very different if Outcome 7 was the single prespecified primary outcome than if it was the only nominally significant finding among 30 outcomes, subgroups, and time-point comparisons.

Defensible reporting should therefore make the search space visible. Readers need to know what was tested, which analyses were primary or exploratory, and what multiplicity strategy—if any—was used.

A Researcher-Facing Multiple Comparisons Decision Guide

Use this sequence before interpreting a collection of p-values.

1. How many questions am I testing?

List the actual inferential claims: outcomes, groups, time points, subgroup effects, predictors, contrasts, and pairwise comparisons.

If there is only one prespecified inferential question, the classical multiple-comparisons problem may not arise.

If there are several, continue.

2. Are these questions primary or exploratory?

Primary or confirmatory

Define the inferential family and error-control objective prospectively.

Exploratory

Report them as exploratory rather than treating every nominal p < .05 as independently confirmatory (Altman, 1991).

3. Were the comparisons prespecified?

Yes: retain the focused set that corresponds to the research objectives.

No: recognize that choosing comparisons after examining the data increases the need for cautious interpretation. Do not disguise data-driven comparisons as planned hypotheses.

4. What error rate matters?

Ask whether the scientific claim concerns:

  • one individual comparison; or
  • a family of inferential decisions for which protection against one or more false rejections is required.

If family-level control is required, ordinary unadjusted thresholds for every comparison do not represent that objective.

5. What adjustment or interpretation is justified?

For a fixed collection of comparisons, a Bonferroni-type approach provides simple family-level protection but may be conservative (Agresti, 2013).

For structured pairwise or treatment-versus-control questions, use a multiple-comparison strategy that corresponds to that comparison structure rather than automatically applying every available test (Hollander et al., 2014).

Where available, report simultaneous confidence intervals alongside or instead of relying solely on adjusted p-values.

6. What should I report?

Research objectives → hypotheses tested → primary versus exploratory status → prespecified versus data-driven comparisons → inferential family → error-control strategy → effect estimates → uncertainty intervals → adjusted or otherwise appropriate inferential results.

Do not report only the comparisons that achieved statistical significance.

Common Multiple-Comparison Mistakes

“Every P < .05 is a separate discovery”

Not when those p-values were obtained from a broader search and the conclusions depend on finding one or more significant results across that search. Family-level error may be the relevant inferential criterion (Agresti, 2013; Hollander et al., 2014).

“The omnibus test was significant, so the pairwise differences are established”

An omnibus rejection and identification of individual differences are different inferential tasks. Multiple comparisons are used precisely to address the latter (Hollander et al., 2014).

“We should compare every possible pair”

Only if every pair is part of the research objective. Treatment-versus-control questions, for example, may require fewer comparisons than all-treatment pairwise analysis (Hollander et al., 2014).

“Bonferroni is always the correct solution”

No. It is a simple, broadly applicable method for controlling a family of inferences, but Agresti describes it as somewhat conservative (Agresti, 2013).

“Adjustment ruined the result”

The observed effect estimate did not disappear because the inferential criterion became more stringent. The relevant issue is whether the evidence is strong enough for the family-level conclusion the study intends to make.

“We only need to report the significant comparisons”

Selective reporting hides the number of opportunities available for obtaining a significant finding and is itself a methodological problem (Altman, 1991).

Practical Reporting Checklist

Before submitting a paper or finalizing an analysis, confirm that:

  • The primary research questions are explicitly identified.
  • The number and structure of inferential comparisons are clear.
  • Primary and exploratory analyses are distinguished.
  • Prespecified and data-driven comparisons are distinguished.
  • The relevant inferential family has been defined.
  • The intended error criterion is stated.
  • Any multiple testing correction is justified by that error criterion.
  • Pairwise comparisons correspond to questions the study actually needs to answer.
  • Effect estimates are reported, not merely significance labels.
  • Simultaneous confidence intervals are reported where supported and useful.
  • The power implications of stringent multiplicity control have been considered.
  • Nonsignificant as well as significant planned comparisons are reported.
  • Exploratory findings are not presented as though they were the sole prespecified hypothesis.

The Key Principle

The multiple comparisons problem is fundamentally a problem of inferential opportunity.

A researcher who tests many outcomes, groups, predictors, subgroups, time points, or pairwise contrasts has created more opportunities for chance findings than a researcher testing one prespecified question. Family-level procedures address this by evaluating a collection of inferences jointly; Bonferroni provides one simple, conservative mechanism, and simultaneous confidence intervals provide an estimation-oriented counterpart (Agresti, 2013; Hollander et al., 2014).

But correction is only part of the solution.

The stronger strategy begins earlier:

Define the research questions → limit unnecessary testing → distinguish primary from exploratory analyses → prespecify important comparisons → define the relevant error rate → choose a matching inferential procedure → report the complete set transparently.

That approach controls false-positive opportunities without reducing statistical practice to a mechanical search for whichever correction leaves the most results below .05.

Frequently Asked Questions

What is the multiple comparisons problem?

The multiple comparisons problem arises when several statistical inferences are made and each is interpreted as though it were the only one. When conclusions depend on a family of comparisons, researchers may need to control the probability of one or more false-positive decisions across that family rather than only the error rate of an individual test (Agresti, 2013; Hollander et al., 2014).

Why does multiple hypothesis testing increase false-positive findings?

Multiple testing creates more opportunities for an apparently significant result to occur by chance. Altman specifically warns that large-scale data searching can reveal apparent relationships even when no real population relationship exists (Altman, 1991).

What is family-wise error?

In the framework used by the approved sources, family-level or experimentwise error concerns making one or more incorrect decisions across a specified collection of comparisons. Multiple-comparison procedures can be designed to bound that probability at a chosen α level (Agresti, 2013; Hollander et al., 2014).

How does the Bonferroni correction work?

For g inferences and desired family error probability α, Bonferroni assigns each inference an error probability of α/g. Agresti shows that the same principle can produce simultaneous confidence intervals with joint coverage of at least 1−α (Agresti, 2013).

Is Bonferroni too conservative?

It can be. Agresti describes Bonferroni as a simple multipurpose but somewhat conservative method, while Hollander et al. show more generally that strong experimentwise protection can make individual differences harder to detect, particularly as the number of treatments grows (Agresti, 2013; Hollander et al., 2014).

Do planned comparisons need multiple testing correction?

Prespecification does not automatically remove multiplicity when several separate confirmatory claims are being made. What it does accomplish is to distinguish scientifically planned questions from comparisons selected after examining the data. The appropriate error-control strategy still depends on the inferential family and the claims the study is intended to support.

Should I run pairwise comparisons after a significant omnibus test?

Only when identifying particular differences answers the research question. An omnibus procedure does not itself identify which groups differ, but the follow-up comparison structure should reflect the scientific objective rather than automatically testing every possible pair (Hollander et al., 2014).

Why is reporting only significant comparisons problematic?

It conceals how many opportunities existed to obtain those significant results. Altman explicitly identifies both multiple comparisons and selective reporting of significant results as methodological problems in published research (Altman, 1991).

References

Agresti, A. (2013). Categorical data analysis (3rd ed.). Wiley.

Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.

Chow, S.-C., Shao, J., Wang, H., & Lokhnygina, Y. (2018). Sample size calculations in clinical research (3rd ed.). Chapman & Hall/CRC.

Hollander, M., Wolfe, D. A., & Chicken, E. (2014). Nonparametric statistical methods (3rd ed.). Wiley.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry