Resource

Five Outcomes, Five P-Values: A Clinical-Trial Multiplicity Case Study

Four of five endpoints have p < .05. Does that mean a treatment worked on four outcomes? This synthetic clinical-trial case shows how multiplicity adjustment, family-wise error control, hypothesis hierarchy, and gatekeeping can lead to very different—but methodologically coherent—conclusions.

A clinical trial can have one treatment comparison and still create a multiple-testing problem. The difficulty arises when the treatment is evaluated on several endpoints and each endpoint generates its own hypothesis test. If the analysis simply looks across the resulting p-values and highlights whichever ones fall below .05, the probability of making at least one false rejection across the family can exceed the error rate attached to any single test. This is the central multiplicity problem addressed by family-wise error rate methods (Dmitrienko & Koch, 2017; Lovric, 2011).

This fully synthetic case study shows why five apparently ordinary p-values cannot automatically be interpreted as five independent opportunities to declare success. It then compares a single-step adjustment with prespecified hierarchical and gatekeeping strategies.

Synthetic case: All treatments, endpoints, effect estimates, and p-values below are illustrative. They are not findings from an actual clinical trial.

The research question

Suppose investigators conduct a randomized clinical trial comparing Treatment A with placebo in adults with a chronic symptomatic condition.

The investigators have several questions about treatment benefit. Their main objective is improvement in overall symptom severity at Week 12, but they also want to examine physical functioning, sleep disturbance, fatigue, and a patient global assessment.

The five planned null hypotheses are:

Five planned endpoints and synthetic treatment results
Hypothesis Planned endpoint Role in the research objectives Synthetic treatment effect Raw p-value
H1 Symptom-severity score Primary −4.8 points .008
H2 Physical-function score Key secondary +3.7 points .018
H3 Sleep-disturbance score Key secondary −2.9 points .031
H4 Fatigue score Secondary −2.5 points .044
H5 Patient global assessment Secondary +1.6 points .071

Lower scores indicate improvement for symptom severity, sleep disturbance, and fatigue; higher scores indicate improvement for physical function and global assessment.

If each hypothesis is examined separately at α = .05, four results are nominally significant:

.008, .018, .031, and .044.

That creates an attractive but potentially misleading summary: Treatment A significantly improved four of five outcomes.

Whether that statement is justified depends on the inferential structure specified for the five tests.

Why five tests change the inferential problem

For one hypothesis, the Type I error rate is the probability of rejecting that null hypothesis when it is true. With a family of multiple hypotheses, the family-wise error rate (FWER) is the probability of incorrectly rejecting at least one true null hypothesis in the family. Testing every member of a family at its ordinary unadjusted significance level can inflate this overall error probability (Dmitrienko & Koch, 2017; Lovric, 2011).

The important unit is therefore not necessarily an isolated p-value. Researchers must ask what collection of hypotheses constitutes the inferential family and what conclusions they intend to draw from that family.

Weak FWER control

Weak control protects the error rate under the global null hypothesis, when all individual null hypotheses are simultaneously true.

Strong FWER control

Strong control limits the probability of incorrectly rejecting any true null hypothesis regardless of which and how many hypotheses in the family are true.

Dmitrienko and Koch (2017) distinguish weak and strong FWER control. This distinction matters when investigators want endpoint-specific conclusions rather than only a global statement that some treatment effect exists somewhere in the family (Dmitrienko & Koch, 2017).

Lovric (2011) similarly describes multiplicity as inflation of the family-wise Type I error rate when a family of hypotheses is tested simultaneously, and identifies multiple endpoints as one setting in which this issue arises.

The tempting analysis: report every raw p-value below .05

Suppose the trial team receives the five results without having committed to a multiplicity strategy.

A results meeting might quickly focus on this pattern:

  • symptom severity: p = .008;
  • physical function: p = .018;
  • sleep disturbance: p = .031;
  • fatigue: p = .044;
  • global assessment: p = .071.

Four p-values are below .05. The temptation is to present those four endpoints as four statistically established treatment benefits.

But that approach treats each p-value as though its inferential meaning were unaffected by the other planned tests. It does not control the FWER for the five-hypothesis family. The fact that investigators planned several opportunities to reject a null hypothesis is precisely what creates the multiplicity problem (Dmitrienko & Koch, 2017; Lovric, 2011).

The solution is not simply to search afterward for whichever adjustment preserves the most significant results. The appropriate strategy depends on the scientific objectives, the logical relationships among the hypotheses, available distributional information, and—when hypotheses have a meaningful hierarchy—the ordering specified for them (Dmitrienko & Koch, 2017).

Option 1: Treat all five endpoints equally with a single-step Bonferroni procedure

A straightforward approach is to place all five hypotheses in one family and apply a Bonferroni adjustment.

For m hypotheses and family-wise significance level α, Bonferroni tests each hypothesis at:

α / m

With five hypotheses and α = .05:

.05 / 5 = .01

Thus, an individual raw p-value must be no greater than .01 to be rejected under this procedure. Equivalently, Bonferroni-adjusted p-values can be compared with .05. The Bonferroni procedure provides strong FWER control without requiring a particular joint distribution for the marginal p-values (Dmitrienko & Koch, 2017). Lovric (2011) likewise describes Bonferroni as a simple method that divides α by the number of comparisons and controls the FWER.

Bonferroni-adjusted results for the synthetic trial
Endpoint Raw p-value Bonferroni-adjusted p-value Decision at family-wise α = .05
Symptom severity .008 .040 Reject H1
Physical function .018 .090 Do not reject H2
Sleep disturbance .031 .155 Do not reject H3
Fatigue .044 .220 Do not reject H4
Global assessment .071 .355 Do not reject H5

The interpretation changes substantially.

The symptom-severity result survives the five-test Bonferroni adjustment. The other four endpoints do not.

That does not mean the observed physical-function, sleep, or fatigue differences disappeared. Their estimates and raw p-values remain what they were. Rather, those raw p-values do not meet the decision rule required for individual rejections while controlling the FWER across this five-hypothesis family.

Why the Bonferroni result may be conservative

Single-step procedures test each hypothesis independently of the outcomes of the other hypothesis tests. Dmitrienko and Koch (2017) discuss Bonferroni, Šidák, and Simes methods in their treatment of single-step procedures and note that basic single-step procedures can have power limitations. Bonferroni can be conservative, although its strong FWER control does not depend on a particular joint distribution of the raw p-values (Dmitrienko & Koch, 2017).

Lovric (2011) similarly characterizes a tradeoff: simultaneous or single-step procedures control FWER but can have relatively low statistical power.

This is an important practical point. Multiplicity adjustment is not free. Tightening the evidence required for individual rejection generally makes it harder to detect genuine effects. A procedure should therefore reflect the research objectives rather than applying the same allocation of error to every hypothesis when the hypotheses clearly have different scientific roles.

The objectives suggest a hierarchy

In the synthetic trial, the five endpoints were not equally important.

The investigators' primary question was:

Does Treatment A improve symptom severity?

Only after establishing that result were they especially interested in broader benefits involving physical function and sleep, followed by fatigue and global assessment.

That scientific hierarchy contains information that an equal five-way Bonferroni split ignores.

Dmitrienko and Koch (2017) distinguish procedures with data-driven ordering from procedures in which hypothesis order is prespecified. Prespecified ordering can reflect the clinical importance of the analyses. Examples include fixed-sequence, fallback, and chain procedures (Dmitrienko & Koch, 2017).

The crucial word is prespecified. The hierarchy must express the intended research logic rather than being reconstructed after investigators see which endpoint happened to produce the smallest p-value.

Option 2: A prespecified fixed-sequence procedure

Suppose the investigators had specified this hierarchy before examining the treatment results:

H1 → H2 → H3 → H4 → H5

Under the fixed-sequence procedure described by Dmitrienko and Koch (2017), testing begins with the first hypothesis. Each hypothesis can be tested at the full α level provided all preceding hypotheses have been rejected. When a hypothesis fails, testing stops and the remaining hypotheses are accepted without testing for purposes of this procedure. This conditional sequence controls the FWER (Dmitrienko & Koch, 2017).

Fixed-sequence decisions for the synthetic trial
Step Hypothesis Raw p-value Testable at α = .05? Sequential decision
1 H1: Symptom severity .008 Yes Reject
2 H2: Physical function .018 Yes, because H1 was rejected Reject
3 H3: Sleep disturbance .031 Yes, because H1–H2 were rejected Reject
4 H4: Fatigue .044 Yes, because H1–H3 were rejected Reject
5 H5: Global assessment .071 Yes, because H1–H4 were rejected Do not reject

Under this particular prespecified hierarchy, four hypotheses can be rejected while maintaining the error-control structure of the fixed-sequence procedure (Dmitrienko & Koch, 2017).

Notice what made this possible: it was not the convenient discovery that four p-values happened to be below .05. It was the scientific ordering imposed on the hypotheses before their observed significance was used to choose the order.

What if the second endpoint had failed?

Consider a different synthetic result:

Alternative synthetic p-values for the fixed sequence
Hypothesis Raw p-value
H1 .008
H2 .083
H3 .004
H4 .006
H5 .009

The later three p-values look exceptionally compelling in isolation. Yet a strict sequence H1 → H2 → H3 → H4 → H5 stops when H2 is not rejected. H3–H5 would therefore not become successful confirmatory tests under that fixed-sequence procedure, despite their small raw p-values (Dmitrienko & Koch, 2017).

This is the price of concentrating testing opportunity according to a hierarchy: the procedure can be powerful when effects follow the anticipated order, but an early failure can block later hypotheses.

Dmitrienko and Koch (2017) specifically note that fixed-sequence procedures test at the full significance level as long as previous hypotheses are rejected, whereas failure to reject one hypothesis causes subsequent hypotheses to be accepted without testing under the procedure.

Option 3: Use a fallback strategy when the hierarchy is important but not absolute

Sometimes the investigators want to prioritize hypotheses without making every later hypothesis completely dependent on every earlier success.

Dmitrienko and Koch (2017) describe the fallback procedure as a more flexible method for hypotheses with a prespecified order. Error-rate weights are assigned to the hypotheses in advance. When an earlier hypothesis is rejected, its error rate can be carried forward. Unlike a strict fixed sequence, testing can continue after a nonsignificant result because later hypotheses may retain some initially assigned error rate (Dmitrienko & Koch, 2017).

This changes the design question.

Instead of asking only:

Which endpoint has the smallest p-value?

the researchers must ask:

Which hypotheses matter most, how should the available Type I error be allocated among them, and how should that error be transferred after a successful test?

The fallback procedure therefore trades some of the concentration available under a strict fixed sequence for greater flexibility if an earlier hypothesis fails.

Dmitrienko and Koch (2017) also describe chain procedures, which extend this idea by allowing prespecified weights and more flexible propagation of error after rejection, including transfer to several other hypotheses. Such procedures can represent more complex scientific decision structures than a single linear sequence.

Option 4: Gatekeeping when endpoints form distinct families

The five outcomes can also be organized into scientifically meaningful families.

Suppose the planned objectives are:

Primary family

H1: symptom severity

Key-secondary family

H2: physical function

H3: sleep disturbance

Additional-secondary family

H4: fatigue

H5: global assessment

This is no longer merely a list of five tests. It is a hierarchy of families.

Gatekeeping procedures are designed for multiplicity problems in which hypotheses are grouped into ordered families and access to a later family depends on results in earlier families. Earlier families therefore act as “gatekeepers” for later inference (Dmitrienko & Koch, 2017).

Serial gatekeeping

The later family becomes testable only after the required hypotheses in the preceding family have been rejected.

Parallel gatekeeping

Access to the later family can depend on at least one successful rejection in the preceding family.

Dmitrienko and Koch (2017) describe serial, parallel, and more general gatekeeping structures. In serial gatekeeping, the later family becomes testable only after the required hypotheses in the preceding family have been rejected. In parallel gatekeeping, access to the later family can depend on at least one successful rejection in the preceding family. The book also describes gatekeeping implementations built from component procedures such as Bonferroni, Holm, Hochberg, and Hommel procedures.

Lovric (2011) similarly describes gatekeeping procedures as methods that exploit ordering and logical relationships among hypotheses or families of hypotheses, with particular usefulness for multiple endpoints and combinations of multiplicity structures.

For this synthetic trial, a gatekeeping strategy could encode the substantive rule that secondary efficacy claims should only be pursued after the primary symptom hypothesis passes its gate. The exact decisions for H2–H5 would then depend on the prespecified component procedure and error-allocation rules within and between the families. It would be incorrect to invent those rules after seeing the five p-values.

Single-step versus ordered testing is a design decision, not a p-value contest

The same five raw p-values can lead to different valid conclusions because the procedures answer different structured inferential questions.

How different prespecified testing structures change the synthetic conclusions
Testing structure Synthetic conclusion
Five-way Bonferroni procedure Only H1 is rejected.
Strict hierarchy H1 → H2 → H3 → H4 → H5 H1 through H4 are rejected before the sequence stops at H5.
Prespecified gatekeeping structure Decisions depend on the component procedure and error-allocation rules specified within and between the endpoint families.

Those conclusions are not contradictory. The procedures encode different assumptions about the scientific relationships among the hypotheses and different allocations of the available Type I error.

Dmitrienko and Koch (2017) emphasize that selection among multiplicity procedures can depend on clinical information, including logical restrictions and dependencies among hypotheses, as well as statistical information, such as the joint distribution of test statistics. Their classification distinguishes single-step methods, procedures with data-driven ordering, and procedures with prespecified ordering.

The correct question is consequently not:

“Which correction gives us the most significant endpoints?”

It is:

“What family of conclusions was the study designed to support, and what testing structure represents those objectives?”

What about Holm and other data-driven stepwise procedures?

A hierarchy does not always have to be determined by scientific priority. Some stepwise procedures order hypotheses according to the observed test statistics or p-values.

Dmitrienko and Koch (2017) distinguish these data-driven hypothesis-ordering procedures from procedures whose ordering is specified in advance. Their discussion includes Holm, Hochberg, and Hommel procedures. Many stepwise multiple-testing procedures can be understood through the closure principle, which provides a general foundation for constructing procedures with strong FWER control (Dmitrienko & Koch, 2017).

Lovric (2011) likewise describes closed testing as a general framework for stepwise multiple-testing procedures and identifies Holm, Hochberg, and Hommel among modified Bonferroni procedures.

This distinction prevents two very different ideas from being confused:

Formal data-driven procedure

A predefined statistical algorithm is used to order observed p-values.

Outcome-driven cherry-picking

Results are examined first and investigators then informally decide which endpoints or hierarchy should count.

Only the first is a prespecified multiple-testing method.

The power consequence of multiplicity adjustment

Multiplicity control inevitably affects the ability to reject individual hypotheses.

A single-step Bonferroni procedure spreads the available error rate across the entire family. In this synthetic example, the per-hypothesis threshold becomes .01, causing three raw p-values between .01 and .05 to lose their rejection status.

More efficient procedures can sometimes improve power by exploiting additional structure. Dmitrienko and Koch (2017) describe stepwise extensions of basic single-step methods and procedures that use prespecified hypothesis ordering. They also discuss parametric procedures that can gain power by using information about the joint distribution of test statistics. Such gains require additional assumptions or structure; there is no universally superior multiplicity procedure independent of the problem being studied (Dmitrienko & Koch, 2017).

Lovric (2011) similarly notes the tradeoff between FWER control and power for simultaneous tests and describes stepwise procedures developed to improve on single-step approaches.

For study planning, this means that multiplicity and power should not be treated as unrelated topics. If success requires several adjusted hypothesis tests, the relevant testing strategy affects how difficult those successes will be to achieve.

Why pre-specification matters

The synthetic results make the importance of pre-specification especially visible.

After seeing:

.008, .018, .031, .044, .071

it would be easy to invent the sequence:

H1 → H2 → H3 → H4 → H5

and observe that four tests pass at .05.

But the evidential justification for an ordered procedure comes from a meaningful a priori ordering of the hypotheses—not from discovering an advantageous ordering after examining the data. Dmitrienko and Koch (2017) explicitly distinguish procedures based on prespecified hypothesis ordering and describe fixed-sequence ordering as normally reflecting the clinical importance of the analyses.

Pre-specification should therefore make the inferential architecture clear: which endpoints form the relevant family, which hypotheses have priority, whether testing is simultaneous or ordered, how α or hypothesis weights are allocated where applicable, when testing stops or continues, and how error can propagate through more complex procedures (Dmitrienko & Koch, 2017).

This is a methodological point rather than a statement about any current regulatory requirement.

Adjusted and unadjusted findings answer different questions

Unadjusted p-value

An unadjusted p-value describes the evidence produced by an individual test under its ordinary testing framework. It should not automatically be interpreted as though that test were the only inferential opportunity in a larger family.

Adjusted p-value

An adjusted p-value incorporates the decision rule of a particular multiplicity procedure. Dmitrienko and Koch (2017) define an adjusted p-value as the smallest significance level at which the corresponding hypothesis would be rejected under the specified multiple-testing procedure.

Consequently, the raw and adjusted values should not be presented as competing estimates of the treatment effect. The effect estimate itself does not become smaller merely because multiplicity is addressed. What changes is the inferential criterion used to decide whether an endpoint supports rejection while controlling the relevant family-level error rate.

For the synthetic physical-function endpoint, for example:

Estimated treatment effect

+3.7 points

Raw p-value

.018

Five-test Bonferroni-adjusted p-value

.090

A defensible interpretation is:

Physical function favored Treatment A, with an unadjusted p-value of .018, but the hypothesis was not rejected after the prespecified five-test Bonferroni adjustment (adjusted p = .090).

It would be misleading to report only “p = .018, statistically significant” if the confirmatory interpretation was supposed to use the five-test Bonferroni family.

Conversely, it would also be misleading to say that multiplicity adjustment demonstrated “no effect.” Failure to reject an adjusted hypothesis does not erase the observed effect estimate. It means that the evidence did not satisfy that procedure's rejection criterion.

What should the investigators conclude?

If the five endpoints had been prespecified as one equally weighted family analyzed with Bonferroni adjustment, the main inferential conclusion would be straightforward:

Treatment A produced evidence of improvement on the primary symptom-severity endpoint after controlling the FWER across the five planned tests. The other endpoints did not meet the Bonferroni-adjusted rejection criterion, even though three had raw p-values below .05.

If instead the clinical objectives had genuinely established the strict sequence H1 → H2 → H3 → H4 → H5 in advance, the interpretation would differ:

The first four hypotheses were rejected sequentially at α = .05, and the sequence stopped when the global-assessment hypothesis was not rejected.

A prespecified gatekeeping design could produce another result because it would represent a different set of scientific priorities and logical relationships among the endpoint families (Dmitrienko & Koch, 2017).

The observed data alone therefore do not determine the multiplicity strategy.

Practical checklist for multiple endpoints in a clinical trial

Before interpreting a table of endpoint p-values, researchers should establish:

  • Research objective: What treatment claims or scientific conclusions are actually being tested?
  • Hypothesis family: Which planned tests belong to the family for which error control is required?
  • Hierarchy: Are the hypotheses equally important, naturally ordered, or divided into ordered families?
  • Procedure: Does the problem call for a single-step procedure, a data-driven stepwise procedure, a prespecified ordered procedure, or a gatekeeping structure?
  • Assumptions: Does the selected procedure require information about dependence or the joint distribution of test statistics?
  • Power implications: How much testing opportunity does each hypothesis receive, and what happens to later hypotheses if an earlier test fails?
  • Pre-specification: Were ordering, weights, gates, and propagation rules established before the observed results were used to choose among them?
  • Reporting: Are raw and multiplicity-adjusted results clearly distinguished?
  • Interpretation: Are effect estimates kept separate from the binary decision to reject or not reject a hypothesis?

These questions reflect the broader principle that multiplicity procedures should be chosen to represent the inferential objectives and logical structure of the study rather than to rescue whichever p-values happen to look favorable (Dmitrienko & Koch, 2017).

Conclusion

Five outcomes do not merely produce five numbers. They create an inferential structure.

In this synthetic trial, four of five raw p-values were below .05. A simple results-first reading could therefore produce the headline that Treatment A improved four outcomes. But a five-test Bonferroni analysis supported only one individual rejection. A genuinely prespecified fixed sequence could support four. A gatekeeping strategy could organize the same endpoints into primary and secondary families and produce decisions according to yet another prespecified structure.

The lesson is not that Bonferroni is always preferable, nor that hierarchical testing is a way around multiplicity. Both are methods for controlling error under different inferential structures. Single-step procedures provide straightforward simultaneous protection but may sacrifice power. Ordered and gatekeeping procedures can use scientifically meaningful hierarchy more efficiently, but their interpretation depends on that hierarchy being established independently of the observed pattern of p-values (Dmitrienko & Koch, 2017; Lovric, 2011).

When a clinical trial has multiple endpoints, the decisive question comes before the p-values:

What conclusions was the study designed to test, and how were those conclusions organized?

References

Dmitrienko, A., & Koch, G. G. (Eds.). (2017). Analysis of clinical trials using SAS: A practical guide (2nd ed.). SAS Institute.

Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry