Resource

Noninferiority vs Equivalence vs Superiority Trials: What Researchers Commonly Get Wrong

Superiority, noninferiority, and equivalence trials answer different clinical questions and require different hypotheses, boundaries, and interpretations. This Resource explains why a nonsignificant superiority result does not establish noninferiority or equivalence and shows how prespecified margins and confidence intervals determine what a study can support.

A common clinical-trial interpretation error is deceptively simple:

“The treatments were not significantly different, so they are equivalent.”

That conclusion does not follow.

A superiority trial, a noninferiority trial, and an equivalence trial are designed to answer different questions. They use different hypotheses and decision boundaries, and those differences affect sample-size planning and interpretation. Chow et al. explicitly distinguish equality, noninferiority/superiority, and equivalence hypotheses and emphasize that valid sample-size calculations must correspond to the study objective and its statistical hypothesis (Chow et al., 2018).

The central mistake is therefore not merely statistical terminology. It is a mismatch between what the trial was designed to demonstrate and what researchers claim after seeing the results.

The three questions are fundamentally different

Suppose the treatment effect is defined as:

θ = new treatment − standard treatment

with larger values favoring the new treatment.

Superiority: Is the new treatment better?

A superiority analysis asks whether the treatment effect crosses a boundary representing no superiority. For a conventional comparison against no difference, that boundary is zero.

Conceptual hypotheses:

H0: θ ≤ 0

HA: θ > 0

A two-sided comparison may instead test equality against any difference, depending on the prespecified objective. The important point is that a superiority analysis seeks evidence of a difference in the favorable direction; it is not designed to establish that two treatments are sufficiently similar.

Noninferiority: Can an unacceptable loss of efficacy be excluded?

Noninferiority changes the boundary.

The new treatment may be allowed to perform somewhat worse than the active comparator, provided the loss does not exceed a prespecified amount considered clinically unacceptable. Chow et al. describe this quantity as the noninferiority margin, reflecting the degree of inferiority that the trial attempts to exclude (Chow et al., 2018).

If Δ > 0 represents the maximum acceptable loss under the sign convention above, the noninferiority boundary is −Δ.

Conceptual hypotheses:

H0: θ ≤ −Δ

HA: θ > −Δ

Rejecting the null supports the conclusion that the treatment is not worse than the comparator by more than the prespecified margin.

Noninferiority therefore does not require proof that the treatments have identical effects. It requires sufficiently precise evidence to exclude an unacceptable degree of inferiority.

Equivalence: Can clinically important differences in either direction be excluded?

Equivalence is two-sided.

Instead of asking only whether excessive inferiority can be excluded, an equivalence trial asks whether the treatment effect lies inside prespecified acceptable limits. For symmetric limits, these can be represented as:

−Δ < θ < +Δ

The null corresponds to a difference outside the acceptable equivalence region, whereas the alternative places the treatment effect inside it. Chow et al. describe equivalence as requiring the absolute treatment difference to remain within a difference of clinical importance and discuss equivalence testing through two one-sided procedures (Chow et al., 2018).

Comparison of the three trial objectives

What each trial objective requires the evidence to support
Trial objective What must the evidence support? Relevant boundary
Superiority Treatment is better Usually the no-effect boundary
Noninferiority Unacceptable inferiority can be excluded Prespecified noninferiority margin
Equivalence Clinically important differences can be excluded in both directions Lower and upper equivalence limits

The most common error: “Not significant” does not mean “equivalent”

Suppose a conventional superiority comparison produces P > 0.05.

What has been established?

Only that the superiority analysis did not provide sufficient evidence to reject its null hypothesis at the chosen significance level. It does not follow that the treatments are equivalent, and it does not automatically establish noninferiority.

Altman warns specifically against interpreting a nonsignificant result as evidence that there is “no difference.” Statistical significance used as a binary decision can obscure both the estimated magnitude of an effect and its uncertainty; confidence intervals provide substantially more information for interpretation (Altman, 1991).

The problem becomes especially serious for equivalence and noninferiority. Piantadosi notes that conventional comparative hypothesis-testing approaches are inadequate for these objectives because low power or low precision can favor an apparent conclusion of equivalence (Piantadosi, 2005).

An imprecise study can fail to find superiority precisely because it cannot distinguish important treatment differences from no difference. That lack of information cannot then be converted into evidence that important differences have been excluded.

The noninferiority margin is part of the scientific question

A noninferiority trial needs a clinically meaningful boundary.

Chow et al. emphasize that the choice of the clinically meaningful difference is critical in equivalence and noninferiority trials. Different choices can change both the calculated sample size and the eventual clinical conclusion. They also state that there is no universal rule for choosing this quantity and describe margin determination as requiring both statistical reasoning and clinical judgment (Chow et al., 2018).

This is why there is no defensible universal noninferiority margin.

A margin must be justified in relation to the endpoint, comparator, clinical context, evidence concerning the active treatment's effect, and what loss of efficacy would remain acceptable.

It is not merely a number inserted into a sample-size formula.

Why the margin must be determined during planning

The noninferiority margin defines what inferiority the study is intended to rule out. It also enters directly into the sample-size calculation.

Selecting or relaxing that boundary after observing the results would change the criterion the study was supposed to evaluate.

Planning logic

  1. Clinical question
  2. Acceptable loss
  3. Prespecified margin
  4. Hypothesis
  5. Sample size
  6. Data
  7. Confidence interval
  8. Conclusion

Do not reverse the planning sequence.

Data → confidence interval → choose a convenient margin → declare noninferiority

Equivalence limits perform a related but two-sided role

An equivalence trial requires an acceptable region rather than a single inferiority boundary.

With symmetric limits, the objective is to establish that:

−Δ < θ < +Δ

Both clinically unacceptable directions must therefore be excluded.

This is why “no statistically significant difference” is especially inadequate as evidence of equivalence. A nonsignificant superiority comparison can arise from a confidence interval that is very wide. Equivalence, by contrast, requires enough precision for the interval to remain inside the prespecified acceptable region.

One-sided versus two-sided reasoning

The hypothesis structure follows the research question.

A noninferiority question is directional: the principal concern is whether the new treatment is unacceptably worse than the comparator. Chow et al. formulate noninferiority/superiority testing as a one-sided problem relative to a prespecified margin (Chow et al., 2018).

Equivalence requires ruling out unacceptable differences on both sides. Chow et al. describe equivalence hypotheses and their evaluation using two one-sided tests (Chow et al., 2018).

These procedures cannot be treated as interchangeable versions of an ordinary superiority test. The hypotheses and rejection boundaries are different because the scientific claims are different.

Confidence intervals make the distinction visible

Confidence intervals provide a practical way to understand what a trial can support.

Assume again that larger effects favor the new treatment and that −Δ is the prespecified noninferiority boundary.

How confidence intervals relate to superiority, noninferiority, and equivalence
Objective Boundary or region What the confidence interval must establish
Superiority No-effect boundary, usually 0 The confidence interval excludes the no-effect boundary in the prespecified favorable direction, at a confidence level corresponding to the prespecified testing procedure.
Noninferiority Unacceptable-inferiority boundary, −Δ The relevant confidence bound excludes effects at or beyond the unacceptable-inferiority boundary. The interval does not necessarily need to exclude zero.
Equivalence Acceptable region from −Δ to +Δ Both relevant confidence limits must fall inside the prespecified equivalence limits under the planned equivalence procedure.

Superiority

Evidence supports superiority when the confidence interval excludes the no-effect boundary in the favorable direction, at a confidence level corresponding to the prespecified testing procedure.

Noninferiority

Evidence supports noninferiority when the relevant confidence bound excludes effects at or beyond the unacceptable inferiority boundary.

The interval does not necessarily need to exclude zero.

A result can therefore support noninferiority while failing to demonstrate superiority.

Piantadosi describes this logic in terms of requiring the appropriate confidence limit to remain beyond a clinically selected tolerance for inferiority rather than relying on the point estimate alone (Piantadosi, 2005).

Equivalence

For equivalence, the confidence interval must satisfy both acceptable boundaries under the prespecified equivalence procedure.

Conceptually:

−Δ < CIlower

and

CIupper < +Δ

Superiority

Boundary to exclude: 0

The confidence interval lies wholly on the prespecified favorable side of the no-effect boundary.

Noninferiority

Boundary to exclude: −Δ

The relevant confidence bound lies beyond the unacceptable-inferiority margin. Zero may still be inside the interval.

Equivalence

Region to satisfy: −Δ to +Δ

Both relevant confidence limits fall inside the prespecified equivalence limits.

The exact orientation reverses when lower values represent better outcomes. Researchers should therefore define the effect measure and favorable direction explicitly before applying these rules.

What conclusion is the study actually designed to support?

Use this decision framework before interpreting the P value.

Claim: “The new treatment is better”

Design: Superiority trial

  1. Was superiority demonstrated?
  2. If yes, a superiority conclusion may be supported.
  3. If no, conclude that superiority was not demonstrated.

Do not automatically conclude noninferiority or equivalence.

Claim: “The new treatment is not unacceptably worse”

Design: Noninferiority trial

  1. Was a clinically justified noninferiority margin prespecified?
  2. Was the trial designed and sized for that hypothesis?
  3. Does the relevant confidence bound exclude unacceptable inferiority?

Decision: If yes, noninferiority may be supported. If no, noninferiority is not demonstrated.

Claim: “The treatments differ by no more than an acceptable amount in either direction”

Design: Equivalence trial

  1. Were lower and upper equivalence limits prespecified?
  2. Was the trial designed and sized for equivalence?
  3. Are both relevant confidence limits inside the equivalence region?

Decision: If yes, equivalence may be supported. If no, equivalence is not demonstrated.

The key decision is not “Was P > 0.05?”

It is “Which effects has this design and confidence interval actually ruled out?”

Sample size follows the hypothesis—not the other way around

Superiority, noninferiority, and equivalence are not labels that can simply be attached after recruitment.

Chow et al. state that valid sample-size calculations depend on statistical tests appropriate to the hypotheses reflecting the study objectives, and that the different hypotheses have different sample-size requirements for achieving the desired power or precision (Chow et al., 2018).

For noninferiority calculations, the required sample size depends on quantities such as the significance level, desired power, anticipated treatment effect, outcome variability or event probabilities, allocation, and the noninferiority margin. The distance between the anticipated treatment effect and the inferiority boundary directly affects the information required (Chow et al., 2018).

Tighter margins demand greater precision

Holding other planning assumptions fixed, moving the noninferiority boundary closer to the expected treatment effect leaves less separation between the anticipated effect and the unacceptable boundary.

The study therefore needs greater precision to distinguish them.

For equivalence, both limits must be addressed. Chow et al. show explicitly that equivalence and one-sided noninferiority can have materially different sample-size requirements and recommend sensitivity analysis because changes in assumed treatment differences, variability, and margins or equivalence limits can substantially alter the required sample size (Chow et al., 2018).

Practical lesson: Do not power a superiority trial and assume that a nonsignificant result can later serve as an equivalence or noninferiority result.

Active-control noninferiority creates an additional problem

Noninferiority commonly compares a new treatment with an established active treatment rather than placebo.

That creates an inferential complication.

Piantadosi distinguishes superiority trials, which principally concern relative treatment effects, from noninferiority trials, which must address both relative and absolute treatment effects. If the control has an established absolute effect, the comparison is more direct. But if the standard treatment is merely presumed to outperform placebo, demonstrating similarity to that standard has limited meaning unless an appropriate degree of efficacy of the standard can also be supported (Piantadosi, 2005).

This matters because similarity between two treatments is not automatically reassuring.

If the active comparator performs weakly in the current setting, the new treatment may appear similar without providing adequate evidence that meaningful efficacy has been preserved.

Chow et al. likewise discuss the margin in relation to the effect that the active treatment could reliably be expected to have compared with placebo under conditions relevant to the planned trial, while emphasizing uncertainty and clinical judgment in margin selection (Chow et al., 2018).

An active-control noninferiority trial cannot be interpreted solely by asking whether the two randomized groups look statistically similar.

What researchers commonly get wrong

1. “The superiority test was nonsignificant, therefore the treatments are equivalent.”

Incorrect.

A nonsignificant result can reflect inadequate precision. It does not demonstrate that clinically important differences have been excluded (Altman, 1991; Piantadosi, 2005).

2. “The confidence interval includes zero, therefore noninferiority failed.”

Not necessarily.

Noninferiority is judged against the prespecified inferiority margin, not solely against zero. A confidence interval can include zero yet exclude unacceptable inferiority.

3. “Noninferiority and equivalence mean approximately the same thing.”

They do not.

Noninferiority excludes unacceptable loss in one direction. Equivalence requires acceptable differences to be established in both directions (Chow et al., 2018).

4. “We can decide on the margin once we see the data.”

That undermines the role of the margin as a planning-stage definition of clinically acceptable loss. Chow et al. emphasize that the choice is critical during study planning because it affects both sample size and conclusions (Chow et al., 2018).

5. “There should be a standard percentage for the noninferiority margin.”

There is no universal margin supported by these sources. Chow et al. explicitly state that there is no general “golden rule”; determination requires statistical reasoning and clinical judgment appropriate to the specific setting (Chow et al., 2018).

6. “The P value tells us what conclusion the trial supports.”

A P value is tied to a particular null hypothesis. Changing from superiority to noninferiority or equivalence changes the hypothesis itself.

Altman argues against reducing results to “significant” versus “not significant” and emphasizes confidence intervals for understanding the magnitude and uncertainty of treatment effects (Altman, 1991).

A practical interpretation checklist

Before writing “superior,” “noninferior,” or “equivalent,” check:

  • What was the prespecified primary trial objective?
  • What treatment effect and direction were defined?
  • Was the hypothesis superiority, noninferiority, or equivalence?
  • For noninferiority, was a noninferiority margin prespecified and clinically justified?
  • For equivalence, were acceptable lower and upper limits prespecified?
  • Was the sample size calculated for the actual hypothesis?
  • Is the directionality of the test appropriate to that hypothesis?
  • Does the confidence interval exclude the boundary or boundaries required for the intended conclusion?
  • For an active-control trial, is the comparator's efficacy sufficiently supported for the intended noninferiority interpretation?
  • Is the conclusion based on the effect estimate and its uncertainty, rather than merely whether P crossed 0.05?

How to report a failed superiority comparison

When a superiority trial produces a nonsignificant result, avoid:

“There was no difference between treatments.”

Also avoid:

“The treatments were equivalent.”

A more defensible structure is:

“The trial did not demonstrate superiority. The estimated treatment effect and confidence interval should be used to describe the range of effects compatible with the data. Unless the study was prospectively designed with an appropriate noninferiority margin or equivalence limits and the corresponding hypothesis, the result should not be interpreted as demonstrating noninferiority or equivalence.”

This keeps the conclusion aligned with the question the study was actually designed to answer.

The key takeaway

Noninferiority vs equivalence is not a semantic distinction, and neither conclusion is the fallback result of a failed superiority trial.

A superiority trial asks whether one treatment is better. A noninferiority trial asks whether an unacceptable degree of inferiority can be excluded. An equivalence trial asks whether unacceptable differences can be excluded in both directions.

Those questions require different hypotheses, boundaries, interpretations, and potentially different sample sizes (Chow et al., 2018).

Most importantly, failure to demonstrate superiority is not evidence of equivalence. Low precision can itself produce nonsignificance, which is exactly why equivalence and noninferiority require purpose-built hypotheses and clinically meaningful boundaries (Piantadosi, 2005).

The safest interpretive question is therefore:

What clinically important effects does the confidence interval actually exclude, relative to the boundary the study specified before observing the data?

That question moves interpretation away from P-value dichotomization and back toward the treatment effect, its uncertainty, and the clinical claim the trial was designed to support (Altman, 1991).

FAQs

Is a nonsignificant superiority test evidence of equivalence?

No. Failure to demonstrate superiority does not establish that clinically important differences are absent. A nonsignificant result can occur because the study is insufficiently precise. Equivalence requires a design and analysis capable of showing that the treatment effect lies within prespecified acceptable limits (Altman, 1991; Piantadosi, 2005).

What is the difference between noninferiority and equivalence?

Noninferiority asks whether unacceptable inferiority can be excluded in a specified direction. Equivalence asks whether clinically unacceptable differences can be excluded in both directions. The latter can be evaluated using two one-sided tests (Chow et al., 2018).

What is a noninferiority margin?

The noninferiority margin is the prespecified degree of inferiority relative to the comparator that the trial attempts to exclude. Its selection affects both interpretation and sample-size requirements and should reflect statistical reasoning and clinical judgment rather than a universal rule (Chow et al., 2018).

Can a trial be noninferior but not superior?

Yes. Because the noninferiority boundary differs from the no-effect boundary, a confidence interval may exclude unacceptable inferiority while still including no difference. Such a result can support noninferiority without demonstrating superiority, provided the trial was appropriately designed and the margin was prespecified.

Does a confidence interval have to exclude zero for noninferiority?

Not necessarily. The relevant comparison is with the prespecified noninferiority boundary. Zero is central to a conventional superiority comparison, whereas the noninferiority margin defines the unacceptable-loss boundary.

What confidence interval demonstrates equivalence?

Conceptually, equivalence requires both relevant confidence limits to fall within the prespecified equivalence region, using a confidence level consistent with the planned testing procedure. Simply having a confidence interval that includes zero does not establish equivalence.

Why can an underpowered trial falsely look “equivalent”?

An imprecise study may generate a wide confidence interval and a nonsignificant conventional comparison because it cannot distinguish among materially different effects. Piantadosi specifically warns that low power or low precision can favor an apparent equivalence conclusion when conventional hypothesis-testing logic is misapplied (Piantadosi, 2005).

How should a noninferiority margin be chosen?

The approved sources do not support a universal recommended margin. Chow et al. emphasize that margin selection is a critical planning decision requiring both statistical reasoning and clinical judgment and consideration of the evidence supporting the active comparator's effect (Chow et al., 2018).

References

Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.

Chow, S.-C., Shao, J., Wang, H., & Lokhnygina, Y. (2018). Sample size calculations in clinical research (3rd ed.). Chapman & Hall/CRC.

Piantadosi, S. (2005). Clinical trials: A methodologic perspective (2nd ed.). John Wiley & Sons.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry