Type I Error, Type II Error and Statistical Power: A Researcher’s Practical Guide
Type I and Type II errors, alpha, beta, and statistical power are parts of one hypothesis-testing decision framework. This guide explains how they relate, why nonsignificant results do not prove the null hypothesis, and how power connects to effect size, variability, sample size, and study design.
Researchers often learn Type I error, Type II error, alpha, beta, P values, and statistical power as separate definitions. The more useful way to understand them is as parts of a single decision framework.
A hypothesis test can reject or not reject a null hypothesis. Either decision can be correct or incorrect depending on the true state represented by the statistical model. Type I and Type II errors describe those two possible errors; alpha and beta quantify their probabilities under specified conditions; and power describes the probability of detecting a specified alternative when it is true (Altman, 1991; Chow et al., 2018; Piantadosi, 2005; Rosner, 2016).
Key interpretive point: Failure to reject the null hypothesis is not proof that the null hypothesis is true. A study may simply have been too imprecise to distinguish an important effect from the null value (Altman, 1991).
Type 1 vs Type 2 error: the core framework
Start with the hypotheses.
The null hypothesis, H0, specifies the value or set of values against which the evidence is evaluated. In a conventional comparison, it commonly represents no difference or a specified reference value. The alternative hypothesis, H1 or Ha, specifies values that differ from the null in the direction or directions relevant to the research question. A two-sided alternative permits departures on either side of the null value, whereas a one-sided alternative concerns a specified direction (Rosner, 2016).
The decision structure can be summarized as follows (Chow et al., 2018; Piantadosi, 2005; Rosner, 2016):
| Statistical decision | Null hypothesis is true | Alternative is true |
|---|---|---|
| Reject H0 | Type I error | Correct detection |
| Do not reject H0 | Correct non-rejection | Type II error |
Statistical significance and truth are not the same thing. A significant result can occur when the null hypothesis is true, and a nonsignificant result can occur when an important alternative is true.
What is a Type I error?
A Type I error occurs when the null hypothesis is rejected even though it is true. Altman describes this as a false-positive result, while Chow et al. define it as rejecting H0 when H0 is true (Altman, 1991; Chow et al., 2018).
The probability of a Type I error is conventionally denoted by:
α
Thus, alpha is defined under the assumption that the null hypothesis is true.
If a testing procedure is conducted at a prespecified significance level α, that level controls the probability of the relevant Type I error under the assumptions of the test. Alpha is therefore an operating characteristic of the testing procedure—not the probability that a particular significant result is false and not the probability that the null hypothesis is true.
False positive does not mean “the P value is alpha”
Alpha and the observed P value have different roles.
Alpha is a criterion specified for the testing procedure. The P value is calculated from the observed data under the null hypothesis and is compared with that criterion when making a conventional reject/do-not-reject decision.
P < α
may lead to rejection of H0, but this does not mean that the probability that H0 is true equals P, nor that the probability the finding is a false positive equals P.
What is a Type II error?
A Type II error occurs when the null hypothesis is not rejected even though the relevant alternative is true. Altman characterizes this as a false-negative finding, and Chow et al. define it as failing to reject the null hypothesis when it is false (Altman, 1991; Chow et al., 2018).
Its probability is denoted by:
β
Unlike alpha, beta cannot be understood without specifying an alternative. The chance of missing a very small effect is not generally the same as the chance of missing a large effect.
A statement such as “the study had a beta of 0.20” is incomplete unless the effect, design, test, and other assumptions underlying that calculation are clear.
Alpha, beta and power are related—but answer different questions
Alpha (α)
Probability of rejecting H0 when H0 is true.
Beta (β)
Probability of not rejecting H0 under the specified alternative.
Power (1 − β)
Probability of rejecting H0 under that specified alternative.
Power is therefore:
Power = 1 − β
(Altman, 1991; Chow et al., 2018; Piantadosi, 2005).
For example, describing a design as having 90% power implies β = 0.10 for the particular alternative and design assumptions used in the calculation.
That final qualification is essential.
Statistical power explained: power is conditional on an effect
Statistical power is not a universal score attached to a study.
Piantadosi emphasizes that power is defined by assuming that a treatment effect of a specified magnitude is present and then considering the probability that the resulting test will declare an effect statistically significant. Consequently, a statement such as “this trial has 90% power” is ambiguous without specifying the effect against which that power was calculated (Piantadosi, 2005).
The planned test has 90% power to detect the prespecified effect under the stated assumptions and significance criterion.
Power can differ substantially across possible alternatives. An effect far from the null is generally easier to distinguish than an effect close to the null.
Why power is a design property—not the probability that a hypothesis is true
One of the most important misconceptions about statistical power is treating it as a probability statement about the truth of a hypothesis.
It is not.
Piantadosi explicitly distinguishes these ideas. Power assumes a treatment effect of a specified size and asks about the behavior of the statistical procedure across hypothetical repetitions of the experiment. It is neither the probability that the null hypothesis is true nor a probability statement about the unknown treatment effect itself (Piantadosi, 2005).
In frequentist study planning, therefore, power is best understood as a design operating characteristic:
P(reject H0 | specified alternative and design assumptions)
It does not reverse the conditioning to give:
P(alternative is true | observed data)
Those are fundamentally different probability statements.
This distinction also explains why post-result reasoning such as “the study had 80% power, so there is an 80% probability the alternative is true” is incorrect.
What determines statistical power?
Power cannot be separated from the effect being targeted, the amount of information in the study, and the variability of the outcome. Piantadosi specifically notes that power, study size, and the magnitude of the hypothetical treatment effect must be considered together (Piantadosi, 2005).
For conventional comparisons, several relationships are particularly important.
Effect size
Effects farther from the null are easier to distinguish statistically, other design features being equal.
A study therefore has greater power for a larger true difference than for a smaller true difference. This does not justify choosing an unrealistically large effect merely to make the sample-size calculation easier. Altman emphasizes planning around a difference that would be scientifically or clinically worthwhile (Altman, 1991).
Variability
Greater variability makes a given signal harder to distinguish.
For continuous outcomes, the standard deviation is therefore an important planning input. Altman describes sample-size planning for comparisons in terms of quantities including the clinically relevant difference and the expected standard deviation (Altman, 1991).
The relevant question is not simply “How large is the effect?” It is also “How large is that effect relative to the variation in the measurements?”
Sample size
Increasing sample size generally increases the information available to distinguish the specified alternative from the null.
Altman notes that beta depends on the effect of interest and the sample size and that sample size can be selected in advance to give a high probability of detecting an effect of a specified magnitude (Altman, 1991).
Chow et al. similarly frame prestudy power analysis as choosing the sample size required to achieve desired power for detecting a clinically or scientifically meaningful difference at a fixed Type I error rate (Chow et al., 2018).
Alpha
Changing alpha changes the rejection criterion.
With a fixed sample size, making the Type I error criterion more stringent generally makes rejection harder and affects power. Chow et al. note the trade-off between Type I and Type II errors at fixed sample size and explain that increasing sample size provides a way to reduce both error probabilities rather than simply trading one against the other (Chow et al., 2018).
A practical summary
Holding the other assumptions fixed:
| Change | Typical implication for power |
|---|---|
| Larger true effect relative to the null | Higher power |
| Greater variability | Lower power |
| Larger sample size | Higher power |
| More stringent rejection criterion | Lower power unless additional information is added |
| Alternative closer to the null | Lower power |
These relationships are why power and sample size cannot be discussed independently of the target effect and data-generating assumptions.
One-sided versus two-sided testing
The alternative hypothesis determines whether evidence is being sought in one direction or both.
Rosner defines a two-sided alternative as allowing the parameter under the alternative to lie either above or below the null value. In a conventional two-sided test, the Type I error probability is distributed across both tails of the relevant null distribution. A one-sided test instead focuses the rejection region in the prespecified direction (Rosner, 2016).
Sidedness affects power and sample-size requirements because it changes the rejection boundary. Chow et al. explicitly discuss different sample-size requirements for one-sided and two-sided procedures and show that changing from a two-sided to an appropriate one-sided test can reduce the calculated sample size (Chow et al., 2018).
Sample-size convenience is not a scientific justification for choosing a one-sided hypothesis. Directionality should follow the research objective and should be specified as part of study planning—not chosen after observing which direction the data happened to favor.
Why a nonsignificant result does not prove the null hypothesis
This is the most important interpretive consequence of Type II error.
Suppose a study produces:
P > α
The appropriate conventional decision is to not reject the null hypothesis.
That is not equivalent to demonstrating that the null hypothesis is true.
Altman emphasizes that a nonsignificant study may have low power and a sample size too small to detect a real difference. A wide confidence interval can remain compatible with substantial effects even though the test is nonsignificant (Altman, 1991).
Appropriate statement
“We did not detect convincing evidence against the null.”
Stronger claim requiring more evidence
“We demonstrated that there is no effect.”
The second statement requires much more than a nonsignificant P value.
Look at the confidence interval
A confidence interval makes the uncertainty that surrounds an estimate explicit.
A nonsignificant estimate with a narrow confidence interval excluding effects of practical importance conveys something very different from a nonsignificant estimate with a wide interval compatible with important benefit and important harm.
Altman argues against reducing results to the binary labels “significant” and “not significant” and emphasizes the additional information provided by confidence intervals about effect magnitude and uncertainty (Altman, 1991).
The practical question after a nonsignificant result is not merely “Was P > 0.05?” It is “Which scientifically important effects are still compatible with the estimate and its uncertainty?”
Power and sample size belong together
Power analysis is principally useful before the study, when investigators can still change the design.
How much information is required to give the study a prespecified probability of detecting an effect that would matter, while controlling the Type I error criterion?
Chow et al. describe prestudy power analysis precisely in these terms: selecting the required sample size to achieve the desired power for a clinically or scientifically meaningful difference at a fixed Type I error rate (Chow et al., 2018).
Piantadosi likewise describes comparative-trial planning by choosing an alternative corresponding to the smallest difference of clinical importance and selecting the study size to provide a high probability of rejecting the null if that alternative is true (Piantadosi, 2005).
For a simple continuous endpoint, the planning inputs commonly include:
Target difference → expected variability → alpha → desired power → sidedness → study design → required sample size
(Altman, 1991; Chow et al., 2018).
The exact formula depends on the endpoint and design, so there is no universal sample-size number associated with “80% power” or “90% power.”
Power is not meaningful without the target alternative
A common reporting problem is:
“Our study had 80% power.”
Power to detect what?
A study can have high power for a large effect and low power for a smaller but still scientifically important effect. Piantadosi explicitly cautions that statements about power made without the magnitude of the hypothetical treatment effect are ambiguous or meaningless (Piantadosi, 2005).
A defensible power statement should therefore identify, as applicable:
- the effect of interest;
- variability or event assumptions;
- sample size;
- alpha;
- sidedness;
- statistical test; and
- study design.
This makes the operating characteristic interpretable rather than presenting “power” as an isolated percentage.
Common misconceptions about type 1 vs type 2 error and statistical power
| Common statement | Why it is wrong | Correct interpretation |
|---|---|---|
| “A Type I error means the P value is wrong.” | Type I error concerns the decision to reject a true null hypothesis, not a computational error in the P value. | A Type I error is a false-positive rejection of H0. |
| “Alpha is the probability that the null hypothesis is true.” | Alpha is defined conditional on the null hypothesis being true. | α is the probability of rejecting H0 when H0 is true. |
| “P = 0.03 means there is a 3% probability the null hypothesis is true.” | A conventional P value is calculated under the null model; it is not the posterior probability of H0. | Interpret the P value according to its sampling definition, not as P(H0 | data). |
| “A Type II error is just an insignificant result.” | A nonsignificant result is a Type II error only when the relevant alternative is actually true. | Type II error is failure to reject H0 under the specified alternative. |
| “Beta is always 0.20.” | Beta depends on the specified alternative and design. | β is the Type II error probability for a particular alternative and analysis plan. |
| “80% power means there is an 80% probability the hypothesis is true.” | Power is conditional on an assumed alternative; it does not assign probabilities to hypotheses. | 80% power means an 80% probability of rejecting H0 under the specified alternative and assumptions. |
| “The study has 90% power.” | Power without an effect size or alternative is incomplete. | State the effect the study has 90% power to detect and the assumptions used. |
| “P > 0.05 proves there is no effect.” | Failure to reject H0 can result from inadequate information or substantial uncertainty. | A nonsignificant result means the chosen test did not reject H0; examine the effect estimate and confidence interval. |
| “Nonsignificant means the treatments are equivalent.” | Failure to establish a difference is not evidence that meaningful differences have been excluded. | Equivalence requires an appropriately formulated hypothesis, design, and inferential procedure. |
| “A larger sample reduces the chance of false positives by itself.” | The nominal Type I error is governed by the testing rule; increasing sample size principally changes precision and power under the planned procedure. | Sample size is chosen jointly with alpha, target effect, variability, power, and design. |
| “Use a one-sided test because it needs fewer participants.” | Sidedness must follow the scientific hypothesis, not recruitment convenience. | Prespecify a one-sided procedure only when the directional hypothesis is scientifically appropriate. |
| “Power should be calculated after a nonsignificant result to tell us whether the null is true.” | Power is not a probability that the null or alternative is true. Its main utility is in study design under prespecified alternatives. | Interpret the observed effect and its uncertainty; use power prospectively to plan the study. |
A practical researcher workflow
Before running a hypothesis test—or calculating sample size—work through the problem in this order:
1. Define the study problem
Research question → study design → primary endpoint
2. Specify the hypotheses
Null hypothesis → scientifically relevant alternative
3. Define the testing rule
One- or two-sided test → alpha
4. Specify information and power
Expected variability or event characteristics → desired power
5. Derive sample size
Required sample size follows from the preceding scientific and statistical specifications.
This order matters.
Sample size is the consequence of the scientific and statistical specification. It should not be the starting number around which the effect size, alpha, or power assumptions are subsequently manipulated.
Altman describes sample-size planning in terms of obtaining a high probability of finding a true effect of a specified magnitude, while Chow et al. emphasize alignment among the study objective, design, hypotheses, statistical test, and sample-size calculation (Altman, 1991; Chow et al., 2018).
What to report when a study is nonsignificant
Do not stop at:
“The result was not statistically significant.”
Instead, report and interpret the estimated effect and its uncertainty.
Ask:
- What was the estimated magnitude and direction of the effect?
- How wide is the confidence interval?
- Does the interval include scientifically or clinically important effects?
- Was the study designed with adequate information for the effect it was intended to detect?
- Does the interval rule out effects that would matter in practice?
A wide interval accompanying a nonsignificant result signals uncertainty. It may be compatible with the null value while also remaining compatible with important effects. In that setting, “no evidence of an effect” should not be converted into “evidence of no effect” (Altman, 1991).
Key takeaways
Type I error
Rejecting a true null hypothesis—a false positive.
Type II error
Failing to reject the null when the relevant alternative is true—a false negative.
Alpha and beta
Alpha (α) is the Type I error probability specified for the testing procedure. Beta (β) is the Type II error probability for a specified alternative.
Statistical power
Power is 1 − β: the probability that the planned procedure rejects the null when the specified alternative is true.
Power depends on the alternative effect, sample size, variability, testing criterion, sidedness, and study design. It is therefore a design operating characteristic, not the probability that a hypothesis is true.
Most importantly, a nonsignificant result does not prove the null hypothesis. Interpretation should focus on the estimated effect and its uncertainty, particularly whether the confidence interval excludes or remains compatible with effects that would matter scientifically or clinically.
Frequently Asked Questions
What is the difference between Type 1 and Type 2 error?
A Type I error occurs when a true null hypothesis is rejected; a Type II error occurs when the null is not rejected under the relevant alternative. They correspond to false-positive and false-negative decisions, respectively (Altman, 1991; Chow et al., 2018).
What is alpha in statistics?
Alpha, α, is the probability of a Type I error under the null hypothesis for the specified testing procedure. It is commonly used as the significance criterion against which the observed test result is evaluated (Altman, 1991; Chow et al., 2018).
What is beta in statistics?
Beta, β, is the probability of a Type II error under a specified alternative. Because beta depends on the alternative, it should not be interpreted without stating the effect the study is intended to detect (Altman, 1991; Piantadosi, 2005).
What is statistical power?
Statistical power is 1 − β: the probability of rejecting the null hypothesis when the specified alternative is true. It describes how the planned statistical procedure behaves under that assumed alternative (Chow et al., 2018; Piantadosi, 2005).
Is statistical power the probability that the alternative hypothesis is true?
No. Piantadosi explicitly distinguishes power from the probability that the null is true or from a probability statement about the treatment effect. Power assumes a specified alternative and calculates the probability of obtaining a statistically significant result under that assumption (Piantadosi, 2005).
Does P > 0.05 mean the null hypothesis is true?
No. It means that the chosen test did not reject the null hypothesis at the specified criterion. A nonsignificant result can occur even when a real and important effect exists, particularly when the study is imprecise or has low power for that effect (Altman, 1991).
What is the relationship between power and sample size?
For a specified effect, alpha, variability, test, and design, increasing sample size generally increases power. Conversely, sample size can be selected prospectively to achieve a desired power for detecting a prespecified meaningful effect (Altman, 1991; Chow et al., 2018).
Does a larger effect increase statistical power?
Yes, under otherwise fixed assumptions, alternatives farther from the null are easier for the test to distinguish and therefore have greater power. The planning effect should nevertheless be scientifically justified rather than selected merely to obtain a convenient sample size (Altman, 1991; Piantadosi, 2005).
Does greater variability reduce statistical power?
For conventional comparisons of continuous outcomes, greater variability makes a fixed difference harder to distinguish and therefore reduces power unless additional information is obtained. Expected variability is consequently an important input to power and sample-size calculations (Altman, 1991; Chow et al., 2018).
Is a one-sided test more powerful than a two-sided test?
For an appropriately specified directional alternative and the same nominal alpha, concentrating the rejection criterion in one direction changes the power and can reduce the required sample size relative to a two-sided procedure. But sidedness should be determined by the scientific hypothesis, not selected simply to increase power or reduce recruitment (Chow et al., 2018; Rosner, 2016).
References
Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.
Chow, S.-C., Shao, J., Wang, H., & Lokhnygina, Y. (2018). Sample size calculations in clinical research (3rd ed.). Chapman & Hall/CRC.
Piantadosi, S. (2005). Clinical trials: A methodologic perspective (2nd ed.). John Wiley & Sons.
Rosner, B. (2016). Fundamentals of biostatistics (8th ed.). Cengage Learning.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.