Resource

One-Tailed vs Two-Tailed Tests: Decide Before You See the Result

Learn how to decide between one-tailed and two-tailed tests based on the research question, how the choice changes rejection regions, p-values, and power, and why the direction must be specified before examining the result.

Choosing a one tailed vs two tailed test is not a way to make a borderline result statistically significant. It is a decision about the hypothesis being tested.

A one-sided test asks whether a parameter differs from its null value in a specified direction. A two-sided test asks whether it differs in either direction. Because those alternatives are different, they produce different rejection regions, different p-values, and different power properties. The direction should therefore come from the research question and be specified before examining the result—not chosen afterward because one analysis crosses a significance threshold (Moore et al., 2021; Newton & Rudestam, 1999).

This Resource provides a practical framework for deciding between a one sided vs two sided test, interpreting a one tailed p value or two tailed p value, and understanding what changes when the direction of the alternative changes.

Core principle: Choose the direction of the hypothesis from the research question before examining the estimate, test statistic, or p-value. Do not choose the tail afterward because one version produces statistical significance.

Start with the research question, not the observed statistic

Suppose a researcher is comparing a new programme with an existing programme. Let

Δ = μnew − μexisting.

The statistical test begins by specifying a null hypothesis and an alternative hypothesis. Hypotheses concern population parameters or models rather than the particular sample outcome that happens to be observed (Moore et al., 2021; Devore & Berk, 2012).

Nondirectional research question

If the research question is simply whether the programmes differ, the hypotheses can be written as:

H0: Δ = 0

HA: Δ ≠ 0

The alternative is nondirectional or two-sided: positive and negative departures from zero both count as evidence against the null hypothesis (Moore et al., 2021).

Directional research question

If the research question specifically concerns whether the new programme produces a higher population mean, the hypotheses can instead take the form:

H0: Δ ≤ 0

HA: Δ > 0

This is a directional, or one-sided, alternative.

Newton and Rudestam (1999) describe a one-tailed test as appropriate when the research hypothesis specifies the direction of the hypothesized difference; a two-tailed formulation states that a difference exists without specifying its direction. Meier et al. (2014) likewise distinguish a one-tailed test for a hypothesis that specifies a direction from a two-tailed test for a problem in which departure in either direction matters.

The practical distinction

Research questions and corresponding alternatives
Research question Alternative Test
Is the parameter different from the null value? HA: θ ≠ θ0 Two-sided
Is the parameter greater than the null value? HA: θ > θ0 One-sided, upper tail
Is the parameter less than the null value? HA: θ < θ0 One-sided, lower tail

The wording of the substantive question should determine which row applies—not the sign of the estimate after the data have been analyzed (Moore et al., 2021).

What changes when you choose one tail or two?

The choice changes which sample outcomes count as sufficiently strong evidence against H0.

A rejection region is the set of values of the test statistic for which the null hypothesis will be rejected. Devore and Berk (2012) describe a hypothesis-testing procedure in terms of a test statistic and a rejection region: H0 is rejected when the observed test statistic falls within that region.

Upper-tail test

Only sufficiently large positive values belong to the rejection region.

Lower-tail test

Only sufficiently large negative values belong to the rejection region.

Two-sided test

Sufficiently extreme values in either direction belong to the rejection region.

These distinctions follow from the alternative hypothesis (Devore & Berk, 2012; Moore et al., 2021).

The direction of the alternative is not cosmetic. It changes the formal rule used to decide what counts as evidence against the null hypothesis.

Alpha is allocated differently

The significance level α controls the probability of rejecting the null hypothesis when it is true under the specified testing procedure. Changing from a one-sided to a two-sided alternative changes where that probability is placed in the null distribution (Moore et al., 2021; Devore & Berk, 2012).

Consider a standard Normal test with α = .05.

Allocation of α in one-sided and two-sided standard Normal tests
Testing procedure Allocation of α Approximate critical value
Upper-tailed one-sided test The entire .05 rejection probability is placed in the upper tail. z = 1.645
Two-sided test The .05 significance level is divided between the two tails, giving .025 in each tail. z = −1.96 and z = 1.96

Devore and Berk (2012) formulate two-tailed rejection regions using upper and lower critical values based on α/2, whereas one-tailed tests place the rejection region in the direction specified by the alternative.

The practical consequence is important: with the same overall α, an observation does not have to travel as far into the specified tail to enter the rejection region of a one-sided test.

The tradeoff: That advantage is obtained by asking a narrower question.

One-tailed and two-tailed p-values answer different questions

A p-value is calculated assuming the null hypothesis and measures how extreme the observed test statistic is in the direction or directions specified by the hypotheses. The alternative hypothesis determines which outcomes count as “as extreme or more extreme” than the observed result (Moore et al., 2021).

For a positive observed z-statistic and an upper-tailed alternative,

HA: θ > θ0,

the one tailed p value is the probability in the upper tail at or beyond the observed statistic.

For a symmetric null distribution and a two-sided alternative,

HA: θ ≠ θ0,

extreme results in both directions count. Moore et al. (2021), for example, calculate a two-sided Normal p-value by adding the probabilities in the two symmetric tails; equivalently in that setting, it is twice the corresponding one-tail probability.

Do not turn this into a general rule to “divide the reported two-tailed p-value by two.” That calculation is appropriate only when the statistical setting and the prespecified directional alternative justify it. The hypotheses determine the p-value that should be calculated (Moore et al., 2021).

Numerical example: one-sided significant, two-sided not significant

Illustrative example: Consider a standard Normal test producing z = 1.75 with α = .05.

Directional hypothesis

Suppose first that the prespecified research question is directional:

H0: θ ≤ θ0

HA: θ > θ0.

The upper-tail probability beyond z = 1.75 is approximately

pone-sided ≈ .040.

Because .040 < .05, the result falls in the rejection region for the one-sided test. The researcher rejects H0 in favor of the prespecified positive-direction alternative.

Nondirectional hypothesis

Now suppose instead that the research question is nondirectional:

H0: θ = θ0

HA: θ ≠ θ0.

For the symmetric standard Normal distribution, the corresponding two tailed p value is approximately

ptwo-sided ≈ 2(.040) = .080.

Because .080 > .05, the result does not meet the .05 criterion for the two-sided test. The researcher therefore fails to reject H0.

Observed statistic

z = 1.75

One-sided p-value

p ≈ .040

Decision: Reject H0 at α = .05.

Two-sided p-value

p ≈ .080

Decision: Fail to reject H0 at α = .05.

Comparison of the illustrative one-sided and two-sided tests
Test Alternative Rejection criterion at α = .05 Observed z Approx. p-value Decision
One-sided HA: θ > θ0 z ≥ 1.645 1.75 .040 Reject H0
Two-sided HA: θ ≠ θ0 |z| ≥ 1.96 1.75 .080 Fail to reject H0

Why these conclusions are not contradictory

The conclusions differ because the hypotheses are different.

One-sided question

Is there sufficient evidence that θ is greater than θ0?

Two-sided question

Is there sufficient evidence that θ differs from θ0 in either direction?

Those are not the same inferential question. They assign the significance level differently, define different rejection regions, and calculate extremeness according to different alternatives (Moore et al., 2021; Devore & Berk, 2012).

It would therefore be misleading to describe the example as one statistical test giving two contradictory answers. Two different hypotheses were tested.

Fail to reject is not the same as accept. Because p = .080 > .05 in the two-sided test, the decision is to fail to reject the null hypothesis. It is not to “accept the null.” Newton and Rudestam (1999) explicitly distinguish failure to reject from accepting the null: failure to reject means that the evidence is insufficient for rejection, not that equality or absence of an effect has been proved.

What happens to power?

Power is the probability that a test rejects the null hypothesis when a specified alternative is true; equivalently, it is 1 − β, where β is the probability of a Type II error for that alternative (Newton & Rudestam, 1999; Devore & Berk, 2012).

The one-sided design concentrates its rejection region in the direction specified by the research hypothesis. At the same overall α, this produces a less extreme critical threshold in that direction than the corresponding symmetric two-sided test. Consequently, a properly specified one-sided test can have greater ability to detect alternatives in its designated direction. The tradeoff is that the testing rule is specifically constructed around that direction rather than departures on both sides (Devore & Berk, 2012).

This is the source of the apparent attraction of a one-sided test: if the true effect lies in the prespecified direction, concentrating the rejection region there can increase power.

The power advantage is legitimate only when the narrower directional hypothesis is the research question that was intended to be tested.

Power itself also depends on the particular alternative being considered. Effects farther from the null are easier to distinguish, and sample size, variability, significance level, and the alternative effect all influence a study's ability to reject the null (Moore et al., 2021; Devore & Berk, 2012; Newton & Rudestam, 1999).

When to use a one tailed test

The defensible question is not “Would a one-tailed test give me significance?” It is:

Before seeing the data, was the research hypothesis genuinely directional?

A one-sided test is most defensible when the research question specifies a direction in advance and the inferential claim being tested is genuinely about that direction. A two-sided test is appropriate when departures in either direction answer the research question or when no specific direction is firmly justified before examining the data (Meier et al., 2014; Moore et al., 2021; Newton & Rudestam, 1999).

Moore et al. (2021) make the timing requirement especially explicit: the alternative should express the hopes or suspicions brought to the data, and framing the alternative after looking at the observations to match what they show is improper. If a specific direction is not firmly established beforehand, they instruct the researcher to use a two-sided alternative.

Newton and Rudestam (1999) similarly emphasize objectivity in significance testing and discuss selecting the cutoff before “peeking at the results.”

Why “I predicted the direction” is not enough after the fact

Suppose a researcher originally asks whether an intervention changes an outcome. The data then show a positive estimate with a two-sided p-value of .08. Recasting the analysis afterward as a positive one-sided test because the resulting p-value is approximately .04 changes the question after seeing which direction makes the evidence look strongest.

The problem is not that .04 was calculated incorrectly. The problem is that the directional alternative was selected using information from the result it is supposed to test.

That undermines the prespecified interpretation of the significance level. Moore et al. (2021) explicitly warn against looking at the data first and then framing the alternative to fit the observations. More generally, their discussion of the use and abuse of significance tests cautions against searching for significance (Moore et al., 2021).

A directional theoretical expectation can justify a one-sided test when it genuinely defines the research question beforehand. An observed favorable direction cannot retroactively create that justification.

A practical decision framework

Before running the test, document these decisions:

  1. What population parameter is being tested? State the parameter or contrast clearly.
  2. What is the null hypothesis? Specify the null value or boundary.
  3. What substantive claim would answer the research question? Is the question about any difference, or specifically an increase or decrease?
  4. Would a result in the opposite direction still be scientifically or practically relevant to the question? If so, a two-sided alternative is usually aligned with that question.
  5. Was the direction specified before examining the result? Do not choose the tail after seeing the estimate, test statistic, or p-value.
  6. What significance level will be used? State α before testing and recognize that one- and two-sided procedures allocate it differently.
  7. What p-value corresponds to the actual alternative? Interpret the p-value for the hypothesis that was prespecified rather than selecting whichever version is smaller.
  8. What is the correct decision language? Reject H0 when the evidence satisfies the prespecified criterion; otherwise fail to reject H0. Do not translate p > α into “accept the null” (Moore et al., 2021; Devore & Berk, 2012; Newton & Rudestam, 1999).

Decision sequence: Research question → hypotheses → direction → alpha → test → p-value → interpretation.

Common mistakes

Choosing the smaller p-value afterward

“The two-tailed p-value is .08, so I will report the one-tailed p-value of .04.”

That is defensible only if the one-sided alternative was justified independently of the observed result. Choosing it because the two-sided result was nonsignificant reverses the proper order of hypothesis specification and testing (Moore et al., 2021).

Choosing one tail only for power

“A one-tailed test is better because it has more power.”

It concentrates the rejection region in one specified direction, which can improve power there, but it does so because it tests a narrower directional hypothesis. Power is not a justification for changing the scientific question (Devore & Berk, 2012).

Using the observed sign as justification

“My theory predicts a positive effect, so any positive estimate warrants a one-sided test.”

The relevant issue is whether the directional hypothesis defined the test before the data were examined. The observed sign itself cannot supply the justification (Moore et al., 2021; Newton & Rudestam, 1999).

Treating different tests as inconsistent

“The one-sided test is significant but the two-sided test is not, so statistics are inconsistent.”

No. The tests have different alternatives, different rejection regions, and different p-values. They are answering different questions (Moore et al., 2021; Devore & Berk, 2012).

Accepting the null after p > .05

“The two-sided test gave p > .05, so we accept the null.”

No. The appropriate conclusion is that the test failed to reject the null hypothesis at the chosen significance level. Failure to reject does not establish that the null is true (Devore & Berk, 2012; Newton & Rudestam, 1999).

The decision rule to remember

The choice between a one tailed vs two tailed test belongs at the hypothesis-design stage.

Use a one-sided alternative when the research question itself is genuinely directional and that direction is specified before examining the result. Use a two-sided alternative when the question concerns differences in either direction or when a directional alternative cannot be justified in advance (Meier et al., 2014; Moore et al., 2021; Newton & Rudestam, 1999).

A one-sided test can produce a smaller p-value and greater power in its specified direction because the rejection region is concentrated in that tail. That is not a statistical shortcut. It is the consequence of asking a narrower question (Devore & Berk, 2012).

Use this sequence:
Research question → hypotheses → direction → alpha → test → p-value → interpretation.

Do not use this sequence:
Result → p-value → choose the tail that makes it significant.

References

Devore, J. L., & Berk, K. N. (2012). Modern mathematical statistics with applications (2nd ed.). Springer.

Meier, K. J., Brudney, J. L., & Bohte, J. (2014). Applied statistics for public and nonprofit administration (9th ed.). Cengage Learning.

Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). Macmillan Learning.

Newton, R. R., & Rudestam, K. E. (1999). Your statistical consultant: Answers to your data analysis questions. SAGE Publications.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry