When Should You Use a Nonparametric Test? A Practical Guide for Researchers
A significant normality test is not, by itself, a reason to switch automatically to a nonparametric procedure. This guide explains how to choose among parametric, rank-based, and transformed analyses by considering the research question, study design, inferential target, and assumptions.
A significant normality test is not, by itself, a reason to switch automatically from a t test or ANOVA to Mann–Whitney, Wilcoxon, Kruskal–Wallis, or Friedman.
That shortcut confuses two separate decisions: whether a normal-theory procedure is appropriate for the scientific question and data, and whether a particular nonparametric procedure answers the question you actually want to ask.
Nonparametric methods can be valuable when normal-theory assumptions are unsuitable, when observations can be meaningfully ordered but their numerical distances are less defensible, or when a rank-based location or distributional comparison matches the research objective. But nonparametric does not mean assumption-free. Rank procedures have their own requirements concerning independence, pairing, distributional structure, continuity, symmetry, or the form of the alternative, depending on the procedure and interpretation (Hollander et al., 2014).
The practical decision is not: “Was my normality test significant?”
It is: “What am I trying to estimate or compare, what is my study design, and which procedure has assumptions and an interpretation that fit that problem?”
The Key Principle: Choose the Question Before the Test
Before deciding when to use nonparametric tests, establish four things:
1. Target of inference
Are you interested in means, a location shift, a median, or a broader difference between distributions?
2. Study design
Are observations independent, paired, matched, repeated, or blocked?
3. Outcome information
Are numerical distances scientifically meaningful, or is ordering the more defensible information?
4. Procedure assumptions
What assumptions does the proposed procedure require?
This order matters because common parametric and nonparametric procedures do not necessarily test the same hypothesis or estimate the same quantity.
Rosner distinguishes nonparametric procedures from their normal-theory counterparts and notes that a major advantage is the ability to relax normality assumptions when those assumptions are unreasonable. He also identifies a potential cost: when normal-theory assumptions are appropriate, nonparametric procedures can lose power and typically replace original measurements with ranks (Rosner, 2016).
Hollander et al. similarly emphasize that many nonparametric methods use ranks rather than the original magnitudes and can be relatively insensitive to outlying observations. Their treatment also shows why the more precise term distribution-free is useful: under specified null hypotheses, the sampling distribution of some test statistics does not depend on a particular parametric population distribution such as the Normal distribution (Hollander et al., 2014).
What Does “Distribution-Free” Actually Mean?
“Nonparametric” is sometimes interpreted as meaning that the method makes no assumptions about the data. That is incorrect.
Hollander et al. explicitly describe the term nonparametric as imprecise and distinguish it from the mathematically more specific concept of a distribution-free procedure. For example, their Wilcoxon rank-sum framework starts with an independent random sample from one continuous population and an independent random sample from another. Under the null hypothesis, the two distributions are the same but that common distribution is left unspecified (Hollander et al., 2014).
The procedure therefore avoids specifying that the common distribution must be Normal, but it has not eliminated assumptions about sampling, independence, continuity, or the null hypothesis.
Key distinction: fewer or different distributional assumptions ≠ no assumptions.
Do Not Use a Normality Test as an Automatic Switch
Formal assessment of Normality can contribute to an analysis, but it should be combined with descriptive and graphical investigation.
Altman discusses graphical assessment using Normal plots and emphasizes that samples from Normal populations will not themselves look exactly Normal because of sampling variation. He also demonstrates that strongly skewed data may sometimes become approximately Normal after an appropriate transformation (Altman, 1991).
The consequence is important. A significant normality test identifies evidence against a particular Normal model; it does not determine which alternative analysis answers your scientific question.
Before switching methods, investigate:
- the distribution's shape and skewness;
- unusual or extreme observations;
- group-specific distributions;
- whether observations are paired or independent;
- whether quantitative distances are meaningful;
- whether a scientifically sensible transformation is available; and
- the assumptions and target of the alternative procedure.
Interpretation boundary: The question is not whether the data are perfectly Normal. The question is whether the planned analysis is appropriate for the inferential target and the observed data structure.
Parametric vs Nonparametric Decision Checklist
Use this checklist before replacing a parametric procedure with a rank-based test.
- State the scientific target. Do I want to compare means, locations, medians under an appropriate model, or distributions/rank tendencies?
- Identify the design. Are observations independent, paired, matched, repeated, or blocked?
- Check the measurement scale. Are numerical distances meaningful, or is ordering the main reliable information?
- Describe and plot the data. Examine center, spread, skewness, tails, ties, and unusual observations rather than relying only on a normality-test P value.
- Investigate unusual observations. Determine whether they are errors, observations from the wrong population, or legitimate extreme values.
- Evaluate the assumptions of the parametric candidate. Do not reduce this step to “Normality passed/failed.”
- Consider whether transformation is scientifically meaningful. A transformation can sometimes provide a more suitable analysis scale rather than requiring an immediate switch to ranks.
- Identify the exact nonparametric procedure required by the design. Independent and paired observations require different methods.
- Check the assumptions of that nonparametric procedure. In particular, do not assume that “nonparametric” means assumption-free.
- Check the intended interpretation. Does the test support a distributional statement, a location-shift statement, or a median statement under additional assumptions?
- Consider estimation as well as testing. Where supported, report an associated rank-based estimate and interval rather than only a P value.
- Make the final choice from question + design + estimand + assumptions, not from the normality-test result alone.
Which Nonparametric Test Fits Which Design?
| Research structure | Rank/nonparametric procedure | Core distinction |
|---|---|---|
| One sample or paired observations; location inference using signs | Sign procedure | Uses direction of differences |
| One sample or paired observations; location inference using signed ranks | Wilcoxon signed-rank | Uses direction and ranks of absolute differences |
| Two independent groups | Wilcoxon rank-sum / Mann–Whitney | Ranks observations jointly across groups |
| Three or more independent groups | Kruskal–Wallis | Rank-based one-way independent-groups comparison |
| Three or more related conditions or randomized blocks | Friedman procedure | Ranks treatments within the dependent/block structure |
Design comes first. These procedures are not interchangeable. The distinction between paired versus independent samples must be established from the study design before a test is selected.
Sign Test vs Wilcoxon Signed-Rank Test
The sign and signed-rank procedures are both useful for paired or one-sample location problems, but they do not use the data in the same way.
Sign procedures
For paired data, a sign procedure focuses on whether the difference within each pair is positive or negative. Rosner describes the sign test as requiring only determination of which member of a matched pair has the higher or lower score (Rosner, 2016).
Hollander et al. develop separate sign-based procedures for paired-replicate and one-sample location problems, including associated estimation and confidence-interval methods (Hollander et al., 2014).
Core distinction: The sign approach uses relatively limited information and does not use the magnitude of each nonzero difference in the way the signed-rank procedure does.
Wilcoxon signed-rank procedures
The Wilcoxon test for paired observations uses more information. It ranks the absolute within-pair differences and then incorporates their signs.
In Hollander et al.'s paired-replicate formulation, the differences are mutually independent and arise from continuous distributions symmetric about a common median treatment effect. The null hypothesis of zero treatment effect corresponds to the difference distributions being symmetric about zero (Hollander et al., 2014).
Core distinction: Additional use of magnitude ordering comes with additional structure.
This is an important example of why a Wilcoxon signed-rank test is not simply “the paired t test without assumptions.”
If symmetry of the difference distributions is central to the signed-rank interpretation, replacing a paired t test with signed ranks solely because a normality test rejected Normality may simply exchange one set of assumptions for another.
Mann Whitney vs t Test: What Question Are You Asking?
For two independent groups, the familiar choice is often framed as:
independent t test vs Mann–Whitney.
But the procedures should not be treated as identical analyses operating on different scales.
The two-sample t procedure
A two-sample t procedure is fundamentally a comparison involving population means. If the mean difference is the scientifically important effect, abandoning that target requires justification.
Wilcoxon–Mann–Whitney
The Wilcoxon rank-sum or Mann–Whitney procedure combines observations from the two independent groups, ranks them, and bases inference on the ranks. Hollander et al.'s distribution-free formulation assumes independent random samples from continuous populations; under the general null hypothesis, the two population distributions are equal but otherwise unspecified (Hollander et al., 2014).
Altman likewise presents Mann–Whitney as the nonparametric procedure for comparing two independent groups, based on ranking all observations as though they came from a single sample (Altman, 1991).
| Question | More directly aligned procedure |
|---|---|
| Are the population means different? | Two-sample t framework |
| Is a rank/distribution-based comparison appropriate for two independent groups? | Mann–Whitney/Wilcoxon rank-sum |
| Is a location-shift interpretation justified by additional distributional structure? | Mann–Whitney with that additional model explicitly defended |
Do not automatically call Mann–Whitney a test of medians
One of the most common interpretation errors in nonparametric statistics is:
“Mann–Whitney was significant, therefore the medians differ.”
That statement needs more support.
Altman describes a nonparametric confidence interval for a difference between medians that requires the restrictive assumption that the two population distributions have identical shapes and differ only by a location shift (Altman, 1991).
Hollander et al. likewise distinguish general distribution-free rank-sum testing from the additional location framework used for associated estimation (Hollander et al., 2014).
Therefore, when the groups differ in shape, spread, skewness, or other distributional features, a rank-based result should not automatically be translated into a pure statement about medians.
Rank-Based Estimation: Do More Than Report a P Value
Nonparametric analysis does not have to mean reporting only a test statistic and P value.
Hollander et al. explicitly pair several rank procedures with estimators and confidence intervals. Their one-sample/paired chapter includes Hodges–Lehmann estimators associated with signed-rank and sign statistics, while their two-sample location chapter includes an estimator associated with the Wilcoxon rank-sum statistic and a distribution-free confidence interval (Hollander et al., 2014).
Reporting principle: When the relevant location model and estimator are appropriate, accompany the hypothesis test with an effect estimate and uncertainty interval.
Do not describe that estimate more broadly than its assumptions justify.
When to Use Kruskal Wallis
The Kruskal Wallis procedure extends rank-based comparison to a one-way layout with multiple independent groups.
It is appropriate to consider when:
- there are three or more independent groups;
- the outcome can be meaningfully ranked;
- a rank-based comparison matches the scientific question; and
- the assumptions of the chosen Kruskal–Wallis interpretation are defensible.
Hollander et al. treat Kruskal–Wallis within the one-way layout, not as a generic solution to every failed ANOVA normality diagnostic. Their discussion also shows that the standard distribution-free property does not persist under arbitrary differences in population scale: allowing unequal scale parameters changes the problem and can require modified procedures (Hollander et al., 2014).
That point matters because different group spreads are not merely cosmetic. If groups differ in shape or scale as well as location, the interpretation of a rank-based omnibus result becomes more complicated.
A significant Kruskal–Wallis result should therefore not automatically be written as “At least one median differs.”
The inferential statement must reflect the assumptions under which a location or median interpretation is justified.
When to Use the Friedman Procedure
Kruskal–Wallis is for an independent-groups one-way structure. Friedman procedures address a different design: related observations organized through blocks or repeated conditions.
Hollander et al. develop Friedman procedures in their treatment of the two-way/randomized-block layout (Hollander et al., 2014).
Independent groups
Three or more independent groups → consider Kruskal–Wallis.
Related conditions or blocks
Three or more related conditions/blocks → consider Friedman-type procedures.
Using Kruskal–Wallis on repeated measurements because it is familiar discards the dependence structure. Conversely, using Friedman on unrelated groups invents a blocking structure that does not exist.
Nonparametric never overrides the design.
Why Nonparametric Does Not Mean Assumption-Free
The major procedures in this guide illustrate different assumptions.
| Procedure | Important requirements or interpretation boundaries |
|---|---|
| Mann–Whitney / Wilcoxon rank-sum | Requires an appropriate independent-sample structure and, in Hollander et al.'s basic distribution-free formulation, continuous populations. Stronger location or median interpretations require additional distributional structure (Hollander et al., 2014). |
| Wilcoxon signed-rank | Requires paired or one-sample structure and, in Hollander et al.'s paired-replicate formulation, mutually independent continuous difference variables symmetric about a common median treatment effect (Hollander et al., 2014). |
| Kruskal–Wallis | Is built for the independent-samples one-way layout, and its standard distribution-free interpretation does not simply remain unchanged when arbitrary scale differences are introduced (Hollander et al., 2014). |
| Friedman procedures | Require the block/repeated structure for which they were developed. |
So a nonparametric procedure can eliminate a Normal-distribution assumption while still requiring assumptions about:
sampling + independence/dependence + continuity + symmetry + distributional form under the null + treatment/block structure + interpretation.
That is why “assumption-free test” is a misleading description.
Where Transformations Fit Into the Decision
A non-Normal distribution does not force a choice between “use the raw data parametrically” and “rank everything.”
Transformation can sometimes provide another defensible analysis.
Altman illustrates strongly skewed bilirubin measurements that become approximately Normal after logarithmic transformation and discusses transformation as part of preparing data for statistical analysis (Altman, 1991).
A transformation should still have an analytical purpose. It changes the scale on which effects are estimated and interpreted. The relevant question is therefore not merely whether a transformation makes a normality test nonsignificant.
Ask instead: Does the transformed scale produce an analysis that is both statistically appropriate and scientifically interpretable?
If yes, transformation may preserve a useful parametric framework. If no, a rank-based or other method may be more appropriate.
A Practical Decision Path
When confronted with a significant normality test, use this sequence:
-
Define the target.
If the scientific question is explicitly about means, recognize that switching to ranks may change the question. -
Confirm the design.
Independent, paired, repeated, and blocked data require different procedures. -
Examine the observations descriptively.
Look at distributions, skewness, spreads, ties, and unusual values. Do not outsource this decision to one P value. -
Ask whether the parametric method remains defensible.
Assess its relevant assumptions rather than requiring the raw observations to look perfectly Normal. -
Consider a scientifically meaningful transformation where appropriate.
Reassess the distribution and interpretation after transformation. -
If considering a nonparametric method, identify its exact target and assumptions.
Do not choose it merely because it appears in software under “nonparametric tests.” -
Match the rank procedure to the design.
Mann–Whitney for two independent groups; signed-rank or sign procedures for appropriate paired/one-sample questions; Kruskal–Wallis for a one-way independent layout; Friedman for a block/repeated layout. -
Decide what you can actually claim.
Do not translate a general rank/distribution result into a median difference unless the additional assumptions needed for that interpretation are supported. -
Report estimation where supported.
Use an appropriate rank-based estimator and uncertainty interval when the model and procedure provide one.
Decision flow: research question → design → data structure → target of inference → descriptive investigation → assumptions → transformation or method choice → estimation and interpretation.
Common Mistakes
“Shapiro–Wilk is significant, so I must use Mann–Whitney”
A normality result does not identify the correct alternative procedure. Mann–Whitney has its own target and assumptions.
“Nonparametric means there are no assumptions”
Distribution-free procedures avoid specifying some aspects of the population distribution. They still rely on sampling and structural assumptions.
“Mann–Whitney tests medians”
Not automatically. A median/location interpretation requires additional distributional structure.
“Wilcoxon is always safer than a paired t test”
Wilcoxon signed-rank has its own assumptions, including symmetry in the paired-replicate location formulation described by Hollander et al. (2014).
“Kruskal–Wallis is simply ANOVA on medians”
It is a rank-based procedure with its own hypotheses and assumptions. Its interpretation should not automatically be converted into a statement about medians.
“Friedman is just Kruskal–Wallis for non-Normal data”
The crucial difference is data structure. Friedman addresses blocked/related observations; Kruskal–Wallis addresses independent groups.
“A nonparametric analysis means I only report a P value”
Rank procedures can have associated effect estimators and confidence intervals. Hollander et al. explicitly develop these for sign, signed-rank, and rank-sum location procedures (Hollander et al., 2014).
Bottom Line
The best answer to when to use nonparametric tests is not “whenever the normality test is significant.”
Use a nonparametric procedure when its inferential target, measurement requirements, study design, and assumptions are better aligned with the research problem than the relevant parametric alternative.
For two independent groups, that may mean Mann–Whitney rather than a t procedure. For paired observations, it may mean a sign or Wilcoxon signed-rank procedure. For several independent groups, Kruskal–Wallis may fit. For several related conditions or blocks, a Friedman procedure may fit.
But every one of those choices requires more reasoning than checking whether P < .05 in a normality test.
A stronger workflow is:
research question → design → data structure → target of inference → descriptive investigation → assumptions → transformation or method choice → estimation and interpretation.
That approach treats nonparametric statistics as a substantive analytical choice rather than an automatic fallback.
Frequently Asked Questions
Does a significant normality test mean I should use a nonparametric test?
No. A departure from Normality is one piece of diagnostic information. Examine the data, clarify the target of inference, evaluate the assumptions relevant to the planned parametric procedure, consider transformations where appropriate, and understand the assumptions and interpretation of the proposed nonparametric alternative (Altman, 1991; Hollander et al., 2014).
What is the difference between Mann Whitney vs t test?
A two-sample t framework directly addresses population means. Mann–Whitney/Wilcoxon rank-sum uses ranks from two independent samples and provides distribution-free inference under its specified null conditions. The procedures should not be assumed to test exactly the same population feature (Altman, 1991; Hollander et al., 2014).
When should I use the Wilcoxon test?
For paired or one-sample location problems, the Wilcoxon signed-rank procedure can be considered when its rank-based target and assumptions are appropriate. In the paired-replicate formulation developed by Hollander et al., the within-pair differences are independent, continuous, and symmetric about a common median treatment effect (Hollander et al., 2014).
When should I use Kruskal Wallis?
Consider Kruskal–Wallis for a rank-based comparison involving multiple independent groups in a one-way layout. Do not use it for repeated or matched observations, and do not automatically interpret its result as proof of different medians without the required distributional assumptions (Hollander et al., 2014).
When should I use the Friedman test?
Use a Friedman-type procedure when the design contains multiple related treatments or conditions organized through subjects or blocks. It is a rank-based approach for a dependent/block structure, not an independent-groups substitute.
Are nonparametric tests assumption-free?
No. The required assumptions differ by procedure. Independence, pairing/blocking, continuity, symmetry, and distributional structure can all matter. “Distribution-free” means that under specified conditions the procedure does not depend on a particular parametric distribution; it does not mean there are no assumptions (Hollander et al., 2014).
Can I report a median difference after Mann–Whitney?
Not automatically. Altman notes that nonparametric inference for a difference between medians under a shift model requires the restrictive condition that the population distributions have identical shapes and differ only in location (Altman, 1991).
Should I transform the data instead of using a nonparametric test?
Sometimes transformation is a reasonable alternative. Altman demonstrates that a highly skewed variable can become approximately Normal after logarithmic transformation. The transformation must still yield a scientifically meaningful analysis and interpretation rather than being performed solely to make a normality test nonsignificant (Altman, 1991).
References
Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.
Hollander, M., Wolfe, D. A., & Chicken, E. (2014). Nonparametric statistical methods (3rd ed.). Wiley.
Rosner, B. (2016). Fundamentals of biostatistics (8th ed.). Cengage Learning.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.