Non-Normal Data: Should You Switch to a Nonparametric Test?
A significant Shapiro–Wilk test does not automatically mean a researcher should switch to a nonparametric test. This synthetic case shows how the research question, measurement scale, design, skewness, outliers, robustness, and the hypothesis tested by Mann–Whitney should guide the decision.
A researcher compares a quantitative outcome between two independent groups. The histograms are right-skewed, several observations sit far above the main body of the data, and Shapiro–Wilk tests are significant in both groups.
The researcher concludes:
“Shapiro–Wilk is significant, therefore I must use a nonparametric test.”
That conclusion is too mechanical.
A significant normality test tells the researcher that the data provide evidence against an exactly Normal population model. It does not, by itself, determine the research question, measurement level, design structure, target parameter, seriousness of the departure, influence of unusual observations, or robustness of the candidate parametric procedure. Parametric-versus-nonparametric selection therefore requires more than converting one diagnostic p value into a test-selection rule (Moore et al., 2021; Newton & Rudestam, 1999).
This synthetic case shows why.
The synthetic study
Synthetic example: The study, data, and reported numerical results in this Resource are constructed for illustration.
A community health service wants to compare time to resolve a routine service request under two intake systems:
System A
Standard intake.
System B
Redesigned intake.
Sixty independently sampled requests are observed under each system. Resolution time is recorded in days.
The research question is:
Do resolution times differ between requests handled under System A and System B?
For this synthetic example, the outcome is treated as a quantitative continuous variable. Each request contributes one observation and belongs to only one system, so this is an independent-groups design rather than a paired or repeated-measures design.
Those facts matter before normality is considered. Statistical-test selection should reflect the research question, measurement properties, and study structure rather than beginning with a normality test (Adams & Lawrence, 2018; Newton & Rudestam, 1999).
What the data look like
The synthetic data are deliberately right-skewed and include several unusually long resolution times.
| System | n | Mean days | SD | Median days | Sample skewness |
|---|---|---|---|---|---|
| A | 60 | 6.33 | 4.35 | 5.69 | 2.69 |
| B | 60 | 5.58 | 3.96 | 4.50 | 1.98 |
Three System A observations are especially high at approximately 18, 22, and 26 days. Two System B observations are approximately 17 and 21 days.
The corresponding synthetic Shapiro–Wilk results are:
- System A: W ≈ .728, p < .001
- System B: W ≈ .800, p < .001
So there is no reasonable argument that these particular samples look Normal. The mistake would be turning that observation into the automatic conclusion:
Non-Normal → Mann–Whitney.
The distributions, unusual observations, sample sizes, estimand, and assumptions of both candidate procedures still need examination.
First question: what are we trying to compare?
Suppose the substantive question is specifically about mean resolution time.
That makes the independent-samples t framework directly relevant because the two-sample t procedure compares population means. A rank-based alternative does not simply reproduce the same mean comparison without requiring normality (Moore et al., 2014, 2021).
Changing the statistical test can change the question being tested.
If the practical target is the expected number of resolution days—and the mean is substantively meaningful—then abandoning a mean-based procedure requires more justification than “Shapiro–Wilk was significant.”
Conversely, if the outcome were genuinely ordinal and only order information could defensibly be interpreted, a rank procedure could be much more natural. Adams and Lawrence (2018), for example, present rank-based nonparametric procedures for ordinal outcomes and distinguish the relevant tests according to whether groups are independent or dependent.
Measurement therefore belongs near the beginning of the decision, not after the normality test.
What assumptions matter for the two-sample t procedure?
Normality is not the only consideration.
For an independent-groups comparison, the sampling or design must support independence between observations. Independence cannot be established by Shapiro–Wilk; it comes from understanding how the observations were generated (Moore et al., 2021).
Distributional behavior also matters, but the practical consequences of non-Normality depend on features such as sample size, skewness, outliers, and balance between the groups.
Moore et al. (2021) give deliberately conservative guidance for two-sample t procedures. With very small combined samples, clearly non-Normal data or outliers are problematic. With intermediate sample sizes, outliers or strong skewness remain concerns. For large combined samples, t procedures can be used even with clearly skewed distributions, although unusual observations remain important because means and standard deviations are sensitive to them (Moore et al., 2021).
The same source notes that two-sample t procedures are particularly robust against non-Normality when group sizes are equal (Moore et al., 2021).
Newton and Rudestam (1999) similarly warn against both extremes: researchers should not automatically reject parametric methods whenever data depart from normality, but neither should they assume that every parametric procedure is robust under every violation.
So the right diagnostic question is not:
“Did normality pass?”
It is:
“Are the departures serious enough, given this design and this particular procedure, to make the intended inference unreliable or poorly aligned with the research question?”
Shapiro–Wilk is evidence about distributional form—not a test-selection algorithm
Formal normality tests can be useful components of diagnostic work. Shapiro–Wilk is one formal procedure for assessing departures from normality (Lovric, 2011).
But diagnostics should not stop there. Graphical examination is also informative when evaluating distributional assumptions, and unusual observations deserve separate investigation because they can affect statistical analyses in ways that a generic normality decision does not fully characterize (Lovric, 2011; Newton & Rudestam, 1999).
In this case, the significant Shapiro–Wilk results are consistent with what the histograms, skewness, and unusually high observations already show: the distributions are not Normal.
The remaining question is what that fact does to the proposed analysis.
That is a robustness question, not another normality question.
The unusual observations deserve their own investigation
The values at 18–26 days should not automatically be deleted simply because they make the distribution less Normal.
First ask what they are.
Data-entry errors?
Determine whether the observations were recorded incorrectly.
Wrong population?
Determine whether the measurements belong to the population being studied.
Valid rare delays?
They may be legitimate observations carrying substantive information.
A genuinely long right tail?
The underlying service process may actually produce unusually long delays.
Outliers may reflect errors, but they may also be legitimate observations carrying substantive information. Data should therefore be investigated rather than mechanically altered to make a statistical assumption look better (Adams & Lawrence, 2018; Newton & Rudestam, 1999).
That distinction is especially important here.
If a 26-day resolution really occurred, deleting it because “the boxplot called it an outlier” changes the phenomenon being summarized. Long delays may be precisely what makes one service system practically problematic.
Distinguish data quality from statistical inconvenience.
What happens if we retain the observations?
Using all 120 synthetic observations, a Welch two-sample t analysis gives approximately:
Welch two-sample t
t(117.0) = 0.98
p = .329
Estimated mean difference
6.33 − 5.58 = 0.74 days
For this constructed dataset, there is not strong evidence against equality of the population means.
This calculation does not prove that the t analysis is universally safe whenever n = 60 per group. It illustrates a more defensible process: identify the estimand, examine the data, assess the assumptions and robustness relevant to that procedure, and then interpret the result subject to those conditions.
What if we use Mann–Whitney instead?
The same synthetic observations produce approximately:
Mann–Whitney result
U = 2106
p = .109 for a two-sided comparison.
The conclusion is again not statistically significant at the conventional .05 level.
It would be tempting to say:
“Both tests agree, so it does not matter which one we use.”
But that is not quite right.
The two procedures do not generally test identical hypotheses.
What does Mann–Whitney actually test?
The Wilcoxon rank-sum procedure—also called the Mann–Whitney test—replaces the original measurements with their ranks. Its general null hypothesis concerns equality of the two population distributions, with alternatives concerning one distribution having systematically larger values than the other (Moore et al., 2014, 2021).
That is not automatically a test of two population medians.
Moore et al. (2014, 2021) make the distinction explicit: interpreting the Wilcoxon/Mann–Whitney procedure as a comparison of population medians requires an additional condition that the two population distributions have the same shape, apart from a location shift. Without that condition, the more general interpretation concerns differences between the distributions or systematically larger observations.
This is particularly relevant in a case involving skewness and unusual observations.
If System A has a heavier right tail than System B, a Mann–Whitney result cannot simply be relabeled as:
“The medians differ.”
A rank test can respond to broader distributional differences.
That may be useful—but it means the estimand has changed from the mean difference targeted by the t procedure.
Nonparametric does not mean assumption-free
Calling Mann–Whitney “nonparametric” does not remove the need to understand the design and data.
The observations still need an appropriate independent-samples structure for the conventional independent-groups rank comparison. Ties also matter because tied observations alter rank calculations and the sampling distribution used for inference; statistical software applies appropriate adjustments (Moore et al., 2014, 2021).
Interpretation as a median comparison requires still more structure: the population distributions must have comparable shapes so that a location-shift interpretation is defensible (Moore et al., 2014, 2021).
Nonparametric procedures are statistical models and tests with their own conditions. They are not a methodological escape hatch from assumptions.
Newton and Rudestam (1999) also emphasize an important cost of rank transformation: once raw quantitative observations are replaced by ranks, information about the distances between observations is discarded.
For example, resolution times of 5, 6, and 25 days become ordered ranks. The fact that 25 is dramatically farther from 6 than 6 is from 5 is no longer represented in the same way.
Whether that is acceptable depends on the research question.
Why the t test can be more robust than the researcher expects
The word robust does not mean that an assumption is false but irrelevant.
It means that an inferential procedure may continue to perform reasonably under some departures from its idealized assumptions.
Moore et al. (2021) explicitly describe two-sample t procedures as robust to non-Normality under appropriate conditions, with robustness improving with larger and more balanced samples. They nevertheless retain strong warnings about outliers and, at smaller sample sizes, substantial skewness.
The International Encyclopedia of Statistical Science likewise discusses the robustness of ANOVA-type inference to some departures from normality while emphasizing that the consequences depend on the nature and combination of violations (Lovric, 2011).
So “the population is not exactly Normal” and “the parametric inference is unusable” are different statements.
A significant Shapiro–Wilk test establishes neither equivalence.
But robustness is not permission to ignore the data
The opposite mistake is also possible:
“The t test is robust, so I never need to worry about normality or outliers.”
That is no better.
Newton and Rudestam (1999) explicitly caution against automatically relying on robustness. Small samples, unequal group sizes, severe non-Normality, and other assumption problems can make parametric procedures less defensible. Their practical recommendation is to examine the variables and their distributions and to consider nonparametric procedures when distributions are seriously non-Normal or when the measurement structure better supports rank-based analysis.
Outliers deserve particular attention because the sample mean and standard deviation can be strongly influenced by extreme observations.
Parametric procedures can be robust to some non-Normality, but robustness has limits and must be evaluated in the context of the actual data and design.
A sensitivity analysis clarifies the issue
Suppose the analyst investigates the five observations above 15 days and verifies that all five are legitimate service delays.
They should normally remain in the primary dataset.
For diagnostic purposes only, however, the analyst can ask how strongly the inferential conclusion depends on those unusual observations.
Removing values above 15 days from both groups in this synthetic sensitivity calculation gives:
- System A: n = 57
- System B: n = 57
- Welch t: approximately t = 1.23, p = .220
- Mann–Whitney: approximately p = .087
The substantive conclusion remains similar.
This does not justify deleting the observations. Instead, it tells the analyst that the broad inferential conclusion in this synthetic case is not being produced solely by those five observations.
Sensitivity analysis is more informative than silently removing inconvenient data.
When Mann–Whitney may be the better choice
A rank-based comparison becomes more compelling when the research question and measurement support it—for example, when the outcome is ordinal, when a meaningful ordering is available but quantitative distances are not, or when a small-sample quantitative dataset is so clearly non-Normal that the intended t inference is poorly supported (Adams & Lawrence, 2018; Moore et al., 2021).
Two independent groups
Candidate rank procedure: Mann–Whitney/Wilcoxon rank-sum.
Three or more independent groups
Candidate rank procedure: Kruskal–Wallis.
Paired observations
Candidate rank procedure: Wilcoxon signed-rank, subject to its own interpretive conditions.
For two independent groups, the Mann–Whitney/Wilcoxon rank-sum procedure is the relevant rank-based alternative considered here. For three or more independent groups, the analogous rank-based framework is the Kruskal–Wallis test (Adams & Lawrence, 2018; Moore et al., 2014, 2021).
Kruskal–Wallis is likewise not assumption-free. Moore et al. (2014, 2021) frame it as a rank-based comparison of multiple distributions and explicitly discuss its hypotheses and assumptions. Adams and Lawrence (2018) describe it for ordinal outcomes across three or more independent groups.
If the design were paired instead—for example, resolution times measured on matched service units before and after an intervention—the independent-groups Mann–Whitney procedure would no longer match the design. A paired rank procedure such as the Wilcoxon signed-rank test would address a different dependence structure and carries its own interpretive conditions (Moore et al., 2014, 2021).
The word nonparametric therefore never determines the test by itself. Design still does.
Mean difference and rank difference answer different questions
The practical distinction can be summarized this way:
| Analysis | Primary target in this case | What a significant result supports |
|---|---|---|
| Two-sample t procedure | Population means | Evidence that mean resolution times differ |
| Mann–Whitney/Wilcoxon rank sum | Population distributions/ranks | Evidence that one distribution tends to have systematically larger values |
| Mann–Whitney under a defensible same-shape location-shift model | Location/median difference | Evidence consistent with a difference in population location/medians |
| Kruskal–Wallis | Multiple independent population distributions | Evidence that the distributions are not all the same |
The researcher should choose among these because of the question being asked and the conditions supported by the data, not because one has the label “parametric” and another has the label “nonparametric” (Moore et al., 2014, 2021; Newton & Rudestam, 1999).
A better decision workflow
When a normality test is significant, the next step should be investigation rather than automatic test substitution.
- State the estimand. Are you trying to compare means, medians under a location model, ordered responses, or distributions more generally?
- Confirm the measurement level. Does the scale support quantitative differences, or only ordering?
- Confirm the design structure. Are observations independent, paired, repeated, clustered, or otherwise dependent?
- Plot the data. Examine shape, spread, asymmetry, ties, gaps, and unusual observations.
- Investigate outliers. Distinguish data errors from legitimate extreme cases.
- Assess the assumptions of the proposed parametric procedure. Do not reduce this assessment to one formal normality test.
- Ask how robust that procedure is under the observed departures. Sample size, balance, skewness, and outliers all matter.
- Understand the alternative before switching. Determine what the rank procedure actually tests and what additional assumptions are needed for the interpretation you intend to report.
- Use sensitivity analyses when useful. If defensible alternative analyses answer related questions, examine whether conclusions materially depend on the analytical choice.
- Report the limitation. Neither a parametric nor a nonparametric test repairs weak sampling, dependence, measurement problems, or an observational design that cannot support causal conclusions.
This approach treats statistical assumptions as part of research reasoning rather than as a sequence of pass/fail software outputs (Moore et al., 2021; Newton & Rudestam, 1999).
The decision in this synthetic case
The Shapiro–Wilk tests are clearly significant. The data are visibly right-skewed and contain unusual observations.
Yet those facts alone do not establish that Mann–Whitney is the correct primary analysis.
If the target is mean resolution time
The study has two independent groups, 60 observations per group, and a genuinely quantitative outcome. A two-sample mean comparison remains directly aligned with that research question.
The relatively large, balanced samples provide some protection against non-Normality, although the unusual observations still require investigation and transparent reporting (Moore et al., 2021).
If using Mann–Whitney
The rank-based analysis is useful as a complementary comparison, but it answers a different general question about the distributions.
Calling its result a test of medians without checking the additional distribution-shape condition would overstate what the procedure establishes (Moore et al., 2014, 2021).
In this particular synthetic dataset, neither analysis produces strong evidence of a group difference. More importantly, the two procedures reach that conclusion through different inferential targets.
The main lesson
A significant Shapiro–Wilk test is a diagnostic result, not a command to switch statistical tests.
The parametric-versus-nonparametric decision should follow from the research question, measurement scale, design structure, distributional behavior, unusual observations, assumptions and robustness of the candidate parametric procedure, and the hypothesis actually tested by the proposed nonparametric alternative (Adams & Lawrence, 2018; Moore et al., 2021; Newton & Rudestam, 1999).
Nonparametric methods can be extremely useful. They are not assumption-free, and they are not automatically superior whenever data depart from Normality.
The defensible question is not:
“Is Shapiro–Wilk significant?”
It is:
“Which analysis best answers the research question while remaining credible for the design and data we actually have?”
References
Adams, K. A., & Lawrence, E. K. (2018). Research methods, statistics, and applications (2nd ed.). SAGE Publications.
Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2
Moore, D. S., McCabe, G. P., & Craig, B. A. (2014). Introduction to the practice of statistics (8th ed.). W. H. Freeman.
Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). Macmillan Learning.
Newton, R. R., & Rudestam, K. E. (1999). Your statistical consultant: Answers to your data analysis questions. SAGE Publications.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.