Choosing the Right Statistical Test After Data Collection: From Research Question to Defensible Analysis
A synthetic case study showing how to choose a defensible statistical test after data collection by moving from the research question and target inference through variable type, dependence, study design, assumptions, and interpretation. The case shows why three independent groups with one quantitative outcome lead to one-way independent-groups ANOVA while an observational design limits causal inference.
Synthetic-case disclosure: This case is entirely illustrative. The researcher, study, participants, variables, and numerical values are synthetic and do not represent a real client or dataset. The values are included only to demonstrate a defensible statistical decision process.
Context
A graduate researcher has finished collecting data for a dissertation project and arrives at a common analysis-stage question:
“Which statistical test should I use?”
Her first instinct is to open the analysis menu in her statistical software and compare the available procedures: t tests, ANOVA, correlation, regression, chi-square, and several nonparametric tests.
That is the wrong starting point.
Selecting an analysis requires first identifying what the research question asks, how the variables were measured, how many groups or variables are involved, whether observations are independent or dependent, and what the study design permits the researcher to infer. Newton and Rudestam (1999), for example, organize statistical-test selection around the research question and the characteristics of the dependent and independent variables rather than around software commands. Adams and Lawrence (2018) likewise distinguish analyses according to research design, number and dependence of groups, and the variables being analyzed.
The synthetic study concerns graduate students' confidence in conducting quantitative thesis analysis after using one of three types of statistical-support provision:
- self-directed online materials;
- peer-led statistical support sessions; or
- individual statistical consultation.
Each participant used only one support format. Students chose the support available or preferred rather than being randomly assigned to it. At the end of the support period, each student completed a quantitative-analysis confidence scale scored from 0 to 100.
The dataset contains one record per participant and no repeated measurements.
The question is not yet “Which button should be clicked?”
It is: What comparison does the study actually support, and which analysis estimates that comparison under defensible assumptions?
Research decision
The researcher's original question was loosely stated:
“Does the type of statistical support affect students' quantitative-analysis confidence?”
That wording creates a problem. The study was observational: students were not randomly assigned to support formats. Random assignment is a central feature of a true experimental design because it helps separate the effect of an intervention from preexisting differences and other potential confounding variables (Frey, 2022). Moore et al. (2021) similarly emphasize that the method by which data are produced matters when results are interpreted and that a stronger case for causation comes from randomized experimentation.
Defensible research question:
“Do mean quantitative-analysis confidence scores differ among graduate students who used the three statistical-support formats?”
This changes the target inference.
The analysis will estimate and test differences in group means. It will not establish that selecting one support format caused confidence to increase. In a nonexperimental or correlational multiple-group design, ANOVA concerns the relationship between a grouping variable and a measured outcome rather than establishing the causal effect of an experimentally manipulated independent variable (Adams & Lawrence, 2018).
Interpretation boundary: The distinction between association and causation has to be settled before selecting the statistical test.
Data structure
Once the target inference is clear, the next task is to map the data structure.
| Decision feature | Synthetic study |
|---|---|
| Outcome | Quantitative-analysis confidence score |
| Outcome type | Quantitative |
| Outcome measurement | 0–100 composite score, analyzed as an interval-scale quantitative outcome |
| Predictor/grouping variable | Statistical-support format |
| Predictor type | Categorical |
| Predictor measurement | Nominal |
| Number of predictor levels | Three |
| Number of primary outcomes | One |
| Observations per participant | One |
| Relationship between groups | Independent |
| Study design | Cross-sectional observational group comparison |
| Primary target | Difference among group means |
| Causal inference justified? | No |
Measurement level matters because numbers do not automatically make a variable quantitative in the analytic sense. Nominal measurement classifies observations into categories without implying quantitative magnitude, whereas interval measurement permits meaningful numerical differences between values (Meier et al., 2014). Frey (2022) similarly distinguishes nominal classification, ordinal ranking, and interval measurement when describing levels of measurement.
Here, the three support formats are nominal categories. Coding them as 1, 2, and 3 in software would not make “3” three times “1,” nor would it create a meaningful linear distance between categories.
The outcome is different. For purposes of this synthetic case, the validated composite score is treated as a quantitative interval-scale measure, so differences in mean scores are meaningful for the intended analysis.
The observations are also independent rather than paired. A student appears in one support group only. There is no pre/post pair, no matched partner, and no sequence of repeated measurements from the same participant. Dependent-group procedures are intended for structures such as matched pairs or repeated measurements, whereas independent-group designs compare separate groups of participants (Adams & Lawrence, 2018).
These facts substantially narrow the candidate analyses before any software procedure is opened.
Candidate approaches
Several familiar tests might initially look plausible. They do not answer the same question.
Independent-samples t test
An independent-samples t test compares two independent groups on a quantitative outcome. In the present study, however, the predictor has three groups, not two. Adams and Lawrence (2018) identify two independent groups as the setting for an independent-samples t test and three or more independent groups as the setting in which one-way ANOVA becomes relevant.
Running several separate t tests would also replace one planned three-group question with multiple pairwise tests. The primary research question is whether the three population means differ, so an omnibus multiple-group procedure is the more direct starting point.
Decision: Reject as the primary analysis.
Paired-samples t test
A paired or dependent-samples t test requires dependent observations, such as repeated measurements on the same participants or appropriately matched observations. The current groups contain different students and each student contributes one outcome value (Adams & Lawrence, 2018).
There is therefore no meaningful pair structure.
Decision: Reject.
Repeated-measures ANOVA
Repeated-measures analysis is designed for dependent observations across conditions or measurement occasions. This dataset has one observation per student and three independent groups rather than repeated measurements on the same people (Adams & Lawrence, 2018).
Decision: Reject.
Chi-square test
A chi-square analysis would become relevant to a different question involving categorical frequencies or counts. Here, the primary outcome is a quantitative confidence score and the target is a comparison of means, not frequencies across categorical combinations. Meier et al. (2014) distinguish analyses for categorical data from procedures used for interval-level quantitative relationships and comparisons.
Decision: Reject for the stated outcome and estimand.
Correlation
A conventional correlation analysis is not the natural representation of this research question. The support-format variable is nominal, and the numbers used to code its three categories have no quantitative ordering or distance. Newton and Rudestam (1999) distinguish questions about group differences, typically involving discrete independent variables, from questions about the strength of relationships involving ordered or continuously distributed variables.
Treating support format coded 1–3 as if it were a continuous predictor would impose a numerical structure that the categories do not possess.
Decision: Reject.
Multiple regression with indicator variables
Regression is not inherently wrong here. A three-level categorical predictor can be represented using indicator variables, and a regression model can be formulated to compare group means. The question, however, contains one categorical grouping factor, one quantitative outcome, and no requested covariate adjustment.
For this specific target—an omnibus comparison of three independent group means—one-way ANOVA provides the most direct formulation.
Regression would become more attractive if the research question changed to require adjustment for additional predictors or explicitly parameterized group contrasts.
Decision: Defensible alternative, but unnecessary for the primary unadjusted question.
Kruskal–Wallis test
A nonparametric alternative might be considered if the measurement or distributional conditions required for the planned parametric analysis were not adequately supported. Adams and Lawrence (2018) describe the Kruskal–Wallis H test for ordinal data involving a variable with three or more independent levels.
But choosing a nonparametric procedure simply because it is perceived as “safer” would skip an important step: first determine whether the assumptions relevant to the intended mean comparison are materially problematic. Newton and Rudestam (1999) explicitly treat the parametric-versus-nonparametric decision as part of analysis selection rather than as an automatic preference.
Decision: Retain as a possible alternative if the data structure or diagnostics undermine the planned parametric analysis; do not choose it automatically.
Decision
The final primary analysis is a one-way independent-groups analysis of variance (ANOVA).
Selection logic:
One quantitative outcome + one nominal grouping variable + three independent groups + a research question about mean differences → one-way independent-groups ANOVA.
One-way ANOVA is designed to compare several population means when observations are classified by one categorical explanatory factor (Moore et al., 2014, 2021). Adams and Lawrence (2018) specifically describe one-way ANOVA as an inferential procedure for multiple-group designs and distinguish its interpretation in experiments from its interpretation in correlational or quasi-experimental studies.
For illustration only, suppose the synthetic dataset contains 144 students:
| Support format | n | Mean confidence | SD |
|---|---|---|---|
| Self-directed resources | 48 | 71.8 | 10.2 |
| Peer-led support | 48 | 76.9 | 9.6 |
| Individual consultation | 48 | 79.4 | 10.8 |
These values are synthetic and illustrative.
Using these summary values, the illustrative omnibus result is approximately:
Sample
144 students
Omnibus statistic
F(2, 141) = 6.91
p value
p = .001
Effect size
η² ≈ .089
The statistical conclusion would be that the data provide evidence against the null hypothesis that all three population means are equal.
What the omnibus result does not establish: The conclusion is narrower than “all groups differ from one another.” The omnibus test addresses whether the group means are all equal; determining which specific groups differ requires appropriate follow-up comparisons or prespecified contrasts.
Checks
Choosing ANOVA does not finish the analysis. It defines the model whose conditions now need to be examined.
Moore et al. (2021) formulate the one-way ANOVA model using independent errors that are normally distributed around the group means with a common standard deviation. Newton and Rudestam (1999) likewise identify homogeneity of variance as a central assumption for conventional t tests and one-way ANOVA.
Independence
Independence is primarily a design issue, not something that can be repaired by a normality test.
In the synthetic dataset, each student contributes one record and belongs to only one support group. There are no repeated observations or matched pairs.
That supports the intended independent-groups structure.
A further design audit would still ask whether participants are clustered—for example, students nested within the same supervisor, laboratory, course, or university. If meaningful dependence existed at another level, ordinary one-way ANOVA could underrepresent that structure and a different model might be required.
Outcome distributions and unusual observations
The analyst should inspect the outcome within groups rather than treating an assumption test as a substitute for examining the data. Distributional shape, unusual observations, and potential data errors can affect the suitability and interpretation of an analysis (Newton & Rudestam, 1999).
For this synthetic case, suppose histograms and diagnostic plots show no extreme departures from the intended model and no observations that appear to be obvious recording errors.
That is an illustrative diagnostic result, not an empirical finding.
Equality of variability
The illustrative group standard deviations are 10.2, 9.6, and 10.8, so the observed spreads are similar.
Moore et al. (2021) note that one-way ANOVA assumes a common population standard deviation and recommend examining group standard deviations when assessing this condition. Newton and Rudestam (1999) also discuss checking homogeneity of variance when applying one-way ANOVA.
Do not report “assumptions passed” mechanically. The analyst should evaluate whether departures are substantial enough to threaten the intended inference and, when necessary, consider a more appropriate analysis.
Design and data-quality checks
The analysis should also confirm that:
- each participant appears only once;
- group coding matches the actual support categories;
- the outcome has been scored according to its intended measurement procedure;
- missing values have not inadvertently been coded as valid scores;
- impossible or implausible values have been investigated; and
- the group sizes reported in the analysis match the cleaned analytic dataset.
Statistical procedure selection is therefore inseparable from understanding how the variables were constructed and how the observations entered the dataset.
Interpretation
Suppose the illustrative ANOVA produces F(2, 141) = 6.91, p = .001.
A defensible interpretation is:
Mean quantitative-analysis confidence was not equal across the three statistical-support groups in this synthetic sample.
The descriptive means suggest higher confidence among students who used individual consultation than among those using peer-led support or self-directed resources. But the omnibus ANOVA alone does not establish every pairwise difference.
The result also should not be reduced to the p value.
The illustrative η² of approximately .089 describes the proportion of observed outcome variation associated with the grouping factor in this ANOVA decomposition. Statistical significance and practical importance are not interchangeable; effect magnitude, uncertainty, research context, and the substantive consequences of the observed differences should also inform interpretation (Adams & Lawrence, 2018; Newton & Rudestam, 1999).
Most importantly, the result is associational.
Because students were not randomly assigned to support format, the finding does not show that individual consultation caused greater confidence. Students selecting individual consultation could differ from students selecting self-directed materials in prior statistical experience, dissertation stage, motivation, available funding, supervisor encouragement, severity of their analysis problem, or other unmeasured characteristics.
The statistical test cannot turn an observational design into an experiment.
What would have gone wrong
The case illustrates several ways an apparently reasonable software-first analysis could have become difficult to defend.
Multiple independent-samples t tests
Running three independent-samples t tests would have fragmented one three-group research question into several pairwise analyses instead of beginning with the intended omnibus comparison.
Paired t test
Running a paired t test would have imposed dependence that does not exist in the data.
Repeated-measures ANOVA
Choosing repeated-measures ANOVA because there are “three conditions” would have confused the number of conditions with the structure of the observations. Three independent groups and three repeated measurements are fundamentally different designs.
Chi-square
Running chi-square would have changed the target from a comparison of quantitative means to an analysis of categorical frequencies.
Correlation with nominal codes
Entering support format as 1, 2, and 3 in a correlation would have treated arbitrary nominal codes as if they represented ordered quantitative distances.
Automatic nonparametric selection
Automatically choosing a nonparametric test would have allowed a procedural rule to replace examination of the outcome, measurement, estimand, and model conditions.
Causal overinterpretation
Interpreting a significant ANOVA as proof that consultation causes higher confidence would have exceeded what the observational design can establish. Random assignment, not the choice of an ANOVA menu command, is central to the stronger causal interpretation available from a randomized experiment (Frey, 2022; Moore et al., 2021).
In each case, the error begins before the software calculation.
Limitations
This case is synthetic, so its numerical results are demonstrations rather than empirical evidence.
Even if the same pattern appeared in real data, the observational design would remain a major limitation. Self-selection into support formats creates plausible confounding, so between-group differences cannot automatically be attributed to the support itself.
A one-way ANOVA also answers a deliberately limited question. It compares group means without adjusting for other variables. If the substantive research question required adjustment for prior quantitative experience, degree stage, discipline, baseline confidence, or other defensible covariates, the analysis would need to be reformulated rather than simply adding variables after seeing the initial result.
The quality of the outcome measure also matters. Statistical analysis cannot compensate for a poorly defined or unreliable measure of the construct of interest.
Finally, inference depends on how participants entered the study. An analysis of a convenience or otherwise restricted sample does not automatically justify generalization to all graduate researchers. The sampling process and potential sources of selection bias must remain part of the interpretation (Moore et al., 2021).
Practical takeaway
“Which statistical test should I use?” is rarely the first statistical question that needs answering.
A defensible workflow is:
- Research question
- Target inference
- Outcome type
- Predictor type
- Measurement level
- Number of groups or variables
- Independence/dependence structure
- Study design
- Candidate methods
- Assumptions and diagnostics
- Final analysis
- Interpretation within design limits
Research question → target inference → outcome type → predictor type → measurement level → number of groups or variables → independence/dependence structure → study design → candidate methods → assumptions and diagnostics → final analysis → interpretation within design limits.
Only after those decisions are explicit should the researcher open the statistical software.
In this synthetic case, the answer was one-way independent-groups ANOVA—not because ANOVA happened to appear in a menu, but because the research question concerned one quantitative outcome across three independent categories of one grouping variable. The observational design then determined how far the resulting inference could go.
The statistical test is therefore the consequence of the research question and data structure, not the starting point of the analysis (Adams & Lawrence, 2018; Newton & Rudestam, 1999).
References
Adams, K. A., & Lawrence, E. K. (2018). Research methods, statistics, and applications (2nd ed.). SAGE Publications.
Frey, B. B. (Ed.). (2022). The SAGE encyclopedia of research design (2nd ed.). SAGE Publications.
Meier, K. J., Brudney, J. L., & Bohte, J. (2014). Applied statistics for public and nonprofit administration (9th ed.). Cengage Learning.
Moore, D. S., McCabe, G. P., & Craig, B. A. (2014). Introduction to the practice of statistics (8th ed.). W. H. Freeman.
Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). Macmillan Learning.
Newton, R. R., & Rudestam, K. E. (1999). Your statistical consultant: Answers to your data analysis questions. SAGE Publications.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.