How Many Participants Do We Need? A Power and Sample-Size Case Study Under Effect-Size Uncertainty
A synthetic power and sample-size case study showing how to justify a proposed sample size when the expected effect is uncertain. It demonstrates how meaningful effects, variability, power, attrition, feasibility, and sensitivity analysis combine into a defensible planning rationale.
“Why did you choose this sample size?”
The uncomfortable answer is often that the researcher entered an effect size, alpha level, and desired power into software and accepted the resulting number. The calculation may be numerically correct while the most consequential input—the assumed effect—remains difficult to defend.
This synthetic case study shows a different approach. Instead of pretending that one exact expected effect size is known, the study is planned around the primary research question, the analysis that will answer it, the smallest effect that matters, plausible variability, and sensitivity to uncertain assumptions. Power depends on the significance level, variability, sample size, and the alternative effect that the study is intended to detect, so these inputs need to be explicit rather than hidden inside a single calculation (Moore et al., 2014; Moore et al., 2021).
Synthetic-case disclosure: All study details, planning assumptions, effect sizes, sample sizes, power values, attrition rates, and feasibility limits below are generated for illustration. They do not describe a real study or empirical findings.
The planning problem
A research team is preparing a two-arm randomized study of an eight-week digital fatigue-management programme for employees experiencing persistent work-related fatigue. Participants will be allocated equally to the intervention or usual-support condition.
The primary outcome is a continuous fatigue score measured at the end of the intervention. Lower scores indicate less fatigue.
The team initially asks:
“How many participants do we need for 90% power?”
That question is incomplete. Power cannot be determined from the desired power level alone. For a two-sample comparison of means, planning requires an alternative difference that is important to detect, sample sizes or allocation, alpha, and an assumption about the outcome standard deviation (Moore et al., 2014). More generally, statistical power describes the probability of rejecting the null hypothesis when a specified alternative is true, and higher power corresponds to a lower probability of Type II error for that alternative (Moore et al., 2021).
The planning question therefore becomes: How large must this study be to have adequate probability of detecting a difference that would matter, under a defensible range of assumptions about outcome variability?
That formulation produces a rationale that can be explained to a supervisor, ethics committee, reviewer, or research team.
1. Start with the primary research question
The synthetic primary research question is:
At eight weeks, is mean fatigue different between employees assigned to the digital fatigue-management programme and employees assigned to usual support?
This matters because sample-size planning should correspond to the statistical analysis that addresses the research question. A two-sample problem can arise directly from a randomized comparative experiment in which participants are assigned to two treatments, and the corresponding comparison concerns the difference between the two population means (Moore et al., 2014).
The planned primary analysis is therefore a two-sided independent two-sample comparison of mean eight-week fatigue scores, with equal allocation between groups.
This intentionally simple design makes the connection between question, analysis, assumptions, and sample size transparent. If the eventual study instead used a more complex longitudinal, multilevel, latent-variable, or missing-data model, the power calculation would need to represent that model rather than rely automatically on the calculation shown here. Model-specific power approaches, including Monte Carlo methods, can incorporate parameter values and other features of more complex models (Wang & Wang, 2012).
Research question
Compare mean eight-week fatigue scores between the intervention and usual-support groups.
Primary analysis
Two-sided independent comparison of means with equal group allocation.
Meaningful effect
A four-point between-group difference is the synthetic decision threshold used in this case.
Variability
Outcome standard deviations from 8 to 12 points are examined as illustrative planning assumptions.
Sensitivity and feasibility
The required sample is evaluated across uncertain assumptions and compared with the fictional project's recruitment constraints.
2. State the hypotheses before calculating N
Let μI denote the population mean fatigue score under the intervention and μC the corresponding mean under usual support.
The primary hypotheses are:
H0: μI − μC = 0
HA: μI − μC ≠ 0
Explicitly connecting the research hypothesis and its null counterpart helps make clear what statistical claim is being tested (Meier et al., 2014).
A two-sided alternative is used in this synthetic case because the study is intended to detect a difference in either direction rather than prespecifying that an effect in the unexpected direction should be excluded from the primary test.
3. Do not substitute “expected effect” for “effect that matters”
The team cannot defend a statement such as:
“We expect Cohen's d = 0.40.”
There is no sufficiently strong basis in this synthetic scenario for treating one standardized effect as the true expected value.
What the team can specify is a decision threshold:
A four-point between-group difference on the fatigue outcome is the smallest effect that the fictional research team considers practically meaningful enough to influence a decision about the intervention.
The four-point threshold is a synthetic planning assumption, not an empirical estimate and not a universal benchmark.
This distinction matters. Statistical significance and substantive importance are different concepts, and sufficiently large samples can make very small effects statistically significant (Moore et al., 2014; Newton & Rudestam, 1999). For power planning, Moore et al. (2021) recommend specifying an alternative that matters scientifically or for decision making rather than merely choosing an arbitrary nonzero difference.
Newton and Rudestam (1999) similarly frame power and effect size in the context of the phenomenon under investigation, asking researchers to consider prior research, typical outcome values, and the smallest effect of practical significance.
The defensible input in this case is therefore not:
“The true effect will be d = 0.40.”
It is:
“We want to understand the sample size required to detect a four-point effect, while acknowledging uncertainty about how variable the outcome will be.”
4. Variability turns one meaningful difference into several plausible effect sizes
For a two-sample mean comparison, the same raw mean difference can represent different standardized effects depending on the outcome standard deviation. Power calculations for the two-sample t test therefore require both the important mean difference and an assumption about variability (Moore et al., 2014).
The synthetic team considers standard deviations between 8 and 12 points plausible for planning. These values are illustrative assumptions only; they are not presented as pilot estimates or findings from previous studies.
With a four-point meaningful difference, the corresponding standardized effects are:
| Synthetic SD assumption | Meaningful raw difference | Implied standardized effect, d |
|---|---|---|
| 8 | 4 points | 0.50 |
| 9 | 4 points | 0.44 |
| 10 | 4 points | 0.40 |
| 11 | 4 points | 0.36 |
| 12 | 4 points | 0.33 |
This is the central uncertainty in the case. The research team does not have to claim that one of these standardized effects is “the” expected effect. It can instead show what each defensible assumption implies.
5. Make alpha and desired power explicit
The synthetic primary analysis uses a two-sided significance level of α = .05.
Alpha is part of the power calculation because changing the significance level changes the evidence required to reject the null hypothesis and therefore changes power (Moore et al., 2014; Moore et al., 2021).
The team prefers 90% power for its main planning scenario but also calculates 80% power as a sensitivity benchmark. These values are planning choices for this synthetic study, not universal standards.
There is no sample-size rule that can replace the substantive decision about what error probabilities, effects, and recruitment burden are acceptable for the study being planned.
6. Sensitivity analysis: what happens when the assumed effect changes?
The table below was generated computationally for this case using a two-sided, equal-allocation, independent two-sample t-test power calculation with α = .05.
All values are synthetic planning outputs.
| Standardized effect d | Required N/group, 80% power | Total analyzable N, 80% power | Required N/group, 90% power | Total analyzable N, 90% power |
|---|---|---|---|---|
| 0.30 | 176 | 352 | 235 | 470 |
| 0.35 | 130 | 260 | 173 | 346 |
| 0.40 | 100 | 200 | 133 | 266 |
| 0.45 | 79 | 158 | 105 | 210 |
| 0.50 | 64 | 128 | 86 | 172 |
The calculation illustrates why an unexplained effect-size assumption can dominate a sample-size justification. Under otherwise identical assumptions, moving from d = 0.50 to d = 0.30 changes the 90%-power requirement from 172 to 470 analyzable participants.
That direction is expected: effects farther from the null are easier to distinguish, whereas greater variability makes differences harder to detect; increasing sample size increases the information available to distinguish an alternative from the null (Moore et al., 2014; Moore et al., 2021).
Sensitivity analysis is therefore not an optional decorative table. In this case, it exposes the consequence of the assumption that the team is least certain about.
7. Express the sensitivity analysis in the original outcome units
Because the team has defined practical relevance as a four-point difference, another useful sensitivity analysis holds that difference constant and varies the assumed standard deviation.
Synthetic sensitivity analysis for a four-point difference, α = .05, two-sided, 90% power:
| Assumed outcome SD | Implied d | Required N/group | Total analyzable N | Recruitment target with 15% loss |
|---|---|---|---|---|
| 8 | 0.50 | 86 | 172 | 203 |
| 9 | 0.44 | 108 | 216 | 255 |
| 10 | 0.40 | 133 | 266 | 313 |
| 11 | 0.36 | 160 | 320 | 377 |
| 12 | 0.33 | 191 | 382 | 450 |
This table gives the stakeholder a much better answer to “Why this sample size?” than an isolated output from power software.
The required sample size is not unstable because the software is unreliable. It changes because the underlying planning assumptions change.
8. What about attrition?
The sample size required by the statistical calculation is the number of analyzable participants required under its assumptions. If participants are expected to become unavailable for the primary outcome, recruitment may need to exceed that number.
Attrition is not only a loss of observations. An analysis restricted to study completers reduces sample size, and attrition can also introduce bias when loss is related to study condition or outcome processes (Newton & Rudestam, 1999).
For illustration, suppose the team chooses the scenario with 266 analyzable participants and wants to examine recruitment requirements under several possible loss rates.
The arithmetic recruitment target is:
required recruitment = analyzable N / expected retained proportion
All values below are synthetic.
| Planned analyzable N | Synthetic loss assumption | Recruitment target |
|---|---|---|
| 266 | 10% | 296 |
| 266 | 15% | 313 |
| 266 | 20% | 333 |
These calculations should not be interpreted as a claim that attrition will actually equal 10%, 15%, or 20%. They show how recruitment requirements change if those planning assumptions are used.
Inflating recruitment does not solve the methodological consequences of informative missingness. The missing-data process remains an analysis and interpretation issue even when extra participants were recruited initially (Newton & Rudestam, 1999).
9. Bring feasibility into the decision
Suppose the fictional project team estimates that approximately 320 participants can realistically be recruited within its time and resource constraints.
That feasibility constraint changes the discussion.
At the central synthetic assumption of a four-point meaningful difference and SD = 10 (d = 0.40), 90% power requires 266 analyzable participants. Allowing for 15% loss gives a recruitment target of approximately 313 participants, which is within the hypothetical feasibility limit.
But if variability is SD = 12, the same four-point effect corresponds to only d ≈ 0.33. Approximately 382 analyzable participants—and about 450 recruits under 15% loss—would then be required for 90% power.
The appropriate response is not to hide that inconvenient scenario. It is to report it.
10. What does the feasible sample actually buy us?
Consider several candidate analyzable sample sizes.
Synthetic achieved-power calculations for a four-point difference:
| Total analyzable N | Power if SD = 10 | Power if SD = 12 |
|---|---|---|
| 200 | 0.804 | 0.650 |
| 266 | 0.901 | 0.773 |
| 320 | 0.946 | 0.844 |
| 382 | 0.974 | 0.901 |
The table prevents a common planning mistake: describing a study simply as “90% powered” without saying 90% power for what effect under what variability assumption.
Power belongs to a specified alternative and set of design assumptions rather than to a sample size in isolation (Moore et al., 2021).
11. Why not choose the worst-case number automatically?
The largest sample in the sensitivity range is not automatically the correct answer.
A sample size of 382 analyzable participants would protect the 90%-power objective for the four-point effect under the synthetic SD = 12 scenario. But after allowing 15% loss, approximately 450 recruits would be required—well above the fictional project's feasible recruitment of about 320.
The decision therefore has to expose the trade-off rather than manufacture certainty.
This is consistent with the broader lesson that sample-size requirements depend on the analysis and characteristics of the planned study. Wang and Wang (2012), discussing SEM specifically, caution that no single sample-size rule applies to all situations and show how model-based and Monte Carlo approaches can be used when requirements depend on model characteristics. The same planning principle applies here: the calculation should represent the study actually being proposed.
12. When would this simple calculation stop being adequate?
This case uses an intentionally straightforward two-group comparison. The calculation would need reconsideration if the primary analysis changed materially—for example, if the research question became longitudinal or required a substantially more complex model.
For complex models, simulation can be used to specify hypothesized population parameter values, repeatedly generate data under those assumptions, fit the planned model, and examine the resulting power and precision. Wang and Wang (2012) demonstrate this model-based logic for SEM and emphasize that the necessary population values must themselves be specified from theoretical or empirical assumptions.
Simulation can represent a complicated design more faithfully, but it cannot make uncertain assumptions disappear.
13. The final rationale should answer the reviewer, not just the software
The fictional research team can now write a rationale such as this:
Sample-size rationale. The primary research question concerns the difference in mean eight-week fatigue scores between two equally allocated randomized groups. The primary analysis is therefore planned as a two-sided independent comparison of means with α = .05. Rather than assuming one precisely known expected standardized effect, planning was based on a synthetic minimum practically meaningful difference of four outcome points and sensitivity to uncertainty in the outcome standard deviation. Under an SD of 10 points, the meaningful difference corresponds to d = 0.40, for which 266 analyzable participants (133 per group) provide approximately 90% power. Allowing for a synthetic 15% loss to primary-outcome analysis gives a recruitment target of approximately 313 participants. Sensitivity analysis shows that the required analyzable sample ranges from 172 when SD = 8 (d = 0.50) to 382 when SD = 12 (d ≈ 0.33) for 90% power. The planned target of 266 analyzable participants is therefore conditional on variability being near the central planning assumption; it would not provide 90% power if variability were near the upper end of the plausible range. Because recruitment above approximately 320 participants is assumed infeasible in this synthetic case, the team would document this residual risk rather than claim that the study is adequately powered under every plausible scenario.
That is more defensible than:
“G*Power said we need 266 participants.”
The first statement identifies the question, analysis, meaningful effect, variability assumption, alpha, desired power, attrition allowance, sensitivity range, feasibility constraint, and residual uncertainty.
The second provides only a number.
Analyzable sample
266 participants
133 participants per group under the central planning scenario.
Desired power
Approximately 90%
For a four-point difference when SD = 10.
Recruitment target
Approximately 313 participants
After allowing for a synthetic 15% loss.
Residual uncertainty
If SD = 12, N = 266 provides about 77% power rather than 90% power.
14. A practical checklist for defending a proposed sample size
Before presenting a sample-size calculation to a reviewer, supervisor, ethics committee, or research team, the researcher should be able to answer:
- What is the primary research question?
- Which planned statistical analysis directly answers it?
- What are the null and alternative hypotheses, where hypothesis testing is relevant?
- What effect would be practically or scientifically meaningful to detect?
- Why is that effect meaningful?
- Which variability or nuisance-parameter assumptions enter the calculation?
- What alpha level is being used?
- What power is desired, and for which specified alternative?
- Which inputs are well supported and which remain uncertain?
- How much does required N change across plausible assumptions?
- How will loss to analysis affect recruitment requirements?
- Is the resulting recruitment target feasible?
- What assumption, if wrong, would most change the recommendation?
The objective is not to find a universally correct sample size. It is to make the chain of reasoning inspectable.
Conclusion
The hardest part of sample-size planning is often not solving the power equation. It is deciding what belongs in it.
Power for a comparison of means depends on the alternative difference, variability, significance level, and sample size, so an exact required N necessarily reflects assumptions about those quantities (Moore et al., 2014; Moore et al., 2021). Statistical significance should also not be confused with practical importance; the effect worth detecting has to be interpreted in the substantive context of the research question (Newton & Rudestam, 1999).
For this synthetic case, 266 analyzable participants and a recruitment target of approximately 313 under 15% loss form the preferred planning scenario because they provide approximately 90% power for a four-point meaningful difference when SD = 10. They are not presented as magical numbers. If SD were 12, the same study would have only about 77% power at N = 266, and approximately 382 analyzable participants would be required for 90% power.
The defensible conclusion is therefore conditional:
We chose this sample size because it matches the primary analysis, provides the desired sensitivity for an explicitly defined meaningful effect under the central variability assumption, remains feasible after allowance for expected loss, and has been tested against less favorable assumptions. We also state clearly which plausible scenarios it does not fully protect against.
That is the sample-size rationale a reviewer can evaluate.
References
Meier, K. J., Brudney, J. L., & Bohte, J. (2014). Applied statistics for public and nonprofit administration (9th ed.). Cengage Learning.
Moore, D. S., McCabe, G. P., & Craig, B. A. (2014). Introduction to the practice of statistics (8th ed.). W. H. Freeman.
Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). Macmillan Learning.
Newton, R. R., & Rudestam, K. E. (1999). Your statistical consultant: Answers to your data analysis questions. SAGE Publications.
Wang, J., & Wang, X. (2012). Structural equation modeling: Applications using Mplus. John Wiley & Sons.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.