Resource

How to Determine Sample Size for a Clinical or Health Research Study

A practical framework for determining sample size in clinical and health research by defining the study objective, primary endpoint, meaningful effect or precision target, statistical assumptions, design, allocation, and expected information loss before calculating the required sample.

Researchers often ask, “How many participants do I need?” before the statistical problem has been defined well enough to calculate a defensible answer.

Sample size is not a standalone property of a study. It depends on what the study is designed to estimate or test, the primary endpoint, the effect or precision that matters, assumptions about variability or event occurrence, the study design, the allocation of participants, and the inferential error criteria. Piantadosi explicitly notes that there is no universal approach to sample size and power; the calculation must reflect the outcome, required precision, accrual process, and trial design (Piantadosi, 2005).

Core principle: Do not start with “What sample size do I need?” Start with “What exactly must this study be able to estimate or detect?”

This resource provides a researcher-facing framework for answering that question before requesting a formal sample-size calculation.

Sample size planning starts with the research objective

A sample-size calculation is the quantitative consequence of a study design, not the first design decision.

Before calculating a number, define the primary research question and how it will be answered statistically. For a comparative study, this includes identifying the primary endpoint, the treatment or exposure contrast, and the difference that would be scientifically or clinically important. Altman describes hypothesis-testing sample-size planning in terms of having a high probability of detecting a worthwhile effect if it exists and emphasizes specifying the smallest true difference that would be clinically valuable (Altman, 1991).

Piantadosi similarly lists the smallest treatment effect of interest based on clinical considerations among the quantitative design parameters relevant to sample-size planning (Piantadosi, 2005).

This is why a request such as “calculate my sample size for a clinical study” is incomplete. The same nominal study can require different sample sizes when its endpoint, meaningful effect, variability, hypothesis, allocation, or required precision changes.

First decide: are you testing a hypothesis or estimating with a specified precision?

Two studies can use the same type of endpoint yet require different approaches to sample-size determination because their statistical goals differ.

Sample size for hypothesis testing

Power-based planning asks whether the study will have an adequate probability of rejecting a null hypothesis when a specified alternative effect is true.

Power is 1 − β, where β is the probability of a Type II error under the specified alternative. Alpha (α) controls the Type I error criterion. Sample-size determination commonly coordinates the desired power with the significance level and the effect the study is intended to detect (Chow et al., 2018; Rosner, 2016).

Key requirement: Power is not meaningful without identifying the effect against which the study is being powered.

Sample size for estimation and precision

Not every study should be framed primarily around rejection of a null hypothesis.

An estimation-focused study may instead ask: How precisely must the parameter be estimated?

Sample size can be chosen so that a confidence interval has an acceptable width or maximum error. Chow et al. describe precision-based sample-size determination in terms of the maximum acceptable half-width of a confidence interval, while Piantadosi describes confidence-interval planning as choosing a sample size that makes the expected interval suitably narrow (Chow et al., 2018; Piantadosi, 2005).

Important distinction: Sample size for hypothesis testing and sample size for estimation are not interchangeable planning questions. One emphasizes the probability of detecting a prespecified effect; the other emphasizes uncertainty around an estimate.

The core inputs that determine sample size

Although the exact calculation changes with the endpoint and design, several quantities repeatedly determine how much information a study needs.

Alpha: what Type I error criterion will the study use?

Alpha is the prespecified Type I error probability used in hypothesis testing. Changing alpha changes the statistical evidence required for rejection and therefore affects the sample size needed for a given power and alternative effect (Chow et al., 2018; Rosner, 2016).

Alpha should therefore be specified as part of the study's inferential plan rather than selected after the desired sample size is known.

The calculation also needs to reflect whether the relevant hypothesis is one-sided or two-sided. Chow et al. show explicitly that sample sizes differ between one-sided and two-sided testing at the same nominal alpha (Chow et al., 2018).

Power: what probability of detecting the target effect is required?

Power is the probability that the specified statistical test rejects the null hypothesis when the alternative effect used in the calculation is true.

Increasing desired power, with the other planning assumptions held fixed, requires more information and therefore generally a larger sample. Rosner describes power calculations as a study-planning tool and links them to the assumed effect and projected variability (Rosner, 2016).

Power should not be discussed independently of the effect being targeted. A study can have very different power for small and large alternatives.

Effect size: what difference actually matters?

One of the most consequential sample-size inputs is the treatment difference, association, or other effect the study is intended to distinguish.

For clinical research, the most defensible starting point is usually the effect that would matter scientifically or clinically, not an arbitrary standardized effect selected merely because software requests one.

Altman frames the problem as specifying the smallest treatment difference that would be clinically worthwhile. For continuous outcomes, he then standardizes that meaningful difference relative to the outcome standard deviation for calculation (Altman, 1991).

Planning relationship: meaningful difference in the original outcome scale + expected variability → standardized difference

This distinction matters. A standardized effect is often a derived planning quantity rather than a substitute for defining what magnitude of effect matters.

Altman notes that specifying an effect directly in standard-deviation units can be used when variability is unavailable, but he also characterizes such solutions as involving subjectivity (Altman, 1991).

For a researcher, therefore, “use a medium effect size” is generally a weaker justification than explaining what change in the actual endpoint would be important and what variability is reasonably expected.

Variability: how noisy is the endpoint?

For continuous outcomes, the assumed standard deviation is a central sample-size input. The same clinically meaningful mean difference becomes harder to detect when measurements are more variable.

Altman identifies the standard deviation, clinically relevant difference, significance level, and power as required inputs for a basic two-independent-group continuous-outcome calculation (Altman, 1991).

The relevant variability also depends on design. In paired or within-person studies, for example, Altman notes that the standard deviation of the within-person changes is the relevant quantity rather than simply the marginal standard deviation of the outcome (Altman, 1991).

Assumptions about variability should therefore match the proposed analysis and data structure.

Proportions and event rates: what outcome frequency is expected?

Binary endpoints require assumptions about outcome probabilities rather than only a continuous-outcome standard deviation.

For a comparison of proportions, planning requires assumptions about the expected outcome proportion in the relevant groups and the difference considered important, together with alpha and power (Altman, 1991).

Chow et al. likewise provide separate sample-size procedures for means, proportions, rates, and survival endpoints, reflecting the fact that different endpoints have different sampling structures and nuisance parameters (Chow et al., 2018).

For rare outcomes, seemingly small changes in assumed incidence can materially alter the amount of information available and hence the required sample.

Precision: how narrow must the confidence interval be?

When estimation is the primary goal, the researcher needs to define the degree of uncertainty that is acceptable.

That may mean specifying the maximum acceptable margin of error or confidence-interval width around the parameter of interest. Narrower confidence intervals represent greater precision and generally require more information (Chow et al., 2018; Piantadosi, 2005).

A request for “a precise estimate” is therefore insufficient. The statistician needs to know how precise the estimate needs to be.

Allocation: how will participants be distributed between groups?

Allocation affects efficiency.

Equal allocation is often the reference design for simple two-group comparisons, but researchers may prefer unequal allocation because of costs, recruitment, ethics, exposure to an experimental intervention, or other study considerations.

Unequal allocation changes the information contributed by the groups and therefore changes the total sample-size requirement. Both Altman and Piantadosi show that allocation ratio enters sample-size or power calculations; departures from equal allocation can reduce statistical efficiency, although modest imbalance may have relatively small effects in some settings (Altman, 1991; Piantadosi, 2005).

The planned allocation ratio should therefore be specified before the final calculation.

Endpoint type determines the sample-size method

There is no single “medical research sample-size formula” because different endpoints generate different statistical information.

Continuous endpoints

Examples include laboratory measurements, physiological measurements, symptom scores, or other quantitative outcomes.

Planning commonly requires the meaningful difference and an assumption about variability. The relevant variance may differ depending on whether observations come from independent groups, paired measurements, or another design (Altman, 1991).

Binary endpoints

For outcomes such as response/nonresponse or event/no event, expected proportions are central inputs.

The assumed baseline probability matters because binomial sampling variability depends on the underlying event proportion (Altman, 1991; Chow et al., 2018).

Time-to-event endpoints

For survival or other time-to-event outcomes, participant count alone does not characterize the information available.

Piantadosi includes the number of events, censoring, accrual, loss to follow-up, and follow-up duration among the relevant planning parameters for clinical trials. Event-time calculations can depend explicitly on the required number of events and assumptions about accrual and censoring (Piantadosi, 2005).

A study may therefore recruit many participants yet still generate insufficient event information if the endpoint is uncommon or follow-up is inadequate.

Paired, crossover, clustered, and other structured designs

Study design can change the effective information in the data.

Paired and crossover designs involve within-person variability. Cluster-randomized designs lose information relative to the same number of independent observations because outcomes within clusters are correlated. Piantadosi explicitly describes clustering as reducing effective sample size and increasing variance (Piantadosi, 2005).

Methodological requirement: The calculation must correspond to the actual study design—not a simpler design that happens to be available in a software menu.

Superiority, noninferiority, and equivalence require different planning

A critical clinical-trial decision is what conclusion the study is designed to support.

How the inferential objective changes sample-size planning
Objective Question Planning implication
Superiority Is one treatment better than its comparator according to the prespecified endpoint and hypothesis? The effect to be detected, alpha, power, variability or event assumptions, allocation, and design all enter the planning problem.
Noninferiority Can the new treatment exclude an unacceptable degree of inferiority relative to the comparator? A prespecified noninferiority margin is required. For a continuous-outcome comparison, sample size depends on alpha, power, variance, allocation, the expected treatment difference, and the noninferiority margin (Chow et al., 2018).
Equivalence Can clinically important differences be excluded in both directions? Equality, superiority/noninferiority, and equivalence use different hypotheses and sample-size requirements. Equivalence can be formulated through two one-sided tests because both unacceptable boundaries must be excluded (Chow et al., 2018).

For noninferiority, a tighter margin requires the study to distinguish the anticipated effect from a closer boundary and therefore demands greater precision.

Do not infer equivalence from nonsignificance: A failed superiority test is not evidence that two treatments are equivalent or that one is noninferior. Piantadosi warns that inadequate precision can make treatments appear “not significantly different” without establishing meaningful equivalence (Piantadosi, 2005).

One-sided versus two-sided testing must follow the research question

Whether a calculation uses a one-sided or two-sided test affects the required sample size because the rejection criterion changes.

Chow et al. explicitly show that one-sided and two-sided calculations produce different sample sizes at a fixed alpha and power (Chow et al., 2018).

This is not a reason to choose a one-sided test merely because it produces a smaller sample.

The directionality must follow the hypothesis being investigated. Chow et al. note that a one-sided procedure is appropriate only in settings where concern genuinely lies in one tail and an outcome in the opposite direction is not relevant to the inferential question; their broader discussion also documents why two-sided testing may be preferred in many clinical settings (Chow et al., 2018).

The choice should be made during study design, not after examining results or recruitment feasibility.

Allow for dropout and other loss of information

The number produced by a core statistical calculation is not always the number of participants who must be recruited.

Participants may withdraw, become unavailable for outcome measurement, cross over between treatments, fail to adhere, or be censored before an event occurs. These processes can reduce the amount of usable information.

Piantadosi includes loss to follow-up and censoring explicitly among trial-planning parameters and shows that event-time planning can incorporate loss-to-follow-up assumptions. He also discusses how nonadherence can diminish treatment-group differences and increase the sample-size requirement (Piantadosi, 2005).

Distinguish two quantities:

the amount of information needed for the primary analysis → the number of participants who must initially be recruited to obtain that information

A dropout allowance should not be an unexplained percentage automatically added to every study. It should reflect the study's follow-up structure, endpoint, primary analysis, and plausible mechanism of information loss.

Why sensitivity analysis belongs in sample-size planning

Sample-size calculations rely on assumptions about quantities that are often uncertain before the study begins.

For example, the true standard deviation may be unknown. Rosner notes that power calculations often rely on projected variability, sometimes informed by preliminary data (Rosner, 2016). Chow et al. additionally warn that estimates obtained from pilot studies contain sampling error and can consequently produce unstable sample-size estimates when treated as if they were the true population parameters (Chow et al., 2018).

A defensible plan should therefore examine how the required sample changes across plausible assumptions when uncertainty is material.

Instead of reporting only:

“The required sample size is N.”

a stronger justification explains:

“Under these assumptions about the endpoint, meaningful effect, variability or event frequency, alpha, power, allocation, and information loss, the required sample is N; alternative plausible assumptions produce the following planning implications.”

That makes the assumptions visible rather than allowing a single software output to imply false certainty.

Checklist: what to define before asking for a sample-size calculation

Before sending a request to a statistician, try to answer the following questions.

Research question and objective

  • What is the primary research question?
  • Is the study primarily estimating a parameter or testing a hypothesis?
  • If comparing groups, what specific contrast is primary?
  • Is the objective superiority, noninferiority, equivalence, or another inferential goal?

Primary endpoint

  • What is the single primary endpoint for the calculation?
  • Is it continuous, binary, count/rate, ordinal, or time-to-event?
  • At what time point or follow-up period is it evaluated?
  • Is the endpoint measured once or repeatedly?
  • Is the analysis based on independent participants, paired observations, clusters, or another structure?

Effect or precision target

  • What difference or effect would be clinically or scientifically important?
  • Can that effect be expressed in the endpoint's original units?
  • If the goal is estimation, what confidence-interval width or margin of error is acceptable?
  • For noninferiority or equivalence, what clinically justified margin will define the conclusion?

Nuisance parameters and variability

  • For a continuous endpoint, what standard deviation or other variance quantity is expected?
  • For a binary endpoint, what outcome proportions are expected?
  • For a time-to-event endpoint, what event occurrence, censoring, accrual, and follow-up assumptions are plausible?
  • What evidence supports these assumptions?
  • Are the inputs uncertain enough to require sensitivity analyses?

Inferential criteria

  • What alpha or significance level is intended?
  • What power is required for the specified effect?
  • Is the hypothesis one-sided or two-sided, and why?
  • Are multiple primary comparisons or endpoints relevant to the error-rate strategy?

Study design and allocation

  • How many study groups are there?
  • What allocation ratio is planned?
  • Is the design parallel, paired, crossover, clustered, or otherwise structured?
  • If clustered, what information is available about cluster sizes and within-cluster similarity?
  • Does the proposed calculation match the planned primary analysis?

Recruitment and information loss

  • What dropout or loss-to-follow-up is reasonably expected?
  • Can participants be censored before the endpoint is observed?
  • Is nonadherence or treatment crossover plausible?
  • How many analyzable participants or events are required?
  • How many participants may need to be recruited to achieve that amount of information?

What your biostatistician needs from you

You do not need to arrive with a finished statistical formula. You do need to arrive with the scientific decisions that the formula is supposed to represent.

A useful sample-size request might provide:

Study objective
Compare the primary outcome between the specified study groups.
Design
State whether groups are independent, paired, randomized, clustered, crossover, or another design.
Primary endpoint
Name the outcome, its type, and when it is measured.
Primary effect
State the difference, ratio, margin, or other effect that would matter scientifically or clinically.
Variability/event assumptions
Provide plausible standard deviations, proportions, event rates, or related quantities and their sources.
Inferential objective
State whether the study concerns superiority, noninferiority, equivalence, or precision of estimation.
Alpha and sidedness
State the planned significance criterion and whether the hypothesis is one- or two-sided.
Power or precision
State the desired power for the target alternative or the required confidence-interval precision.
Allocation
State the intended group ratio.
Follow-up
Describe study duration, timing of measurements, and relevant censoring or event processes.
Expected information loss
Describe plausible dropout, missing primary outcomes, nonadherence, or loss to follow-up.
Uncertainty
Identify assumptions that should be varied in sensitivity analyses.

If several of these items are unknown, that does not mean a sample-size calculation is impossible. It means the next task is study-design clarification, not pressing “calculate” in power software.

Common sample-size planning mistakes

Asking for N before defining the endpoint

Different endpoints can require different assumptions and different calculations. “We are measuring several health outcomes” is not enough to identify the primary sample-size problem.

Choosing an effect size because software needs one

The effect should correspond to the scientific objective. For clinical research, start where possible with a difference that would matter in the endpoint's natural scale and then incorporate the required variability assumptions (Altman, 1991).

Treating a pilot estimate as the true population value

Pilot estimates can be uncertain. Chow et al. specifically note that sampling error in pilot-based parameter estimates can make the resulting sample-size estimate unstable (Chow et al., 2018).

Ignoring the study design

A formula for two independent groups does not automatically apply to paired, crossover, clustered, or time-to-event studies.

Confusing “not significant” with equivalence

A study that fails to demonstrate superiority has not automatically established equivalence or noninferiority. Those objectives require different hypotheses and planning (Chow et al., 2018; Piantadosi, 2005).

Planning only for enrolled participants

The statistical requirement may concern analyzable observations or observed events. Recruitment must account for realistic loss of information where appropriate (Piantadosi, 2005).

Changing assumptions until the sample becomes affordable

Altman cautions that it is easy to alter the calculated sample by changing the input assumptions but argues that study requirements should preferably be decided in advance. If the scientifically required study cannot feasibly be conducted, simply weakening the planning assumptions is not a statistical solution (Altman, 1991).

A practical workflow for sample size calculation for research

A defensible planning sequence is:

  1. Research question
  2. Study design
  3. Primary endpoint
  4. Target effect or precision
  5. Endpoint-specific assumptions
  6. Alpha and power
  7. Sidedness and hypothesis
  8. Allocation
  9. Sample-size calculation
  10. Information-loss allowance
  11. Sensitivity analysis
  12. Feasibility review

Research question → study design → primary endpoint → target effect or precision → endpoint-specific assumptions → alpha and power → sidedness/hypothesis → allocation → sample-size calculation → information-loss allowance → sensitivity analysis → feasibility review

This sequence prevents the sample-size number from driving the scientific assumptions that were supposed to determine it.

The final sample-size statement in a protocol should explain not only how many participants are required, but why that number follows from the study's objective and assumptions.

Frequently Asked Questions

How do I calculate sample size for a research study?

First define the research objective, primary endpoint, study design, effect or precision target, relevant variability or event assumptions, alpha, power, sidedness, allocation, and expected information loss. The appropriate calculation then follows from that specification. There is no single universal sample-size calculation applicable to all clinical and health studies (Piantadosi, 2005).

What factors affect sample size in clinical research?

Important factors include the endpoint type, effect of interest, variability or event probability, alpha, desired power, precision requirement, study design, allocation ratio, censoring or follow-up structure, and loss of information. Which factors are relevant depends on the specific design and endpoint (Altman, 1991; Chow et al., 2018; Piantadosi, 2005).

Does a larger effect size require a smaller sample?

For otherwise fixed assumptions in conventional power calculations, effects farther from the null are easier to detect and therefore require less information than smaller effects. But the target effect should be scientifically justified rather than enlarged merely to reduce the required sample (Altman, 1991).

Should I use a standardized effect size?

A standardized effect can be useful computationally, particularly for continuous outcomes, but it should not replace substantive thinking about the effect that matters. Altman constructs the standardized difference from the clinically relevant raw difference and standard deviation. Specifying a difference directly in standard-deviation units when variability is unavailable is possible, but it introduces additional subjectivity (Altman, 1991).

What is the difference between alpha and power?

Alpha concerns the probability of a Type I error under the null hypothesis. Power, 1 − β, is the probability of rejecting the null when the specified alternative is true. Both affect sample-size requirements (Chow et al., 2018; Rosner, 2016).

Is sample size always based on statistical power?

No. Sample size can also be chosen to achieve a desired precision of estimation, such as an acceptable confidence-interval width or margin of error. Power-based and precision-based planning answer different statistical questions (Chow et al., 2018; Piantadosi, 2005).

Does a survival study need a sample size or a number of events?

Time-to-event planning often depends critically on the number of observed events, together with assumptions about accrual, event occurrence, censoring, loss to follow-up, and study duration. Participant count alone may therefore be an incomplete description of the required information (Piantadosi, 2005).

Should dropout be added to the calculated sample size?

Expected loss of primary-outcome information should be considered when translating the required statistical information into a recruitment target. The appropriate allowance depends on the study and should reflect realistic follow-up, censoring, nonadherence, or dropout assumptions rather than an unexplained universal percentage (Piantadosi, 2005).

Can I calculate sample size before deciding whether the study is superiority or noninferiority?

Not defensibly. Superiority, noninferiority, and equivalence involve different hypotheses and can require different inputs and sample-size calculations. A noninferiority study additionally requires a prespecified clinically justified margin (Chow et al., 2018).

The key principle

The most important sample-size input is not the software package or formula.

It is the scientific specification of what the study must learn.

A defensible clinical study sample size follows from a defined endpoint, an explicit effect or precision target, appropriate assumptions about variability or event occurrence, a prespecified hypothesis and error structure, the actual study design and allocation, and realistic consideration of information loss.

Only after those decisions have been made does “How many participants do we need?” become a statistical question that has a meaningful answer.

References

Altman, D. G. (1991). Practical statistics for medical research. Chapman & Hall.

Chow, S.-C., Shao, J., Wang, H., & Lokhnygina, Y. (2018). Sample size calculations in clinical research (3rd ed.). Chapman & Hall/CRC.

Piantadosi, S. (2005). Clinical trials: A methodologic perspective (2nd ed.). John Wiley & Sons.

Rosner, B. (2016). Fundamentals of biostatistics (8th ed.). Cengage Learning.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry