Is 200 Participants Enough for SEM? A Model-Specific Decision Case Study
There is no universal SEM minimum that makes 200 participants automatically adequate or inadequate. This synthetic case shows how latent variables, indicators, model identification, measurement quality, structural effects, estimation, missingness, and model-specific power change the sample-size decision.
A researcher is planning a structural equation model and has been given a simple answer to a difficult question:
“SEM requires at least 200 participants.”
The proposed study can recruit approximately 200 participants, so the researcher initially treats the sample-size problem as solved.
It is not solved.
SEM sample-size requirements cannot be separated from the model being estimated. Wang and Wang (2012) explicitly note that there is no absolute sample-size standard that applies to every SEM situation. Required sample size can depend on factors including the number of parameters, indicators per latent variable, indicator reliability, study design, distributional characteristics, missing-data handling, model complexity, and estimator. Tabachnick and Fidell (2013) similarly emphasize that SEM is sensitive to sample size while noting that stronger expected parameters and more reliable measures can reduce the demands placed on the sample.
The planning question should not be: “Is 200 the minimum for SEM?”
It should be: “Is a target of 200 participants defensible for this particular measurement model, structural model, estimator, data quality, and target effects?”
Synthetic case disclosure: The numerical specifications below are illustrative planning assumptions, not empirical findings or universal SEM benchmarks.
The proposed study
A fictional organizational researcher wants to investigate whether employees who perceive stronger Organizational Support report greater Psychological Safety, whether these constructs are associated with Work Engagement, and whether engagement is subsequently associated with Intention to Remain.
The researcher proposes four latent variables:
| Latent variable | Indicators | Measurement |
|---|---|---|
| Organizational Support | SUP1, SUP2, SUP3 | Three reflective survey indicators |
| Psychological Safety | SAF1, SAF2, SAF3 | Three reflective survey indicators |
| Work Engagement | ENG1, ENG2, ENG3 | Three reflective survey indicators |
| Intention to Remain | RET1, RET2, RET3 | Three reflective survey indicators |
The model therefore contains four latent constructs and 12 observed indicators.
The structural hypotheses are:
- Organizational Support → Psychological Safety
- Organizational Support → Work Engagement
- Psychological Safety → Work Engagement
- Psychological Safety → Intention to Remain
- Work Engagement → Intention to Remain
This is already more informative for sample-size planning than the statement “we are using SEM.” SEM can represent hypothesized relationships among observed and latent variables, with the measurement and structural portions of the model contributing different parameters and different sources of uncertainty (Lovric, 2011; Wang & Wang, 2012).
First decision: What exactly is being estimated?
The researcher is not estimating five regressions between perfectly measured variables.
Each latent variable must first be represented through its indicators. A confirmatory factor model specifies relationships between latent factors and observed indicators while explicitly representing measurement error (Lovric, 2011; Wang & Wang, 2012).
For this synthetic case, the intended measurement model has:
- four latent factors;
- three indicators per factor;
- each indicator loading on its intended factor only;
- no planned cross-loadings;
- no planned correlated indicator errors;
- correlations among the four factors when the measurement model is evaluated as a CFA.
The structural model then replaces unrestricted correlations among the latent variables with the five hypothesized directional paths.
That distinction matters because the credibility of the structural coefficients depends partly on whether the proposed indicators provide an adequate measurement model. A nominal sample size does not compensate for a poorly specified measurement structure.
Model identification comes before sample-size adequacy
Before asking whether 200 observations provide enough statistical information, the researcher needs to establish whether the model is identified.
Identification concerns whether the observed data contain enough information, under the model's constraints, to obtain unique estimates of the unknown parameters. An underidentified model cannot be rescued by collecting a larger sample; Wang and Wang (2012) explicitly distinguish identification from sample size and note that an underidentified model remains underidentified regardless of how large the sample becomes.
A necessary identification check compares the number of distinct observed variances and covariances with the number of free parameters. For p observed variables, the covariance matrix supplies:
p(p + 1) / 2
distinct variance-covariance elements. This counting condition is necessary but is not, by itself, sufficient to establish identification (Lovric, 2011; Wang & Wang, 2012).
With 12 observed indicators, this synthetic model supplies:
12(13) / 2 = 78 distinct covariance elements.
Suppose one loading per factor is fixed to establish each latent variable's scale, as described by Wang and Wang (2012). Under the deliberately simple structural specification above, the model would estimate:
| Parameter type | Free parameters |
|---|---|
| Remaining factor loadings | 8 |
| Indicator residual variances | 12 |
| Organizational Support variance | 1 |
| Residual variances of the three endogenous latent variables | 3 |
| Structural paths | 5 |
| Total | 29 |
Observed covariance elements
78
From 12 observed indicators.
Free parameters
29
Under the deliberately simple parameterization described above.
Degrees of freedom
49
78 − 29 = 49.
This counting exercise supports a necessary identification condition and shows that the proposed model is overidentified by the parameter-count criterion. It does not prove that every possible parameterization is identified. Identification also depends on the actual pattern of fixed, free, and constrained parameters, and software checks do not replace thoughtful specification (Lovric, 2011; Wang & Wang, 2012).
“N = 200” and “the model is identified” answer different questions.
The measurement model changes the sample-size problem
The proposed constructs each have three indicators. Wang and Wang (2012) discuss the number of indicators per factor as one factor affecting both identification and sample-size considerations. They also emphasize indicator reliability as a relevant determinant of sample-size needs.
Consider two possible versions of the same 12-indicator model.
Scenario A: Stronger measurement
The indicators are strong and behave cleanly.
As a synthetic planning assumption, standardized loadings might mostly be around .70–.80, with no important cross-loadings or residual dependencies.
Scenario B: Weaker measurement
Several indicators are weaker.
Some hypothetical loadings might be closer to .50–.60, and two constructs may be difficult to distinguish empirically.
Both scenarios have:
- 200 participants;
- four latent variables;
- 12 indicators;
- five structural paths;
- the same nominal number of free parameters.
Yet they need not have the same statistical behavior.
Tabachnick and Fidell (2013) specifically connect sample-size adequacy with the strength of expected parameters and reliability of the measured variables. Wang and Wang (2012) likewise identify indicator reliability as one of the factors affecting SEM sample-size requirements.
A universal-N rule cannot represent that difference.
Why “observations per parameter” is not enough
The synthetic structural model has 29 free parameters under the parameterization above. With 200 observations, a researcher could calculate:
200 / 29 ≈ 6.9 observations per free parameter.
That number may look reassuring or alarming depending on which rule of thumb the researcher has heard.
But the ratio does not answer the substantive planning question.
Wang and Wang (2012) review several historical rules based on cases per observed variable, cases per free parameter, and fixed minimum sample sizes. They then caution against simply trusting such rules because required sample size also depends on measurement reliability, design, nonnormality, missing-data handling, model complexity, and estimation method.
Tabachnick and Fidell (2013) provide an instructive applied example: a CFA with 177 participants had approximately eight cases per estimated parameter, and the authors judged that ratio in conjunction with the high reliability of the indicators rather than treating the ratio as a universal criterion.
The implication is not that 6.9 observations per parameter is adequate.
Nor is it that 6.9 is inadequate.
The ratio alone does not establish either conclusion.
Estimation also matters
The researcher initially plans covariance-based SEM with continuous indicator scores and maximum-likelihood estimation.
Maximum likelihood is a commonly used SEM estimator. Its large-sample properties and inferential procedures depend on assumptions that include distributional considerations; substantial nonnormality can particularly affect standard errors and model chi-square inference, motivating consideration of transformations, bootstrap procedures, or alternative robust estimators where appropriate (Wang & Wang, 2012).
Tabachnick and Fidell (2013) likewise emphasize that sample size and the plausibility of distributional and independence assumptions should be considered when selecting an SEM estimator.
The researcher therefore cannot justify 200 participants without knowing more about the intended data.
For example:
- Are the 12 indicators being treated as continuous?
- Are they strongly skewed or highly kurtotic?
- Are they actually ordered categorical responses requiring a different estimator?
- Are observations independent?
- Will a robust estimator be needed?
- Are there sparse response categories?
Wang and Wang (2012) explicitly identify the choice among estimators such as ML, robust ML, and categorical-data estimators as one factor relevant to SEM sample-size requirements.
“SEM with 200 people” is therefore still an incomplete description of the analysis.
What about the five structural paths?
Suppose the researcher's primary scientific target is the path:
Psychological Safety → Work Engagement.
For planning purposes, the researcher believes the standardized effect could plausibly be somewhere between .15 and .35.
Those values are synthetic assumptions. They are not benchmarks from the literature and are not claims about organizational research.
They illustrate a crucial point: power depends on what effect the study needs to detect.
Statistical power is the probability of rejecting a false null hypothesis under specified conditions, and power depends on the alternative being considered rather than sample size alone (Wang & Wang, 2012). Model-specific SEM power methods can evaluate power for particular parameters under hypothesized population parameter values, and Monte Carlo approaches can be used to investigate sample-size requirements for specified SEM models (Wang & Wang, 2012).
If the target path is relatively strong, the information requirement may differ from the requirement for reliably detecting a weak path.
The same is true of the measurement model. A simulation assuming strong loadings is not equivalent to one assuming weak loadings.
A defensible power calculation would therefore need assumptions not merely about N, but about the population model generating the data.
Why this case cannot produce a precise required N yet
At this point, the researcher has supplied enough information to define the broad model but not enough information to calculate one defensible model-specific minimum sample size.
The uploaded sources do not support deriving a precise N from the facts given here alone.
A model-specific power or Monte Carlo analysis would require a more complete set of planning values. Wang and Wang's (2012) model-based examples begin by specifying hypothesized population parameter values and then examine statistical power across candidate sample sizes.
For this case, the planning model would need defensible assumptions about at least:
- expected factor loadings;
- latent-variable variances and relationships;
- the structural coefficients of interest;
- residual and measurement-error variances;
- the estimator;
- indicator distributions;
- the parameter or model-fit criterion that defines adequate power;
- the amount and pattern of missing information where missingness is expected.
Without those inputs, returning an answer such as “N = 200 is enough” would imply more precision than the planning information supports.
Missingness can turn recruited N into something different
Suppose the researcher recruits exactly 200 participants but anticipates incomplete questionnaire responses.
The planning question is no longer simply whether 200 people are recruited. It becomes how much information the fitted model will actually contain and how missingness will be handled.
Tabachnick and Fidell (2013) identify missing data as a practical issue in SEM and note that problems can arise from simply deleting incomplete observations as well as from estimating missing values. Wang and Wang (2012) explicitly include missing-data handling among the factors affecting required SEM sample size and demonstrate Monte Carlo sample-size work incorporating missing values.
A model-specific simulation could therefore compare, for example:
- complete data;
- modest missingness;
- more substantial missingness under a specified mechanism.
The purpose would not be to subtract an arbitrary percentage from 200. It would be to determine how the planned analysis behaves under missing-data conditions that reasonably represent the intended study.
Measurement quality is not a post-analysis detail
Another problem with the 200-participant rule is that it encourages researchers to treat measurement quality as something to check after recruitment.
For SEM, measurement quality is part of the planning problem.
The CFA component explicitly represents relations between latent constructs, indicators, and measurement error (Lovric, 2011). Wang and Wang (2012) identify reliability of observed indicators as one determinant of SEM sample-size requirements, while Tabachnick and Fidell (2013) similarly note that reliable variables and stronger expected parameters can reduce sample demands.
For the synthetic study, the research team should therefore investigate what is already known about the proposed instrument before finalizing its power assumptions.
If prior evidence supports strong loadings and clearly distinguishable factors, those values can inform the planning model.
If the measures are new or uncertain, the researcher should not quietly assume ideal measurement. Sensitivity analysis across weaker and stronger measurement conditions would provide a more transparent rationale.
Model complexity means more than counting arrows
Complexity is sometimes reduced to the number of free parameters, but Wang and Wang (2012) treat model complexity as one among several interacting determinants of required sample size.
Consider what would happen if the researcher changed the proposed model.
The original model has:
- four latent constructs;
- three indicators per construct;
- five structural paths;
- one population.
Now suppose the researcher also wants to:
Group structure
Compare the SEM between two groups or test measurement invariance.
Measurement specification
Add correlated residuals or cross-loadings.
Additional effects
Estimate indirect effects or include latent interactions.
Different indicator treatment
Model ordered categorical indicators.
That is no longer the same sample-size problem.
The recruitment target should therefore correspond to the model that will actually answer the research question, not to the generic label “SEM.”
The 200-participant claim: what can actually be concluded?
For this synthetic case, the statement:
“SEM requires at least 200 participants.”
should neither be accepted nor rejected simply because the number 200 appears in a rule of thumb.
Wang and Wang (2012) review recommendations involving samples around 200 alongside smaller and larger rules, but conclude that no absolute standard applies to every SEM situation. Their discussion explicitly moves from such heuristics toward model-based power and sample-size methods.
A target of 200 participants is a candidate sample size to evaluate, not a universal threshold that establishes adequacy.
For this particular model, 200 should be evaluated under the proposed four-factor measurement structure, five structural paths, expected measurement quality, estimator, distributional conditions, missingness, and the smallest effects the researcher needs sufficient power to detect.
Only then can the team decide whether 200 appears adequate, inadequate, or too uncertain to defend.
What a defensible SEM sample-size rationale would require
A stronger planning process for this study would proceed as follows.
-
Freeze the substantive model
Specify the four constructs, 12 indicators, five structural paths, and any planned residual correlations, cross-loadings, covariates, indirect effects, or group comparisons before conducting the power analysis.
Sample size should correspond to the model actually intended for inference.
-
Check identification
Confirm the scaling of every latent variable, count observed information against free parameters, and examine the identification conditions of the complete model. Do not assume that additional participants can repair structural underidentification (Lovric, 2011; Wang & Wang, 2012).
-
Specify realistic measurement parameters
Use defensible evidence to propose plausible loading and measurement-error values.
Where measurement quality is uncertain, evaluate several plausible scenarios rather than assuming an ideal measurement model (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
-
Define the target effect
Identify which parameter or model property the study must be able to detect with acceptable probability.
If the primary concern is a structural path, its plausible magnitude needs to be specified. If overall model fit is the target, a fit-based sample-size method may address a different planning question. Wang and Wang (2012) discuss both parameter-focused power approaches and methods based on model-fit indices.
-
Match the estimator to the data
Decide whether the indicators will be modeled as continuous or categorical and whether distributional conditions require a robust or alternative estimator. Estimation method is part of the sample-size problem, not merely a software option selected after data collection (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
-
Represent anticipated missingness
If incomplete responses are plausible, incorporate the intended missing-data conditions into the planning analysis rather than assuming that recruited N equals analyzable information (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
-
Evaluate candidate sample sizes
Instead of asking software to confirm that 200 is acceptable, compare multiple candidate sample sizes under the same population model.
For example, a model-specific Monte Carlo study could evaluate a sequence of candidate Ns and examine parameter recovery, standard errors, power for the primary structural effect, and other relevant estimation behavior. Wang and Wang (2012) describe Monte Carlo simulation specifically as an approach to sample-size estimation for specified SEM models.
-
Conduct sensitivity analysis
Repeat the planning analysis under less favorable but plausible assumptions—for example, weaker loadings, a smaller structural effect, or greater missingness.
A recruitment target that works only under the most optimistic specification is less persuasive than one whose operating characteristics remain acceptable across a defensible range of conditions.
A sample rationale the researcher could eventually report
Once those analyses have actually been performed, the rationale should describe the model and assumptions rather than cite a generic SEM minimum.
A defensible report might take this form:
“Sample-size planning was based on the prespecified four-factor structural equation model rather than a universal SEM minimum. The planning model represented the 12-indicator measurement structure, five structural paths, anticipated indicator quality, the intended estimator, and plausible missing-data conditions. Candidate sample sizes were evaluated for the primary structural parameter under prespecified population values, with sensitivity analyses examining weaker measurement and effect assumptions. The final recruitment target was selected from those model-specific results.”
The missing element in the present synthetic case is deliberate: there is not yet a numerical final N.
The available information does not justify one.
What this case teaches
The central mistake is treating “SEM” as if it were a single statistical design with one sample-size requirement.
It is not.
The same N can support very different levels of information depending on the measurement model, loading strength, number and nature of parameters, structural effect sizes, estimator, distributions, missingness, and model complexity (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
Identification presents an even more fundamental distinction: an underidentified model is not repaired by increasing N (Lovric, 2011; Wang & Wang, 2012).
For the fictional researcher, therefore, 200 is neither approved nor rejected.
It is a candidate.
The next defensible step is to specify the actual population model and evaluate that candidate sample size—alongside alternatives—using model-specific power or simulation methods supported by the research question and the proposed analysis.
Conclusion
“SEM requires at least 200 participants” sounds useful because it converts a difficult planning problem into a single number.
But the number hides the decisions that actually determine whether an SEM study is adequately planned.
In this synthetic case, the proposed study has four latent variables, 12 indicators, a defined measurement structure, and five structural paths. The model passes an initial parameter-count identification check, but that does not establish sample-size adequacy. The answer still depends on measurement quality, target structural effects, estimator, distributional characteristics, missingness, and the operating characteristics required of the fitted model.
A defensible SEM sample-size rationale does not begin with 200.
It begins with the model.
References
Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2
Tabachnick, B. G., & Fidell, L. S. (2013). Using multivariate statistics (6th ed.). Pearson.
Wang, J., & Wang, X. (2012). Structural equation modeling: Applications using Mplus. John Wiley & Sons.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.