The CFA Did Not Fit: Should We Modify the Model?
Poor CFA model fit does not mean you should keep following modification indices until CFI or RMSEA reaches a preferred value. This synthetic case study shows how to diagnose loadings, residuals, and modification evidence, distinguish theory-supported revisions from model fishing, and validate a revised confirmatory factor analysis.
Synthetic-case disclosure: All variables, sample characteristics, estimates, fit statistics, residuals, and modification indices below are invented for methodological illustration. They are not empirical findings.
A confirmatory factor analysis (CFA) does not become successful simply because enough parameters are freed to produce attractive fit statistics. CFA starts with a hypothesized measurement structure, evaluates how well that structure reproduces the observed relationships among indicators, and then asks whether any evidence-based revisions remain substantively defensible. CFA therefore differs from an unrestricted search for whatever factor structure happens to fit one dataset best (Lovric, 2011; Tabachnick & Fidell, 2013; Wang & Wang, 2012).
That distinction becomes difficult to maintain when the first model fits poorly. The analyst sees a long list of modification indices, notices that freeing parameters improves fit, and proposes repeating the process until the usual fit guidelines are reached.
The stakeholder's practical question: How do I improve CFA model fit without turning CFA into exploratory model fishing?
This synthetic case study shows why the answer is not "never modify a CFA." The better rule is to diagnose the source of misfit, distinguish theory-supported revisions from purely sample-driven ones, and treat substantial post hoc modification as something that requires validation rather than as confirmation of the original model (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
The research problem
A research team develops a 12-item questionnaire intended to measure three related dimensions of workplace adaptation:
- Role Clarity (RC): RC1–RC4
- Team Support (TS): TS1–TS4
- Work Confidence (WC): WC1–WC4
The substantive theory specifies three correlated latent factors. Each item is assigned to one factor, cross-loadings are fixed to zero, and indicator residuals are initially assumed to be uncorrelated.
This is a genuinely confirmatory question: Does the proposed three-factor measurement structure provide a reasonable representation of the observed data?
CFA permits researchers to impose theoretically specified restrictions such as fixing particular factor loadings to zero. More generally, SEM represents hypothesized relationships in a testable model rather than discovering those relationships without prior specification (Lovric, 2011). Tabachnick and Fidell (2013) similarly emphasize that SEM is fundamentally confirmatory and requires prior hypotheses about relationships among variables.
The model should therefore be specified before inspecting which omitted parameters would improve fit.
Step 1: Start with the measurement theory, not the modification indices
The original model is:
Indicator structure
- RC1–RC4 load only on Role Clarity.
- TS1–TS4 load only on Team Support.
- WC1–WC4 load only on Work Confidence.
Model restrictions
- The three latent factors may correlate.
- Cross-loadings are fixed at zero.
- Residual covariances among indicators are fixed at zero.
The last two restrictions matter. A CFA does not merely say which parameters are estimated; it also says which relationships are assumed absent.
Wang and Wang (2012) describe SEM as a sequence beginning with model formulation, followed by identification, estimation, evaluation, and—if needed—modification. They explicitly note that formulation may be based on theory or empirical findings and that modification follows evaluation rather than replacing initial specification.
This ordering protects the scientific question. If the model is repeatedly rewritten after seeing the data, the final model is no longer equivalent to the model that was originally proposed.
Step 2: Check identification before debating fit
A model cannot be meaningfully evaluated if its parameters cannot be uniquely estimated. Model identification asks whether a unique numerical solution exists for the free parameters (Lovric, 2011; Wang & Wang, 2012).
For this synthetic CFA, each factor has four indicators, a scale-setting constraint is imposed for each latent factor, and the overall model has positive degrees of freedom. The software converges without an identification warning or inadmissible estimates.
The model is therefore treated as identified for this demonstration.
Identification is not evidence that the model is correct. Identification establishes whether the proposed model can be estimated uniquely; model fit addresses a different question about correspondence between the model and the observed data (Lovric, 2011; Wang & Wang, 2012).
Step 3: The initial CFA fits poorly
Suppose the CFA is estimated in a synthetic sample of N = 520.
Synthetic initial model-fit output
| Evidence | Initial CFA |
|---|---|
| χ² | 238.6 |
| df | 51 |
| p | < .001 |
| CFI | .872 |
| TLI | .834 |
| RMSEA | .084 |
| RMSEA 90% CI | [.074, .095] |
| SRMR | .071 |
The point is not to turn these values into a mechanical pass/fail checklist.
The model chi-square evaluates discrepancy between the observed and model-implied covariance structures, but it is sensitive to factors including sample size and departures from multivariate normality. Wang and Wang (2012) therefore caution against rejecting a model solely because its chi-square test is significant. Lovric (2011) likewise recommends considering additional descriptive fit criteria and fit indices because chi-square has recognized weaknesses.
Fit indices also require interpretation rather than ritual threshold checking. Wang and Wang (2012), for example, describe .90 and .95 as rule-of-thumb values discussed for CFI rather than as immutable laws, and they note characteristics of the data that can affect fit indices. RMSEA is an approximate-fit measure that can also be reported with a confidence interval (Wang & Wang, 2012).
Here, the evidence is more persuasive because several indicators point in the same direction. The incremental fit indices are relatively weak, RMSEA indicates nontrivial approximation error, and local diagnostics will soon reveal concentrated areas of discrepancy.
Initial conclusion: The prespecified three-factor model does not reproduce the synthetic data sufficiently well to stop the diagnostic process.
That is not yet a justification for adding every parameter suggested by the software.
Step 4: Inspect the factor loadings
Global fit tells us that the model and data disagree. Factor loadings help locate possible measurement weaknesses.
Synthetic standardized loadings
| Factor | Item | Loading |
|---|---|---|
| Role Clarity | RC1 | .78 |
| Role Clarity | RC2 | .74 |
| Role Clarity | RC3 | .69 |
| Role Clarity | RC4 | .43 |
| Team Support | TS1 | .82 |
| Team Support | TS2 | .79 |
| Team Support | TS3 | .76 |
| Team Support | TS4 | .71 |
| Work Confidence | WC1 | .80 |
| Work Confidence | WC2 | .77 |
| Work Confidence | WC3 | .72 |
| Work Confidence | WC4 | .68 |
RC4 stands out as the weakest indicator.
A weak loading deserves investigation because factor loadings describe the relationship between observed indicators and their factors. Wang and Wang (2012) also show that squared standardized loadings can be interpreted in terms of how much indicator variance is accounted for by the factor.
But RC4 should not automatically be deleted.
Review the item content
- RC1: "I understand what is expected of me."
- RC2: "My responsibilities are clearly defined."
- RC3: "I know which tasks have priority."
- RC4: "I know whom to ask when an unusual problem occurs."
RC4 is related to role clarity, but it also has an interpersonal-help component. The lower loading now has a plausible substantive explanation.
This is stronger evidence than saying, "RC4 has the lowest loading, so remove it."
Measurement decisions should preserve the intended construct rather than optimize statistics at the expense of what the scale is supposed to represent. Construct validity concerns whether a measure reflects the intended hypothetical construct, while content validity asks whether the items adequately represent its relevant aspects (Adams & Lawrence, 2018).
Step 5: Inspect residuals as local evidence of misfit
A CFA can have disappointing global fit because particular relationships are poorly reproduced.
Residuals describe discrepancies between observed and model-reproduced relationships. In factor analysis, a good solution should generally leave small residual correlations; large residuals point toward relationships the model has not adequately represented (Tabachnick & Fidell, 2013). In SEM more generally, fit is based on how closely the model-implied covariance matrix corresponds to the observed covariance matrix (Lovric, 2011; Wang & Wang, 2012).
Suppose the largest synthetic standardized residuals include:
| Indicator pair | Standardized residual |
|---|---|
| TS2–TS3 | 3.8 |
| RC4–TS1 | 3.5 |
| RC3–RC4 | 2.9 |
| WC2–WC3 | 2.1 |
Now the misfit has structure.
TS2 and TS3 are both Team Support indicators and have unusually similar wording:
- TS2: "My team gives me useful feedback when I need help."
- TS3: "My team gives me useful guidance when I encounter difficulties."
Their association may therefore contain something beyond the common Team Support factor—for example, shared wording or highly overlapping content.
RC4–TS1 raises a different possibility. RC4 may partly represent seeking interpersonal assistance, which could explain why it shares unexplained covariance with a Team Support indicator.
These are hypotheses about misspecification. They are not yet permission to modify the model.
Step 6: What do the CFA modification indices say?
A modification index (MI) evaluates how much model chi-square would be expected to decrease if a currently constrained parameter were freed. Modification indices are therefore diagnostic statistics for identifying possible misspecification (Wang & Wang, 2012). Tabachnick and Fidell (2013) describe the related Lagrange multiplier logic as asking which fixed parameters, if estimated, would improve model fit.
Suppose the largest synthetic MIs are:
| Proposed change | MI | Expected parameter change | Initial substantive assessment |
|---|---|---|---|
| Correlate residuals TS2 ↔ TS3 | 31.4 | .29 | Plausible shared wording |
| Cross-load RC4 → Team Support | 24.8 | .34 | Plausible, but changes item meaning |
| Correlate residuals RC3 ↔ RC4 | 15.7 | .18 | Possible shared role-information content |
| Cross-load WC3 → Team Support | 12.1 | .21 | No clear theoretical rationale |
| Correlate residuals RC1 ↔ WC4 | 10.9 | .16 | No obvious substantive rationale |
The dangerous workflow
- Free the parameter with MI = 31.4.
- Rerun the CFA.
- Inspect the new MIs.
- Free the next largest parameter.
- Repeat until CFI, TLI, RMSEA, or another statistic reaches a preferred value.
That procedure turns fit statistics into an optimization target.
Wang and Wang (2012) state that there is no strict universal rule for how large an MI must be before a modification becomes meaningful. More importantly, they emphasize that model respecification should be both statistically and theoretically driven and explicitly caution against adding or removing parameters solely to improve model fit.
Step 7: Separate defensible revision from model fishing
Consider the five proposed modifications.
Revision A: TS2 ↔ TS3 residual covariance
This is potentially defensible.
The two items belong to the same intended construct and use unusually parallel language about receiving help from the team. If the research team can justify why these indicators share something not represented by the common factor, a residual covariance has a substantive interpretation.
Decision principle: The rationale comes from item content, not merely from MI = 31.4.
Revision B: RC4 cross-loading on Team Support
This is more consequential.
The item asks whom the respondent approaches when a problem occurs. The MI supports the possibility that RC4 reflects Team Support as well as Role Clarity.
Decision principle: Allowing the cross-loading changes the measurement theory. If the wording confounds two constructs, revising or removing the item in a future instrument may be more interpretable than preserving it through a post hoc cross-loading.
Revision C: RC3 ↔ RC4 residual covariance
There is some substantive rationale because both indicators concern information about performing one's role.
Decision principle: If RC4 already has questionable construct specificity, adding a correlated residual may conceal rather than resolve that measurement problem.
Revision D: WC3 cross-loading on Team Support
The MI is statistically noticeable, but the researchers cannot identify a credible measurement rationale.
Decision: Do not add it merely because it improves fit.
Revision E: RC1 ↔ WC4 residual covariance
There is no clear substantive explanation.
Decision: Do not add it solely because the software suggests it.
Central distinction: Modification indices can tell the analyst where a constraint conflicts with the observed covariance structure. They cannot decide whether the resulting parameter makes substantive sense.
Step 8: Make one theory-supported revision and reassess the whole model
The team decides that the TS2–TS3 residual covariance has the clearest substantive justification because the items have nearly parallel wording and belong to the same construct.
They free one parameter and rerun the CFA.
Synthetic revised model
| Evidence | Initial model | One defensible revision |
|---|---|---|
| χ² | 238.6 | 190.2 |
| df | 51 | 50 |
| CFI | .872 | .904 |
| TLI | .834 | .873 |
| RMSEA | .084 | .074 |
| SRMR | .071 | .060 |
Fit improves, but the model is not suddenly declared "validated."
That matters.
A statistically meaningful improvement from a nested model comparison answers whether freeing the restriction improves correspondence with these data. It does not establish that the new parameter is theoretically correct or that the improvement will reproduce in another sample.
The revised model also still contains the weak RC4 loading and localized residual evidence involving RC4. The diagnostic question therefore remains open.
Step 9: Do not hide a problematic item behind enough extra parameters
At this stage, suppose the largest remaining MI suggests cross-loading RC4 on Team Support.
The team has two competing interpretations:
Model 2A: Retain RC4 and estimate a cross-loading
This recognizes that the item may represent both Role Clarity and Team Support.
Model 2B: Remove RC4
This treats RC4 as an ambiguously worded indicator that does not cleanly operationalize the intended Role Clarity construct.
Neither choice should be made from fit statistics alone.
Removing an indicator changes the content represented by the measurement instrument. Allowing a cross-loading changes the hypothesized measurement structure. Construct and content considerations therefore need to accompany the statistical evidence (Adams & Lawrence, 2018).
For this synthetic case, the research team returns to the questionnaire-development rationale and concludes that RC4 was intended specifically to measure role knowledge, not interpersonal support. Its wording failed to isolate that content adequately.
The team therefore removes RC4 as a provisional post hoc revision, while documenting that the item should be rewritten and tested in future data.
Synthetic revised model after removing RC4
| Evidence | Initial CFA | TS2–TS3 residual | Residual + RC4 removed |
|---|---|---|---|
| χ² | 238.6 | 190.2 | 109.8 |
| df | 51 | 50 | 40 |
| CFI | .872 | .904 | .951 |
| TLI | .834 | .873 | .933 |
| RMSEA | .084 | .074 | .058 |
| SRMR | .071 | .060 | .043 |
The temptation is now to write:
"The final CFA demonstrated good fit, confirming the hypothesized measurement model."
That would overstate what happened.
The final model was partly constructed after inspecting the same sample used to evaluate it.
Step 10: Why repeated same-sample modification changes the evidential status
Tabachnick and Fidell (2013) draw an important line between confirmatory and exploratory SEM. They note that although alternative models and modifications can be examined after estimation, testing numerous modifications in search of the best-fitting model moves the analysis toward exploratory data analysis. They recommend caution about significance levels and cross-validation with another sample whenever possible.
This is the central statistical danger of modification-index fishing.
Each time the analyst examines the sample, selects a promising parameter, refits the model, and selects again, information from that sample is being used to construct the model. The apparent performance of the final specification therefore no longer represents a clean test of an independently specified hypothesis.
A beautifully fitting post hoc model may partly reflect peculiarities of that dataset.
The issue is related to the broader problem of overfitting: a solution can fit the observed sample so well that it does not generalize adequately beyond it (Tabachnick & Fidell, 2013).
Evidential consequence: The more extensive the post hoc modifications, the less defensible it becomes to describe the final CFA as purely confirmatory.
Step 11: Validate the revised model on new information
Suppose the original synthetic dataset had been divided before analysis:
Development sample
n = 260
The initial model is evaluated here, and the TS2–TS3 residual covariance and removal of RC4 are proposed here.
Validation sample
n = 260
The revised specification is evaluated here after it has been frozen.
The revised specification is then frozen.
No new paths are added after examining the validation sample.
Wang and Wang (2012) describe cross-validation modeling as using separate calibration and validation samples to examine whether model estimates replicate. Tabachnick and Fidell (2013) likewise recommend cross-validation when model searching has occurred.
Suppose the frozen revised CFA produces the following synthetic validation results:
| Evidence | Development sample | Validation sample |
|---|---|---|
| CFI | .953 | .946 |
| TLI | .935 | .927 |
| RMSEA | .057 | .061 |
| SRMR | .044 | .048 |
The numbers are not identical, nor should they be expected to be.
More important is that the revised model remains broadly plausible when evaluated on observations that did not determine the modifications. Replication with genuinely new data would provide still stronger evidence. Adams and Lawrence (2018) describe replication as a means of examining whether findings recur when a study is repeated, reinforcing why evidence from new observations matters for generalization.
What if the revised model fails validation?
That outcome is informative.
Suppose the development model reached CFI = .972 after eight modification-index-driven changes, but the frozen specification produced CFI = .881 in the validation sample.
The appropriate conclusion would not be:
"We need eight more modification indices in the validation sample."
Doing that simply starts another sample-specific search.
Instead, failure to replicate should trigger reconsideration of the measurement theory, item wording, population, data characteristics, and whether the hypothesized factor structure is sufficiently stable. The revised model may have captured development-sample idiosyncrasies rather than a reproducible measurement structure.
A practical decision framework for poor CFA model fit
When a CFA does not fit well, use the following sequence rather than chasing a target fit statistic.
1. Return to the prespecified model
Document:
- intended constructs;
- indicator-to-factor assignments;
- allowed factor correlations;
- fixed cross-loadings;
- residual assumptions;
- estimator and data characteristics.
CFA begins with hypotheses about the measurement structure, not with the modification-index table (Lovric, 2011; Tabachnick & Fidell, 2013; Wang & Wang, 2012).
2. Confirm that the model is identified and estimable
Check whether the parameters can be uniquely estimated and investigate convergence problems, improper estimates, or other estimation warnings before interpreting fit (Lovric, 2011; Wang & Wang, 2012).
3. Evaluate multiple forms of fit evidence
Consider the model chi-square, several appropriate fit indices, and approximate-fit information rather than treating one number as decisive. Recognize that fit measures have different properties and limitations (Lovric, 2011; Wang & Wang, 2012).
4. Inspect the measurement parameters
Review factor loadings and other parameter estimates for weak, unstable, implausible, or substantively confusing indicators (Wang & Wang, 2012).
5. Inspect local misfit
Examine residual discrepancies to identify relationships the model is reproducing poorly (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
6. Use modification indices as diagnostic clues
An MI can identify a fixed parameter whose release is expected to improve model fit. It does not establish that the parameter belongs in the measurement theory (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
7. Require a substantive explanation before modifying
Ask:
- Why should these residuals covary?
- Why should this indicator load on another factor?
- Does the item wording support the proposed relationship?
- Does the revision clarify the construct or merely improve a statistic?
- Would we have considered this parameter plausible before seeing its MI?
If the only rationale is "the MI is large," the revision is not sufficiently theory-supported (Wang & Wang, 2012).
8. Keep the original model visible
Report the prespecified model before reporting post hoc revisions. This preserves the distinction between the original confirmatory test and the model developed after examining the data.
9. Treat extensive respecification as exploratory
Once many modifications have been selected from the same observations, the analysis has moved away from a clean confirmatory test. Describe that status explicitly rather than relabeling the resulting model as the originally hypothesized CFA (Tabachnick & Fidell, 2013).
10. Freeze and validate
Evaluate the revised model on separate observations when possible. A model that survives independent validation provides stronger evidence than one optimized and judged entirely on the same sample (Adams & Lawrence, 2018; Tabachnick & Fidell, 2013; Wang & Wang, 2012).
Workflow: theory → identification → estimation → global fit → loadings → residuals → modification evidence → substantive justification → revised model → validation.
What the analyst should report
A defensible report for this synthetic case would say that the prespecified three-factor CFA showed meaningful evidence of misfit. Inspection of factor loadings and residuals identified RC4 as a comparatively weak and conceptually ambiguous indicator and identified excess shared variation between TS2 and TS3. Modification indices were considered as diagnostic evidence, but revisions were restricted to changes that could be justified from item content and measurement theory (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
The revised model improved fit after allowing the theoretically interpretable TS2–TS3 residual covariance and provisionally removing RC4. Because these decisions were made after inspecting the development data, the revised model was treated as a post hoc specification rather than as an untouched confirmation of the original model. Its performance was therefore evaluated separately in validation data.
That description is less dramatic than "we fixed the CFA."
It is also more informative.
Interpretation and limitations
This synthetic case does not establish universal fit-index thresholds, a universal minimum loading, or a universal modification-index cutoff. Wang and Wang (2012) explicitly describe commonly used CFI values as rules of thumb and state that no strict rule determines how large an MI must be to warrant meaningful modification. Fit evidence should therefore be interpreted as part of a larger assessment rather than as a collection of absolute laws.
A well-fitting CFA also does not prove that the measurement theory is uniquely correct. Different models can sometimes imply the same covariance structure and therefore produce identical fit evidence despite having different theoretical interpretations (Lovric, 2011).
Nor does CFA itself establish causal relationships. SEM represents hypothesized dependency structures, but causal interpretation ultimately requires support from the research design and substantive assumptions rather than from model fit alone (Lovric, 2011; Tabachnick & Fidell, 2013).
Finally, validation is not a magical repair for weak theory. A model can replicate and still omit important aspects of a construct, use problematic indicators, or have limited applicability beyond the populations and settings represented by the available samples. Measurement content remains part of the validity argument (Adams & Lawrence, 2018).
Conclusion: Improve the model by improving the explanation
When CFA model fit is poor, modification indices are useful because they can point toward restrictions that conflict with the observed data. They become dangerous when treated as instructions for adding parameters until conventional fit statistics look acceptable.
The stronger workflow is:
theory → identification → estimation → global fit → loadings → residuals → modification evidence → substantive justification → revised model → validation.
The goal is not the smallest RMSEA or largest CFI obtainable from one sample. The goal is a measurement model whose parameters are estimable, whose discrepancies are understood, whose revisions have substantive meaning, and whose performance can survive evaluation beyond the observations that suggested those revisions (Tabachnick & Fidell, 2013; Wang & Wang, 2012).
If repeated modification is required to make the model look convincing, the correct interpretation may be that the original CFA was not confirmed—and that the revised structure is now a hypothesis to test, not a conclusion already established.
References
- Adams, K. A., & Lawrence, E. K. (2018). Research methods, statistics, and applications (2nd ed.). SAGE Publications.
- Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2
- Tabachnick, B. G., & Fidell, L. S. (2013). Using multivariate statistics (6th ed.). Pearson.
- Wang, J., & Wang, X. (2012). Structural equation modeling: Applications using Mplus. John Wiley & Sons.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.