Can You Compare Group Scores Yet? Measurement Invariance Before Group Comparisons
Measurement invariance testing determines whether group comparisons are interpretable by assessing whether a construct is measured sufficiently equivalently across groups. This Resource explains the CFA hierarchy from configural through strict invariance and the distinct MICOM workflow required for PLS-SEM.
Researchers often reach the group-comparison question too early.
A questionnaire has been administered in two countries, two organizational units, or two demographic groups. The next analysis appears obvious: compare the means, compare the latent variables, or estimate whether a structural path is stronger in one group than another.
But those comparisons contain an earlier measurement question:
Does the construct have sufficiently equivalent measurement properties across the groups for the proposed comparison to mean what you think it means?
Measurement invariance testing addresses that question. In multigroup CFA, it evaluates whether increasingly restrictive features of the factor model can reasonably be treated as equivalent across groups. In PLS-SEM, measurement invariance of composite models requires a related but distinct procedure because common-factor invariance procedures cannot simply be transferred to composite models (Hair et al., 2017; Wang & Wang, 2012).
The decision sequence should therefore be:
1. Define the comparison
Specify the substantive group comparison you ultimately intend to make.
2. Establish the measurement model
Establish an acceptable measurement model in the groups.
3. Test the required invariance
Evaluate the level of measurement invariance required by the intended comparison.
4. Make the permitted comparison
Only then make the substantive group comparison that the established level of invariance permits (Wang & Wang, 2012).
The problem: a group difference can be a measurement difference
Suppose researchers measure organizational adaptability with six questionnaire items in two synthetic countries, Country A and Country B.
The investigators want to answer three questions:
- Is average adaptability different between the countries?
- Is the relationship between adaptability and work engagement different between the countries?
- Can the two samples be pooled into one model?
Those are not interchangeable questions. Each requires evidence that the measurement system behaves sufficiently similarly across groups for the intended comparison.
Wang and Wang (2012) describe measurement invariance as the requirement that observed scale indicators measure the same theoretical constructs across groups. Without it, apparent group differences cannot be interpreted unambiguously because the groups may differ partly in how the indicators relate to the underlying construct rather than in the construct itself.
This is why acceptable reliability or a well-fitting CFA in the combined sample is not enough. A pooled model can conceal group-specific measurement behavior, while separate group analyses do not establish that parameters from the two fitted models are actually comparable (Wang & Wang, 2012).
What measurement invariance means
Measurement invariance—also called measurement equivalence in some contexts—concerns whether a measurement system operates equivalently across groups. In a CFA measurement model, this involves properties such as the pattern and magnitude of factor loadings, indicator intercepts, and, at the strictest level, residual variances (Wang & Wang, 2012).
The logic is easiest to see through the measurement equation. An observed item score is not treated as a direct observation of the latent construct. It reflects a relationship between the latent factor and the indicator, together with an intercept and residual component. If those measurement parameters differ across groups, the same observed response pattern need not carry the same latent-variable interpretation in both groups (Wang & Wang, 2012).
Important distinction: Measurement invariance does not mean that the groups must have identical latent means, variances, covariances, or structural relationships. Those substantive differences may be exactly what the study is designed to investigate. The purpose of invariance testing is to separate differences in the measurement system from differences in the constructs and relationships being studied (Wang & Wang, 2012).
Start with multigroup modeling, not two unrelated analyses
A multigroup SEM analyzes two or more observed groups simultaneously and allows equality restrictions to be imposed on selected parameters. The groups may represent countries, regions, demographic categories, or other mutually exclusive observed classifications (Wang & Wang, 2012).
The multigroup framework matters because the inferential question is not merely whether a coefficient is significant in Group A and nonsignificant in Group B. The question is whether the relevant model parameter can be constrained to equality across groups or whether there is evidence of a difference under the multigroup model (Wang & Wang, 2012).
Tabachnick and Fidell (2013) similarly describe multiple-group SEM as a way to determine whether covariance structures, regression coefficients, means, or other aspects of a model differ between groups. They note that latent means are typically considered in a multiple-group framework in which constraints are imposed on the factor structure.
Measurement invariance CFA: the hierarchy of questions
Wang and Wang (2012) describe four levels of measurement invariance in multigroup CFA: configural, weak, strong, and strict invariance. Weak invariance is also called metric invariance, while strong invariance incorporates what is often called scalar invariance.
These levels are hierarchical. Failure at an earlier stage changes what later comparisons can reasonably mean.
1. Configural invariance: are we measuring the same configuration?
Configural invariance asks whether the groups share the same basic factor structure: the same number of factors and the same pattern of freely estimated and fixed factor loadings. Equality restrictions are not yet imposed on the numerical values of the other measurement parameters (Wang & Wang, 2012).
The practical workflow begins by establishing substantively meaningful baseline CFA models for the individual groups and then integrating them into a simultaneous multigroup CFA. Wang and Wang emphasize that configural invariance is a necessary condition for proceeding to equality tests of loadings, intercepts, and residual variances (Wang & Wang, 2012).
If configural invariance fails, the problem is fundamental. The indicators may not be representing the same construct configuration across the groups. Further invariance restrictions are therefore not a defensible way to manufacture comparability (Wang & Wang, 2012).
What becomes defensible? At this stage, you have evidence that broadly the same factor pattern can represent both groups. You do not yet have evidence that the factor operates on the same metric or that group means are comparable.
2. Weak or metric invariance: are factor loadings equivalent?
Weak measurement invariance requires equality of factor loadings across groups. Factor loadings represent the strength of the relationships between indicators and their underlying factors, so loading invariance asks whether the indicators relate to the latent construct in the same way across groups (Wang & Wang, 2012).
Wang and Wang call this metric invariance because invariant factor loadings place the measurement relationships on the same scale across groups. Once loading invariance is established, comparisons involving factor variances and relationships among factors become meaningful in their framework (Wang & Wang, 2012).
For the synthetic adaptability example, suppose Item 2 is strongly related to adaptability in Country A but only weakly related to it in Country B. A difference in a downstream adaptability coefficient could then partly reflect a different measurement relationship rather than a genuine difference in the structural process.
What becomes defensible? Metric invariance supports moving toward comparisons of relationships involving the latent constructs. It does not, by itself, justify a latent-mean comparison because indicator intercept equivalence has not yet been established (Wang & Wang, 2012).
3. Strong or scalar invariance: do the indicators share a common origin?
Strong measurement invariance requires invariant factor loadings and invariant indicator intercepts. Wang and Wang (2012) describe the intercept as the origin or scalar of the measurement.
A noninvariant intercept means that respondents in one group may tend to score systematically higher or lower on an item even when the underlying latent construct is held at the same level. That is precisely the kind of measurement difference that can contaminate a group mean comparison (Wang & Wang, 2012).
Once both loadings and indicator intercepts are invariant, the observed indicators share the same measurement metric and scalar across groups. Wang and Wang therefore identify this stage as the point at which latent factor means can meaningfully be compared (Wang & Wang, 2012).
Tabachnick and Fidell (2013) likewise note that interpretable latent-mean comparisons are made in multiple-group models with constraints on the factor structure; one group's latent mean is commonly fixed for identification and the other group's mean is estimated relative to that reference.
What becomes defensible? With strong/scalar invariance, latent-mean comparisons become interpretable under the fitted model. Before this stage, a difference in apparent group level can remain confounded with differences in indicator intercepts (Wang & Wang, 2012).
4. Strict invariance: are residual variances also invariant?
Strict measurement invariance adds equality of indicator residual variances to metric and scalar invariance. It therefore represents a stronger requirement about the measurement-error portion of the model (Wang & Wang, 2012).
Wang and Wang also note that residual-variance invariance is often not required for the substantive comparisons researchers usually prioritize. They describe strict invariance as relevant to questions involving invariance of indicator reliability, while noting that many applications do not treat equality of residual variances as a necessary condition for the usual group comparisons (Wang & Wang, 2012).
What becomes defensible? Strict invariance supports a stronger claim about equivalence of the measurement-error structure. It should not be imposed merely because it is the next item in a checklist.
CFA invariance levels at a glance
| Level | Measurement requirement | What the evidence supports |
|---|---|---|
| Configural invariance | Same basic factor structure and loading pattern across groups. | Broadly the same factor configuration can represent the groups; equal metric or comparable group means have not yet been established. |
| Weak / metric invariance | Equal factor loadings across groups. | Comparisons involving relationships among latent constructs can be considered; latent-mean comparison is not yet justified by this level alone. |
| Strong / scalar invariance | Equal factor loadings and indicator intercepts. | Latent factor means can be meaningfully compared under the fitted model. |
| Strict invariance | Equal loadings, intercepts, and indicator residual variances. | Supports a stronger claim about equivalence of the measurement-error structure. |
A practical CFA decision rule
The required invariance level should be chosen from the comparison you intend to make, not from a desire to pass every possible test.
If the research question is whether a latent mean differs across groups, strong/scalar invariance is central because equality of loadings alone does not establish a common indicator origin. If the focus is on relationships among latent constructs, metric invariance is a key measurement requirement in Wang and Wang's hierarchy. If the interest extends to equivalent indicator residual properties, strict invariance addresses the additional measurement-error question (Wang & Wang, 2012).
That makes measurement invariance testing a decision problem rather than a ritual sequence of software commands.
Why summed or composite scores do not escape the problem
Computing a total score does not make measurement nonequivalence disappear.
If items represent the underlying construct differently across groups, averaging or summing those items can carry that measurement difference into the resulting score. Wang and Wang (2012) explicitly treat measurement invariance as a prerequisite for meaningful group comparison because otherwise differences in the measured variables cannot be unambiguously separated from differences in the construct.
The conclusion should therefore match the evidence. A researcher who has not examined cross-group measurement equivalence should avoid describing an observed score difference as if it necessarily represented a difference in the underlying construct.
PLS-SEM requires a different invariance workflow
PLS-SEM introduces an important qualification.
Hair et al. (2017) state that measurement-invariance procedures developed for CB-SEM common-factor models cannot simply be transferred to PLS-SEM composite models. For PLS-SEM they describe the measurement invariance of composite models (MICOM) procedure.
MICOM contains three hierarchically related steps:
- Configural invariance
- Compositional invariance
- Equality of composite means and variances (Hair et al., 2017)
These are not merely renamed CFA metric and scalar tests. They address invariance in a composite-model framework.
MICOM Step 1: configural invariance
PLS-SEM configural invariance requires that the composites be specified and estimated consistently across groups. Hair et al. (2017) identify identical indicators, identical data treatment, and identical algorithm settings as core elements of this assessment.
The objective is to establish that the construct has been operationalized in the same way before testing whether the estimated composite itself behaves equivalently across groups (Hair et al., 2017).
MICOM Step 2: compositional invariance
Compositional invariance examines whether the composite is formed equivalently across groups. Hair et al. (2017) describe a permutation-based assessment that evaluates whether group-specific composite scores correspond sufficiently despite potential differences in estimated indicator weights.
This stage is critical for multigroup analysis. Hair et al. state that configural plus compositional invariance establishes partial measurement invariance in the MICOM terminology and permits comparison of path-coefficient estimates across groups (Hair et al., 2017).
An applied cross-national example in Avkiran and Ringle (2018) follows this logic: the study first assesses configural invariance, then uses a permutation procedure to assess compositional invariance, and only after those conditions are supported proceeds to compare structural path coefficients across the two samples.
MICOM Step 3: equality of composite means and variances
The third MICOM step evaluates whether the composites also have equal mean values and variances across groups. When configural invariance, compositional invariance, and equality of composite means and variances are all supported, Hair et al. (2017) refer to full measurement invariance.
Their distinction has a practical consequence: configural plus compositional invariance supports multigroup comparison of path coefficients, whereas the additional equality of composite means and variances supports analysis at the pooled-data level, subject to attention to structural heterogeneity (Hair et al., 2017).
Avkiran and Ringle (2018) provide an applied illustration in which full MICOM invariance supported pooling after the researchers had also examined cross-group structural differences.
CFA and MICOM answer related but different invariance questions
| Framework | Invariance sequence described in the source | Practical implication |
|---|---|---|
| Multigroup CFA | Configural → weak/metric → strong/scalar → strict invariance. | The required level depends on the substantive parameter to be compared, including relationships, latent means, or residual properties. |
| PLS-SEM / MICOM | Configural → compositional → equality of composite means and variances. | Configural plus compositional invariance permits path-coefficient comparisons; the additional third step establishes full MICOM invariance and supports pooling under the conditions described by the procedure. |
Do not mix up two meanings of “partial invariance”
The phrase partial measurement invariance needs special care because the approved sources use it differently across frameworks.
PLS-SEM / MICOM
Hair et al. (2017) use partial measurement invariance in a precise procedural sense: configural and compositional invariance have been established, while equality of composite means and variances is a further condition for full invariance.
CFA literature
In the CFA literature discussed by Wang and Wang (2012), partial measurement invariance instead refers to situations in which full equality of measurement parameters is not obtained and selected noninvariant parameters may be freed or problematic indicators reconsidered. Wang and Wang note that such approaches have been proposed but also explicitly describe them as debatable.
These two uses should not be collapsed into one rule. A statement such as “partial invariance was achieved” is incomplete unless the researcher identifies the modeling framework, the parameters tested, the constraints released or retained, and the substantive comparison that remains justified.
Synthetic example: Country A versus Country B
Return to the fictional adaptability scale.
-
Suppose a multigroup CFA shows that the six indicators load on the same single factor in Country A and Country B. Configural invariance is plausible, but this does not yet establish equal loadings.
-
Next, suppose the equality constraints on the loadings are also supportable. The scale now has metric invariance under the fitted CFA model. Comparisons of relationships involving the adaptability factor can be considered, but a latent-mean comparison still requires the intercept question to be addressed (Wang & Wang, 2012).
-
Now suppose one indicator shows a noticeably different intercept across the groups and the strong-invariance model is not supportable as specified.
The correct conclusion is not:
“Country A has higher adaptability.”
At this point, the analysis has identified a measurement problem directly relevant to that claim. Researchers could investigate the source of the noninvariance and, where substantively justified, consider a partial-invariance model, but Wang and Wang (2012) caution that such strategies are debatable and require careful attention to identification and the invariant measurement parameters.
The available measurement model does not yet justify a straightforward latent-mean comparison under full strong invariance.
If the same study instead used PLS-SEM, the researcher should not reproduce the CFA metric/scalar sequence mechanically. MICOM would first establish common specification and estimation, then test compositional invariance, and only then assess equality of composite means and variances as required by the intended analysis (Hair et al., 2017).
What multigroup analysis can and cannot tell you
Measurement invariance and substantive invariance are different questions.
Once the required measurement invariance is established, researchers may test whether factor covariances, latent means, or other structural parameters differ across groups. A difference in a structural parameter does not imply that the instrument is defective; it can represent the population heterogeneity that motivated the research question in the first place (Wang & Wang, 2012).
Conversely, a path estimated as significant in one group and nonsignificant in another does not by itself establish a group difference. Multigroup modeling is designed to evaluate the cross-group parameter comparison directly rather than infer a difference from two separate significance decisions (Wang & Wang, 2012).
Measurement equivalence first → substantive heterogeneity second.
Common mistakes in measurement invariance testing
Comparing means too early
A mean difference may be entangled with differences in indicator loadings or intercepts (Wang & Wang, 2012).
Assuming pooled CFA fit proves invariance
Overall fit does not demonstrate that measurement parameters are equivalent across groups (Wang & Wang, 2012).
Overinterpreting configural invariance
Configural invariance establishes the common pattern, not equality of the metric or scalar (Wang & Wang, 2012).
Comparing latent means after metric invariance alone
Wang and Wang require invariant loadings and intercepts for interpretable factor-mean comparisons (Wang & Wang, 2012).
Treating MICOM as CFA with different labels
Common-factor and composite-model invariance require different tests (Hair et al., 2017).
Running PLS multigroup analysis before MICOM
Hair et al. state that without MICOM configural and compositional invariance, group differences from the multigroup analysis are not validly interpretable (Hair et al., 2017).
Using “partial invariance” without defining it
The term has framework-specific meanings and should be accompanied by the actual invariant and noninvariant parameters (Hair et al., 2017; Wang & Wang, 2012).
A StatsAlly measurement invariance testing workflow
-
Define the intended comparison. Before reporting a group difference, document the comparison you ultimately want to make: observed score, latent mean, structural relationship, or pooled model.
-
Establish the measurement model in the groups. Establish an acceptable measurement model for each group and evaluate whether the same basic construct configuration is plausible.
-
Apply the framework-specific invariance sequence. In a CFA framework, proceed through the invariance constraints required for the intended inferential claim: configural structure, equal loadings, equal intercepts when latent means are to be compared, and residual equality when the research question requires that stronger condition (Wang & Wang, 2012).
-
Use MICOM for PLS-SEM. In PLS-SEM, use the composite-model logic instead: common specification and estimation, compositional invariance, and—when required—equality of composite means and variances. Path-coefficient comparisons should follow rather than precede the required MICOM stages (Hair et al., 2017; Avkiran & Ringle, 2018).
-
Interpret the group difference within the established model. Interpret any detected group difference as a parameter difference under a specified and sufficiently invariant measurement model—not as proof of why the groups differ and not as a causal conclusion unless the research design independently supports causation.
Limitations and cautions
Measurement invariance is model-dependent. Establishing invariance under one factor specification does not prove that every alternative measurement model is equivalent across populations. Poorly specified baseline models can therefore undermine the meaning of later equality constraints (Wang & Wang, 2012).
Invariance is comparison-specific. Evidence obtained for two groups does not automatically establish equivalence for additional groups, different populations, or different measurement occasions. Multigroup modeling evaluates the populations represented in the specified analysis (Wang & Wang, 2012).
Full invariance should not become an automatic badge of measurement quality. Wang and Wang (2012) note that strict residual invariance is not required for many substantive applications, while partial CFA invariance strategies require caution. The appropriate requirement depends on the parameter the researcher intends to compare.
PLS-SEM introduces its own limits. MICOM addresses measurement invariance for composite models and should not be represented as an interchangeable implementation of multigroup CFA. Likewise, evidence of full MICOM invariance does not imply that all structural coefficients are equal; structural heterogeneity remains a separate empirical question (Hair et al., 2017; Avkiran & Ringle, 2018).
Conclusion
The question is not simply:
“Do the groups have different scores?”
It is:
“Have we established enough measurement equivalence for this particular group comparison to have a stable interpretation?”
For measurement invariance CFA, configural invariance establishes a common factor pattern, metric invariance establishes equivalent factor loadings, and strong/scalar invariance adds equal indicator intercepts so that latent means can be meaningfully compared. Strict invariance adds equality of residual variances when that stronger measurement claim is required (Wang & Wang, 2012).
For PLS-SEM, MICOM asks a different sequence of questions: configural invariance, compositional invariance, and equality of composite means and variances. Configural plus compositional invariance permits comparison of structural path coefficients; full MICOM invariance additionally supports pooling of the groups under the conditions described by the procedure (Hair et al., 2017).
Do not interpret a group difference until you have shown that the measurement system is sufficiently comparable for the difference you want to interpret.
That is what measurement invariance testing contributes: not another preliminary checkbox, but the evidential bridge between “these estimates are different” and “these groups differ on the construct or relationship we intended to study.”
References
Avkiran, N. K., & Ringle, C. M. (Eds.). (2018). Partial least squares structural equation modeling: Recent advances in banking and finance. Springer. https://doi.org/10.1007/978-3-319-71691-6
Hair, J. F., Jr., Hult, G. T. M., Ringle, C. M., & Sarstedt, M. (2017). A primer on partial least squares structural equation modeling (PLS-SEM) (2nd ed.). SAGE Publications.
Tabachnick, B. G., & Fidell, L. S. (2013). Using multivariate statistics (6th ed.). Pearson.
Wang, J., & Wang, X. (2012). Structural equation modeling: Applications using Mplus. John Wiley & Sons.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.