Resource

Strong Path Coefficients, Weak Measurement: A PLS-SEM Model Evaluation Case Study

Strong PLS-SEM path coefficients do not rescue weak measurement. This synthetic case study follows the model-assessment sequence from specification and data examination through reflective measurement quality, validity, collinearity, bootstrapping, structural interpretation, and prediction.

A large path coefficient can be persuasive. It can also be premature.

In partial least squares structural equation modeling (PLS-SEM), structural relationships are interpreted only after the measurement models have been evaluated. Hair et al. (2017) explicitly sequence structural-model assessment after establishing that the construct measures are sufficiently reliable and valid. Hair et al. (2021) likewise organize PLS-SEM evaluation from measurement assessment to structural-model collinearity, path relationships, explanatory power, and predictive power.

This synthetic case study shows why that sequence matters. An initial model appears impressive: satisfaction strongly predicts loyalty intention, service quality strongly predicts satisfaction, and the endogenous constructs have substantial-looking explained variance. But closer inspection reveals that the satisfaction construct—the construct carrying some of the strongest structural relationships—has weak convergent measurement and two problematic indicators.

Synthetic-data disclosure: The numbers below are synthetic illustrative values constructed for this case study. They are not empirical findings and are not SmartPLS or SEMinR output.

The research problem

A fictional subscription-service company wants to understand how customers move from service evaluations to loyalty intentions. The proposed behavioral model contains five reflectively measured constructs:

  • Service Quality (SQ): perceived reliability, responsiveness, competence, and consistency of service.
  • Perceived Value (PV): customers' evaluation of what they receive relative to what they give up.
  • Trust (TR): confidence in the provider's dependability and integrity.
  • Satisfaction (SAT): the customer's overall favorable evaluation of the service experience.
  • Loyalty Intention (LOY): intention to continue, recommend, and prefer the provider.

The synthetic study contains 420 survey responses. The purpose is explanatory and predictive rather than causal: the cross-sectional design does not establish temporal ordering or experimental identification.

The substantive question is:

Do service quality, perceived value, and trust explain satisfaction, and does satisfaction in turn explain loyalty intention?

A tempting workflow would estimate the model, inspect the path coefficients, and begin discussing the strongest relationships. PLS-SEM evaluation requires a different order. Measurement quality must first be established because structural relationships connect construct scores whose meaning depends on how adequately their indicators represent the intended constructs (Hair et al., 2017, 2021).

1. Model specification comes before model results

The structural model was specified as:

  • SQ → SAT
  • PV → SAT
  • TR → SAT
  • SAT → LOY
  • TR → LOY

All five constructs were specified reflectively. This means the indicators are treated as manifestations of their respective constructs, so reflective measurement-model criteria—not formative criteria—are appropriate for evaluating their measurement quality (Hair et al., 2017, 2021).

The distinction matters. Reflective measurement assessment considers indicator reliability, internal-consistency reliability, convergent validity, and discriminant validity. Formative measurement requires a different assessment logic involving convergent validity, indicator collinearity, and the significance and relevance of indicator weights (Hair et al., 2017, 2021).

Reflective measurement

Indicators are treated as manifestations of their construct.

Relevant evidence: indicator reliability, internal-consistency reliability, convergent validity, and discriminant validity.

Formative measurement

A different measurement-assessment logic applies.

Relevant evidence: convergent validity, indicator collinearity, and the significance and relevance of indicator weights.

No formative construct is included in this case because the central teaching problem is clearer when the apparent structural success depends on a weak reflective construct.

Synthetic measurement specification

Synthetic measurement specification
Construct Indicators Measurement
Service Quality SQ1–SQ4 Reflective
Perceived Value PV1–PV4 Reflective
Trust TR1–TR4 Reflective
Satisfaction SAT1–SAT4 Reflective
Loyalty Intention LOY1–LOY3 Reflective

Specification is not a cosmetic software step. The type of measurement model determines which diagnostics are meaningful and therefore which evidence can justify interpreting the construct.

2. Examine the data before evaluating the model

PLS-SEM does not remove the need to understand the dataset. Hair et al. (2017) place data examination alongside path-model specification before result assessment, while Hair et al. (2021) discuss missing values, non-normal data, scale characteristics, and other data considerations when applying PLS-SEM.

For the synthetic dataset, assume that screening produced the following deliberately uncomplicated situation:

Synthetic data-screening findings
Data issue Synthetic finding Decision
Sample size 420 complete survey records after screening Proceed
Indicator ranges All observations within the intended response range Proceed
Missingness No missing values in the analysis dataset No imputation required
Coding No incorrect reversals or impossible values detected Proceed
Unusual response patterns No deliberately inserted data-entry anomalies Proceed

These statements describe the constructed case, not general PLS-SEM requirements.

The purpose of this stage is simple: a measurement anomaly should not be interpreted as a substantive psychometric problem until basic coding and data problems have been considered.

3. The structural model initially looks excellent

Suppose an analyst jumps immediately to the structural results and sees this:

Initial synthetic structural path coefficients
Structural path Synthetic β
SQ → SAT .51
PV → SAT .36
TR → SAT .18
SAT → LOY .69
TR → LOY .17

The corresponding synthetic coefficients of determination are:

Satisfaction

R² = .71

Loyalty Intention

R² = .66

At first glance, this is an appealing story. Service quality and perceived value are strongly associated with satisfaction, and satisfaction has a particularly strong positive relationship with loyalty intention.

That is precisely where the analysis should not begin.

Hair et al. (2017) state that structural-model assessment follows confirmation that the construct measures are reliable and valid. The structural assessment then proceeds through collinearity, path coefficients, the coefficient of determination, effect sizes, and predictive-relevance assessment. Hair et al. (2021) similarly separate measurement-model evaluation from assessment of structural collinearity, structural relationships, explanatory power, and out-of-sample predictive power.

The question is therefore not yet, “How important is satisfaction?”

Have we measured satisfaction well enough to interpret its structural relationships?

4. Reflective measurement-model assessment changes the story

Indicator reliability

For reflective constructs, high outer loadings indicate that the indicators share substantial variance with the construct. Hair et al. (2017) give .708 as a commonly used target because squaring .708 yields approximately .50, meaning that the construct explains about half of an indicator's variance. Hair et al. (2021) use the same rationale.

Importantly, a loading below .708 is not an automatic deletion instruction. Indicators between approximately .40 and .70 should be considered in relation to their effect on reliability, convergent validity, and content validity rather than mechanically removed (Hair et al., 2017, 2021).

The synthetic loadings are:

Synthetic reflective indicator loadings
Construct Indicator Loading Initial assessment
Service Quality SQ1 .86 Strong
SQ2 .83 Strong
SQ3 .81 Strong
SQ4 .79 Strong
Perceived Value PV1 .84 Strong
PV2 .82 Strong
PV3 .80 Strong
PV4 .78 Strong
Trust TR1 .88 Strong
TR2 .85 Strong
TR3 .82 Strong
TR4 .80 Strong
Satisfaction SAT1 .82 Strong
SAT2 .79 Strong
SAT3 .55 Requires review
SAT4 .48 Requires review
Loyalty Intention LOY1 .90 Strong
LOY2 .88 Strong
LOY3 .84 Strong

Now the structural story is less secure.

SAT3 and SAT4 do not justify automatic deletion, but they require investigation. Removing indicators solely to improve statistics can damage content validity, so their wording and conceptual role would need to be examined alongside their statistical performance (Hair et al., 2017, 2021).

Why this matters

SAT is not a peripheral construct. It is the principal endogenous construct in one equation and the strongest predictor of LOY in another.

The model's most impressive path therefore depends directly on one of its weakest measurement blocks.

5. Reliability is not enough: convergent validity fails

Composite reliability accounts for the indicators' different loadings and is a central internal-consistency measure in reflective PLS-SEM assessment. Hair et al. (2017) describe values between .70 and .90 as satisfactory in more advanced research contexts, while also warning that very high values can indicate indicator redundancy. Hair et al. (2021) similarly distinguish satisfactory reliability from excessively high reliability.

Average variance extracted (AVE) addresses convergent validity. An AVE of .50 means that the construct explains, on average, half of the variance in its indicators; values below .50 indicate that more indicator variance remains unexplained than is captured by the construct on average (Hair et al., 2017, 2021).

For this synthetic example:

Synthetic reliability and convergent-validity results
Construct Illustrative composite reliability AVE
Service Quality .893 .677
Perceived Value .884 .657
Trust .904 .702
Satisfaction .763 .457
Loyalty Intention .906 .763

The key result is not that Satisfaction's composite reliability is above .70. It is that Satisfaction has AVE = .457.

Its four loadings were constructed so that the AVE equals approximately:

(.82² + .79² + .55² + .48²) / 4 = .457

Thus, the construct does not meet the .50 convergent-validity criterion described by Hair et al. (2017, 2021).

This demonstrates an important reporting principle: acceptable internal consistency cannot be used to conceal weak convergent validity. Reliability and validity answer different questions.

The strong SAT → LOY = .69 coefficient should therefore remain substantively provisional.

6. Discriminant validity creates a second warning

Reflective measurement assessment must also ask whether empirically distinct constructs are sufficiently distinguishable. Hair et al. (2017) recommend the heterotrait-monotrait ratio of correlations (HTMT) for this purpose and note that the appropriate threshold depends partly on conceptual similarity. They discuss .90 for conceptually similar constructs and a more conservative .85 criterion for more distinct constructs. Hair et al. (2021) likewise treat discriminant validity as a separate part of reflective measurement assessment.

Assume the synthetic HTMT matrix contains the following highest values:

Selected synthetic HTMT values
Construct pair Synthetic HTMT
SQ–PV .72
SQ–TR .68
SQ–SAT .76
PV–TR .74
PV–SAT .92
TR–SAT .79
SAT–LOY .81
TR–LOY .75

The PV–SAT HTMT of .92 is deliberately problematic.

Under the guidance in Hair et al. (2017), a value above .90 raises evidence of insufficient discriminant validity even when constructs are conceptually similar.

The problem is consequential. The model reports both PV → SAT = .36 and SAT → LOY = .69, yet the distinction between perceived value and satisfaction is empirically questionable in this synthetic measurement system.

A path diagram cannot resolve that problem.

The measurement decision

The analyst should return to construct definitions and indicator content before interpreting the paths.

Possible explanations include:

  • SAT3 or SAT4 may not represent satisfaction adequately.
  • One or both weak satisfaction indicators may overlap conceptually with perceived value.
  • The conceptual distinction between PV and SAT may not have been translated adequately into the instrument.
  • The specified measurement model itself may require reconsideration.

These are diagnostic possibilities, not findings from the synthetic numbers alone.

The correct response is therefore investigation, not automatic deletion until HTMT falls below a desired value.

7. Collinearity must be checked before interpreting structural coefficients

Suppose the analyst resolves the measurement problem on substantive grounds and estimates a defensible revised measurement model. Structural assessment can then begin.

Hair et al. (2017) place structural collinearity first because PLS-SEM structural paths are based on regressions of endogenous constructs on their predictor constructs, so problematic predictor collinearity can distort coefficient estimation. Hair et al. (2021) similarly recommend examining VIF values for each set of predictor constructs and note that values above 5 indicate probable collinearity problems, while collinearity can also become relevant at values between 3 and 5.

Assume the predictors of SAT produce these synthetic inner VIFs:

Synthetic inner VIF values for predictors of Satisfaction
Predictor of Satisfaction Synthetic inner VIF
Service Quality 2.7
Perceived Value 3.6
Trust 2.4

The VIF of 3.6 for PV is not treated as an automatic rejection. It does, however, deserve attention under Hair et al.'s (2021) guidance because the 3–5 range can already signal collinearity concerns.

This reinforces the earlier measurement finding: PV and SAT are not the only constructs with potential overlap in the model. The predictors themselves also require examination before their relative path magnitudes are interpreted as cleanly separable contributions.

8. Only now should the structural model be evaluated

Once satisfactory measurement has been demonstrated, structural assessment can examine the sign, magnitude, and inferential evidence for path coefficients, followed by explanatory and predictive criteria (Hair et al., 2017, 2021).

To keep the teaching example focused, assume the researcher revises the satisfaction measurement only after reviewing item content and determining that SAT4 is conceptually inconsistent with the intended satisfaction domain. SAT3 is retained because it represents an important facet that is not duplicated by the stronger items.

The model is then re-estimated.

The revised structural estimates cannot simply be assumed to equal the original estimates. Measurement changes alter the estimated model. The analyst must rerun and report the model rather than carry forward the earlier coefficients.

For illustration, suppose the revised synthetic structural estimates are:

Revised synthetic structural estimates
Path Revised β Synthetic 95% bootstrap CI Interpretation
SQ → SAT .48 [.39, .57] Positive association
PV → SAT .31 [.21, .41] Positive association
TR → SAT .15 [.06, .24] Smaller positive association
SAT → LOY .63 [.54, .71] Strong positive association
TR → LOY .14 [.06, .22] Smaller positive association

These values remain fairly strong, but they now belong to a different, revised model.

That distinction belongs in the report.

9. Bootstrapping supports inference; it does not repair measurement

PLS-SEM uses nonparametric bootstrapping to estimate uncertainty for parameters such as structural paths and measurement-model coefficients. Hair et al. (2017) recommend a large number of bootstrap samples and discuss 5,000 as a recommended rule in their treatment. Hair et al. (2021), drawing on subsequent methodological work, recommend at least 10,000 bootstrap samples for PLS-SEM applications.

For this synthetic case, assume a 10,000-sample bootstrap was used for the illustrative revised-model confidence intervals.

The values remain fictional and are not presented as output from SmartPLS, SEMinR, or another PLS-SEM implementation.

The critical point is conceptual:

A narrow bootstrap confidence interval around SAT → LOY does not establish that SAT was measured adequately.

Bootstrapping addresses sampling uncertainty around estimated parameters. Reflective measurement quality still has to be established through the relevant measurement evidence (Hair et al., 2017, 2021).

This is another reason not to start with the bootstrap path table.

10. Explanatory power is not the same as predictive power

R² describes in-sample explanatory power

The coefficient of determination indicates how much variance in an endogenous construct is explained by its predictors. Hair et al. (2017) caution that acceptable R² values depend on model complexity and research context; the same numerical value can be considered substantial in one research setting and modest in another. Avkiran and Ringle (2018) likewise emphasize contextual interpretation rather than treating R² categories as universal standards.

Suppose the revised synthetic model produces:

Satisfaction

R² = .67

Loyalty Intention

R² = .61

These values indicate considerable in-sample explained variance in this synthetic dataset. They do not demonstrate that the model will predict new observations well.

Predictive assessment requires additional evidence

Hair et al. (2021) distinguish the model's in-sample explanatory power from its out-of-sample predictive power and describe PLSpredict as a procedure for predictive assessment. Prediction-error measures such as RMSE and MAE can then be examined for the relevant endogenous construct.

This case does not manufacture PLSpredict results.

Accordingly, the defensible conclusion is:

The synthetic revised model shows substantial-looking explanatory performance within the constructed sample, but no claim of out-of-sample predictive performance is made because a PLSpredict analysis was not actually conducted.

That distinction is more informative than labeling a high R² as evidence that the model “predicts well.”

11. What would have happened if we interpreted the paths first?

The initial structural model encouraged a straightforward business narrative:

Better service quality and perceived value increase satisfaction, and satisfaction strongly increases loyalty.

The model did not justify that statement.

First, the data are cross-sectional and synthetic, so causal language such as “increase” is not supported by the design.

Second, and more importantly for this case, the satisfaction construct had AVE = .457, two comparatively weak indicators, and an HTMT of .92 with perceived value.

The initial structural estimates therefore connected constructs whose measurement evidence was not yet adequate.

The appropriate interpretation was:

The initial structural coefficients appear strong, but substantive interpretation is deferred because the satisfaction measurement model shows inadequate convergent validity and a discriminant-validity concern with perceived value.

That sentence captures the core lesson of the case.

12. A decision-first PLS-SEM assessment sequence

This case suggests a practical workflow consistent with the staged evaluation described by Hair et al. (2017, 2021):

1. Specification

Question: What does each construct mean, and is it reflective or formative?

Evidence: Theory and measurement specification.

2. Data examination

Question: Are coding, missingness, ranges, or unusual observations compromising the analysis?

Evidence: Data screening.

3. Measurement assessment

Question: Do indicators adequately represent their constructs?

Evidence: Reflective or formative criteria as appropriate.

4. Reliability and validity

Question: Is the measurement evidence adequate for the intended interpretation?

Evidence: Loadings, reliability, AVE, and discriminant validity for reflective models.

5. Collinearity

Question: Are predictors sufficiently distinguishable for stable structural estimation?

Evidence: Inner VIF.

6. Structural relationships

Question: What are the signs and magnitudes of the paths?

Evidence: Path coefficients.

7. Inference

Question: How uncertain are the estimated relationships?

Evidence: Bootstrap confidence intervals.

8. Explanation

Question: How much endogenous variance is explained?

Evidence: R² and related structural evidence.

9. Prediction

Question: Does the model predict observations outside its estimation sample?

Evidence: PLSpredict and prediction errors where conducted.

10. Reporting

Question: Can readers see the measurement problems, decisions, and resulting changes?

Evidence: Transparent model and decision reporting.

The sequence prevents an attractive path coefficient from becoming the starting point for an interpretation that the measurement model cannot support.

13. What if the construct were formative?

This case uses reflective constructs, but the logic cannot simply be transferred to formative measurement.

Hair et al. (2017, 2021) assess formative measurement through a different sequence involving convergent validity, collinearity among formative indicators, and the statistical significance and substantive relevance of indicator weights. Internal-consistency reliability and AVE are reflective-model criteria and should not be used as if they validate a formative composite.

Avkiran and Ringle (2018) likewise distinguish reflective from formative assessment and discuss VIF and bootstrapped outer weights when evaluating formative measurement.

A researcher should not ask whether a formative construct “passes Cronbach's alpha and AVE” and then move to its paths. The measurement specification determines the appropriate evidence.

14. Limitations of the case study

This case is intentionally synthetic. Every coefficient, loading, reliability statistic, AVE, HTMT value, VIF, R², and confidence interval was created for methodological illustration. None represents a real company, respondent sample, survey instrument, SmartPLS project, or empirical result.

The case also simplifies the broader PLS-SEM workflow. It focuses on a small reflective model and does not demonstrate formative redundancy analysis, higher-order constructs, mediation, moderation, measurement invariance, model comparison, or other extensions discussed in the PLS-SEM literature (Hair et al., 2017, 2021).

The revised measurement decision is also intentionally simplified. In real research, indicator deletion requires substantive consideration of construct definition and content validity rather than optimization of reliability or validity statistics alone (Hair et al., 2017, 2021).

Finally, the structural coefficients describe associations within a cross-sectional synthetic model. They should not be interpreted as causal effects.

15. How to report a case like this

A transparent report should preserve the chronology of the analytical decision.

A concise results narrative could read:

Measurement-model assessment preceded interpretation of the structural relationships. Most reflective indicators showed strong loadings, but two Satisfaction indicators were comparatively weak (SAT3 = .55; SAT4 = .48). Although Satisfaction's illustrative composite reliability was .763, its AVE was .457, below the .50 criterion for convergent validity (Hair et al., 2017, 2021). Discriminant validity also required attention because the illustrative HTMT between Perceived Value and Satisfaction was .92, exceeding the .90 value discussed for conceptually similar constructs by Hair et al. (2017). Structural-path interpretation was therefore deferred. After substantive review of indicator content, SAT4 was removed from the synthetic model and the complete model was re-estimated. The revised structural estimates, rather than the initial coefficients, were used for interpretation.

The report should then present:

  1. the prespecified measurement and structural models;
  2. data-examination decisions;
  3. complete measurement-model evidence;
  4. any measurement problems and the substantive reasoning behind revisions;
  5. structural collinearity;
  6. revised path estimates and bootstrap intervals;
  7. R² and any other supported explanatory measures;
  8. separate predictive assessment if actually performed; and
  9. limitations and sensitivity to model modification.

The reporting objective is not to make every cell in a diagnostic table appear satisfactory. It is to make the analytical reasoning visible.

Conclusion

The headline result in this synthetic study could easily have been SAT → LOY = .69.

It should not have been.

The more important initial finding was that Satisfaction—the construct carrying that strong relationship—had weak convergent measurement and questionable empirical separation from Perceived Value.

PLS-SEM model assessment is therefore not a search for the largest path coefficient. Reflective and formative measurement models require their own quality assessments, structural collinearity must be considered before relative path effects are interpreted, bootstrap inference does not compensate for weak measurement, and explanatory performance should not be confused with out-of-sample prediction (Hair et al., 2017, 2021; Avkiran & Ringle, 2018).

A strong structural model built on weak measurement is not strong evidence.

Defensible sequence: specification → data examination → measurement quality → collinearity → structural relationships → bootstrap inference → explanatory assessment → predictive assessment where performed → limitations → transparent reporting.

References

Avkiran, N. K., & Ringle, C. M. (Eds.). (2018). Partial least squares structural equation modeling: Recent advances in banking and finance. Springer. https://doi.org/10.1007/978-3-319-71691-6

Hair, J. F., Jr., Hult, G. T. M., Ringle, C. M., & Sarstedt, M. (2017). A primer on partial least squares structural equation modeling (PLS-SEM) (2nd ed.). SAGE Publications.

Hair, J. F., Jr., Hult, G. T. M., Ringle, C. M., Sarstedt, M., Danks, N. P., & Ray, S. (2021). Partial least squares structural equation modeling (PLS-SEM) using R: A workbook. Springer. https://doi.org/10.1007/978-3-030-80519-7

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry