Resource

Reflective or Formative? A PLS-SEM Measurement Model Case Study

Reflective and formative constructs represent different relationships between indicators and constructs. This synthetic PLS-SEM case study shows how construct conceptualization, indicator direction, interchangeability, collinearity, outer weights, loadings, reliability, and validity criteria lead to different measurement decisions.

A PLS-SEM measurement model should not be classified as reflective simply because the software makes reflective measurement easy to specify. The measurement direction has to follow the conceptualization of the construct: are the indicators manifestations of an underlying concept, or are they components that jointly form the concept? Hair et al. (2017) emphasize that constructs are not inherently reflective or formative; the appropriate specification depends on how the construct is conceptualized and on the study objective.

This synthetic case study shows why that distinction matters. A researcher initially specifies every construct reflectively, then discovers that one construct—Digital Work Enablement—is better understood as a formative composite. The correction changes both the measurement criteria and the interpretation of its indicators.

Synthetic-case disclosure: The organization, survey, observations, and numerical results below are entirely fictional and are included only to illustrate measurement-model decisions. They are not empirical findings from any cited source.

The research problem

Northstar Services, a fictional multi-site professional-services organization, wants to understand why some business units adapt their work processes more effectively after a major digital-transformation program.

The research team surveys 420 employees and proposes a PLS-SEM model containing four constructs:

Construct Intended meaning Indicators
Manager Support Employees' perception that their manager provides encouragement, assistance, and backing MS1–MS4
Digital Work Enablement The bundle of organizational conditions that make digitally supported work possible DWE1–DWE5
Change Confidence Employees' confidence in coping successfully with organizational change CC1–CC4
Operational Agility Perceived ability of the unit to respond and adjust rapidly OA1–OA4

The structural model proposes that Manager Support and Digital Work Enablement are associated with Change Confidence, which in turn is associated with Operational Agility.

The initial mistake occurs before any structural relationship is interpreted: all four constructs are entered as reflective because that is the researcher's default model setup.

That apparently minor choice changes what the model assumes about Digital Work Enablement.

Start with construct conceptualization, not software settings

Reflective and formative measurement represent different relationships between constructs and their indicators.

Reflective measurement

Indicators are treated as manifestations or effects of the underlying construct. The causal direction in the measurement model therefore runs from the construct to its indicators.

Because the indicators originate from the same conceptual domain, they are expected to be associated, and individual reflective indicators are generally interchangeable provided that the construct remains measured with sufficient reliability (Hair et al., 2017).

Formative measurement

Indicators jointly form the construct, so the measurement relationships run from the indicators toward the construct.

The indicators capture potentially different aspects of the construct's domain, need not be interchangeable, and removing one can alter what the construct represents (Hair et al., 2017).

Avkiran and Ringle (2018) similarly describe reflective indicators as consequences or manifestations of an underlying construct and formative indicators as complementary sources used to form a composite.

That distinction immediately raises different questions for the two focal constructs in the Northstar study.

Why Manager Support is plausibly reflective

Consider the four synthetic Manager Support items:

  • MS1: My manager supports me when work problems arise.
  • MS2: My manager gives me the assistance I need to perform effectively.
  • MS3: My manager provides encouragement when work becomes difficult.
  • MS4: I can rely on my manager for work-related support.

The researcher's conceptual definition treats Manager Support as an underlying employee perception expressed through these responses.

If an employee's underlying perception of managerial support became more favorable, responses to all four indicators would generally be expected to move in the same direction. The indicators deliberately overlap in content and are intended to provide alternative manifestations of the same underlying concept.

That reasoning fits reflective measurement: the construct gives rise to the observed responses, and the indicators are expected to share substantial commonality (Hair et al., 2017; Avkiran & Ringle, 2018).

The interchangeability test

Conceptual question: Would removing one indicator substantially change what the construct means?

For Manager Support, removing MS3 would reduce the amount of information available, but the remaining items would still represent the same general construct. This is consistent with the interchangeability expected of reflective indicators (Hair et al., 2017).

Interchangeability does not mean that item deletion is statistically or substantively harmless in every application. It means that reflective indicators are conceived as alternative manifestations sampled from the same conceptual domain rather than as unique components that collectively define that domain (Hair et al., 2017).

Why Digital Work Enablement is plausibly formative

Now consider the Digital Work Enablement indicators:

  • DWE1: Employees have reliable access to appropriate digital hardware.
  • DWE2: Core software systems are integrated sufficiently for routine work.
  • DWE3: Employees receive adequate training in the digital tools they use.
  • DWE4: Technical support is available when digital problems occur.
  • DWE5: Employees have appropriate access permissions to the data and systems required for their roles.

These indicators do not look like interchangeable manifestations of one underlying psychological state.

Instead, the researcher defines Digital Work Enablement as a composite of distinct organizational conditions. Hardware, integration, training, technical support, and system access contribute different content to the construct.

Under this conceptualization, the indicators jointly determine what Digital Work Enablement means. Formative indicators are intended to cover different aspects of a construct's domain, and omitting an indicator can therefore change the nature of the construct (Hair et al., 2017).

For example, a unit might have excellent hardware but poor training. Another might have extensive training but weak systems integration. There is no measurement requirement that every component move together. Formative indicators need not exhibit the strong intercorrelations expected under reflective measurement (Hair et al., 2017).

The direction test

Manager Support

Direction:

Manager Support → MS1, MS2, MS3, MS4

The construct is conceptualized as giving rise to its indicators.

Digital Work Enablement

Direction:

DWE1, DWE2, DWE3, DWE4, DWE5 → Digital Work Enablement

The indicators are conceptualized as jointly forming the construct.

This direction should follow the construct definition rather than a software default. Hair et al. (2017) explicitly argue that the specification of construct content and the objective of the research should guide the measurement perspective.

What goes wrong when Digital Work Enablement is forced to be reflective?

Suppose the researcher leaves every construct in reflective mode and obtains these illustrative synthetic results:

Digital Work Enablement indicator Reflective outer loading
DWE1 Hardware .73
DWE2 Systems integration .68
DWE3 Training .35
DWE4 Technical support .42
DWE5 Access permissions .31

A threshold-driven reflective analysis could tempt the researcher to remove DWE3 and DWE5 and perhaps reconsider DWE4.

But that would solve the wrong measurement problem.

Hair et al. (2017) warn specifically that specifying a construct reflectively when its conceptualization calls for formative measurement can produce lower reflective loadings because formative indicators are not necessarily highly correlated. Applying reflective item-purification rules can then delete indicators that should have been retained, damaging the content validity of the construct. Measurement-model misspecification can consequently alter both measurement and structural-model results (Hair et al., 2017).

In the Northstar example, deleting training and access permissions would redefine Digital Work Enablement as something closer to technology infrastructure availability. The resulting construct would no longer represent the broader concept originally proposed.

The problem is therefore not simply that some loadings are low.

The problem is that loadings are being used to answer a question that the construct conceptualization does not ask.

Correcting the measurement specification

After revisiting the construct definitions, the researcher specifies:

Manager Support

Specification: Reflective

Digital Work Enablement

Specification: Formative

Change Confidence

Specification: Reflective

Operational Agility

Specification: Reflective

The researcher then evaluates the reflective and formative portions separately.

That separation is essential because the criteria appropriate for reflective measurement cannot simply be transferred to formative measurement. Formative indicators need not be strongly correlated, and Hair et al. (2017) warn that applying internal-consistency logic to them can remove important indicators and reduce the validity of the index.

Avkiran and Ringle (2018) likewise state that composite reliability and AVE are criteria for reflective measurement rather than formative measurement models.

Evaluating the reflective constructs

Reflective measurement evaluation focuses on whether the indicators perform as manifestations of their intended constructs.

Hair et al. (2021) organize reflective measurement assessment around indicator reliability, internal consistency reliability, convergent validity, and discriminant validity. Avkiran and Ringle (2018) similarly discuss outer loadings, composite reliability, AVE, and discriminant validity for reflective models.

1. Examine outer loadings

For the synthetic Manager Support construct, suppose the estimated standardized loadings are:

Indicator Synthetic loading
MS1 .82
MS2 .87
MS3 .78
MS4 .84

Reflective outer loadings show how strongly individual indicators are associated with the construct. Hair et al. (2021) recommend examining indicator reliability before moving to construct-level reliability and validity criteria. Very weak reflective loadings warrant particular attention, but indicators with intermediate loadings should not automatically be removed without considering their effect on other measurement criteria and the content represented by the item (Hair et al., 2021).

2. Assess internal consistency reliability

Suppose the synthetic Manager Support results are:

Cronbach's alpha

.85

Composite reliability

.90

Internal consistency reliability assesses the degree to which indicators intended to measure the same reflective construct are associated with one another. Composite reliability is a principal PLS-SEM reliability measure, while Cronbach's alpha provides another internal-consistency assessment (Hair et al., 2021).

Extremely high reliability is not automatically desirable. Hair et al. (2021) note that very high reliability values can indicate indicator redundancy, meaning that multiple items may be measuring essentially the same content rather than providing useful breadth.

3. Assess convergent validity

Suppose Manager Support has a synthetic AVE of .69.

Average variance extracted (AVE) summarizes the amount of indicator variance explained by the reflective construct. An AVE of .50 or more indicates that the construct accounts for at least half of the variance of its indicators (Hair et al., 2021; Avkiran & Ringle, 2018).

4. Assess discriminant validity

The researcher must also examine whether each reflective construct is empirically distinct from the other reflective constructs.

Hair et al. (2021) recommend the heterotrait–monotrait ratio of correlations (HTMT) for discriminant-validity assessment and note shortcomings of relying on the traditional Fornell–Larcker criterion alone. The interpretation of HTMT should take the conceptual similarity of the constructs into account, and bootstrap confidence intervals can provide additional inferential evidence (Hair et al., 2021).

The reflective measurement assessment therefore asks whether indicators behave as expected when they are manifestations of a common construct.

Evaluating Digital Work Enablement as formative

Once Digital Work Enablement is correctly specified as formative, the evaluation changes.

Hair et al. (2021) organize formative measurement assessment into three central steps: convergent validity, indicator collinearity, and the statistical significance and relevance of indicator weights. Hair et al. (2017) provide the same general sequence, while Avkiran and Ringle (2018) emphasize theoretical content validity before empirical formative assessment.

1. Protect content validity first

Before interpreting numerical output, the researcher asks whether hardware, integration, training, technical support, and system access adequately represent the intended domain of Digital Work Enablement.

This question is particularly important for formative measurement because the indicators collectively determine the construct's content. Removing a formative indicator can change what is being measured (Hair et al., 2017).

Interpretation boundary: An inconvenient coefficient is not by itself sufficient reason to delete a formative indicator.

2. Assess convergent validity with redundancy analysis

A formative construct can be compared with an alternative reflective or global measure of the same concept through redundancy analysis (Hair et al., 2017, 2021).

The Northstar questionnaire therefore includes a global item that was deliberately kept outside the formative block:

GLOBAL-DWE: Overall, our unit provides the digital conditions employees need to work effectively.

Suppose the synthetic path from formative Digital Work Enablement to this global criterion is .76.

Hair et al. (2017) describe a value of at least .70 as a minimum desirable level in this redundancy-analysis setting, while Hair et al. (2021) use .708 as the corresponding benchmark indicating that the formative construct explains about half of the alternative measure's variance.

The synthetic .76 result would therefore provide supportive evidence that the five components jointly capture the intended overall concept.

Check collinearity among formative indicators

Collinearity has a different role in formative measurement than correlation has in reflective measurement.

Formative indicators do not have to correlate strongly. Indeed, excessive collinearity can make their unique contributions difficult to distinguish. High formative-indicator collinearity increases the standard errors of indicator weights and can even produce unexpected sign changes (Hair et al., 2021).

The standard diagnostic is the variance inflation factor (VIF). Hair et al. (2021) identify VIF values of 5 or more as indicative of critical collinearity concerns while noting that problems can sometimes arise at lower values. Avkiran and Ringle (2018) similarly recommend VIF values below 5 and identify values below 3 as preferable for interpreting formative outer weights.

Suppose the synthetic results are:

Indicator Synthetic VIF
DWE1 Hardware 1.72
DWE2 Systems integration 2.31
DWE3 Training 1.44
DWE4 Technical support 2.08
DWE5 Access permissions 1.63

These results would not indicate a critical formative-indicator collinearity problem under the guidance above.

That conclusion is very different from asking whether the indicators correlate strongly enough to produce a high Cronbach's alpha. For formative indicators, strong internal consistency is not the measurement objective.

Interpret formative weights and loadings differently

The formative outer weight represents an indicator's relative contribution to forming the construct, conditional on the other formative indicators. The corresponding loading provides information about the indicator's absolute contribution (Hair et al., 2017; Avkiran & Ringle, 2018).

Suppose bootstrapping produces the following entirely synthetic results:

Indicator Outer weight 95% bootstrap CI Outer loading Initial interpretation
DWE1 Hardware .29 [.13, .44] .67 Significant relative contribution
DWE2 Integration .41 [.25, .55] .76 Significant relative contribution
DWE3 Training .18 [.02, .34] .58 Significant relative contribution
DWE4 Technical support .24 [.07, .39] .63 Significant relative contribution
DWE5 Access permissions .12 [-.03, .27] .55 Weight nonsignificant; retain for further substantive review

Bootstrapping is used to assess whether formative indicator weights differ significantly from zero (Hair et al., 2017, 2021). A significant weight provides evidence that the indicator makes a unique relative contribution to the composite.

But DWE5 illustrates an important decision point.

Its synthetic weight is not statistically significant because its confidence interval includes zero. That does not automatically establish that access permissions are irrelevant.

Hair et al. (2017, 2021) recommend examining the formative indicator's loading when its weight is nonsignificant. A sufficiently strong and statistically significant loading can support retaining an indicator even when its unique relative weight is nonsignificant. The decision should also respect the content that the indicator contributes to the formative construct (Hair et al., 2017).

In this case, removing access permissions would eliminate a distinct component of the original definition of Digital Work Enablement. With no serious collinearity problem and a synthetic loading of .55, the researcher retains DWE5 rather than deleting it solely because its outer weight is nonsignificant.

Why Cronbach's alpha cannot decide whether Digital Work Enablement is adequate

Suppose the five Digital Work Enablement components produce a synthetic Cronbach's alpha of only .58.

If the researcher still believed the construct was reflective, that result could be interpreted as an internal-consistency concern.

For the correctly conceptualized formative composite, however, the same conclusion does not follow.

Reflective internal-consistency criteria depend on the expectation that indicators are manifestations of a common construct and should therefore share variance. Formative indicators represent complementary components and are not required to exhibit that same correlation pattern (Hair et al., 2017).

Hair et al. (2017) explicitly warn that applying correlation-based reliability analysis to formative indicators can encourage researchers to remove important components and thereby reduce content validity. Avkiran and Ringle (2018) similarly restrict composite reliability and AVE to reflective measurement-model evaluation.

The wrong question for this formative composite

“Are these five items internally consistent?”

The relevant formative question

“Do these five components adequately form the intended construct, without problematic collinearity, and what evidence supports their contributions?”

Those are different measurement questions requiring different diagnostics.

The revised measurement decision

The researcher's final measurement-model reasoning can be summarized as follows:

Decision question Manager Support Digital Work Enablement
What does the construct represent? Underlying employee perception Composite of organizational conditions
Direction Construct → indicators Indicators → construct
Indicators expected to overlap? Yes Not necessarily
Indicators conceptually interchangeable? Largely yes No
Does removing an indicator potentially redefine the construct? Usually less directly Yes
Primary coefficients interpreted Outer loadings Outer weights, with loadings also considered
Internal consistency relevant? Yes Not as a formative validation criterion
AVE relevant? Yes Not as a formative validation criterion
Formative-indicator VIF required? No Yes
Weight significance required for every indicator? Not the reflective decision criterion Important, but nonsignificance does not automatically imply deletion
Content coverage especially critical? Important Central

The key decision was made before looking at these diagnostics: the theoretical and conceptual meaning of each construct determined its measurement specification (Hair et al., 2017).

What the original misspecification would have changed

Leaving Digital Work Enablement reflective would have produced several avoidable problems.

Wrong direction

The model would have imposed the wrong indicator–construct direction.

Wrong indicator logic

It would have expected complementary organizational conditions to behave like interchangeable manifestations.

Wrong diagnosis

Low intercorrelations could have been misdiagnosed as measurement failure.

Content loss

Reflective item-purification procedures could have removed substantively important components.

Hair et al. (2017) specifically caution that this type of misspecification can damage content validity and materially alter measurement and structural results.

Correcting the model does not mean that every formative indicator must be retained regardless of evidence. Formative models still require empirical scrutiny. Researchers should examine content coverage, convergent validity, collinearity, weights, their statistical significance, and the absolute contribution represented by loadings (Hair et al., 2017, 2021; Avkiran & Ringle, 2018).

What changes is the logic used to make those decisions.

A practical reflective-versus-formative decision sequence

  1. Define the construct before drawing arrows. State precisely what the construct is intended to represent.

  2. Determine the conceptual direction. Ask whether the construct produces the observed indicators or whether the indicators combine to form the construct.

  3. Consider interchangeability. If indicators are alternative manifestations from the same domain, reflective measurement may be plausible. If they capture distinct components whose omission changes the construct, formative measurement may be more appropriate (Hair et al., 2017).

  4. Specify the model from theory rather than software convenience. Empirical information can supplement this decision, but the primary basis is theoretical reasoning and construct conceptualization (Hair et al., 2017).

  5. Apply measurement-type-specific evaluation criteria. Reflective models require assessment of indicator reliability, internal consistency, convergent validity, and discriminant validity. Formative models require attention to content coverage, convergent validity, indicator collinearity, and indicator-weight significance and relevance (Hair et al., 2017, 2021).

  6. Do not delete formative indicators mechanically. A nonsignificant weight should prompt examination of the indicator's loading and substantive contribution before removal (Hair et al., 2017, 2021).

  7. Only then interpret the structural model. A structural coefficient cannot repair a poorly conceptualized measurement model.

Decision workflow: Define the construct → determine direction → assess interchangeability → specify from theory → apply measurement-type-specific criteria → review indicator evidence → interpret the structural model.

Conclusion

The reflective-versus-formative decision is not a cosmetic PLS-SEM setting. It defines what the relationship between a construct and its observed indicators is assumed to mean.

In this synthetic case, Manager Support is plausibly reflective because its indicators are treated as correlated manifestations of an underlying employee perception. Digital Work Enablement is plausibly formative because hardware, integration, training, technical support, and access permissions are distinct components that jointly determine the construct.

Modeling both reflectively would subject fundamentally different measurement structures to the same reliability and validity criteria. That can lead to inappropriate indicator deletion, loss of construct content, and changed model results (Hair et al., 2017).

The practical rule is simple but demanding: conceptualize first, specify second, evaluate third. Reflective and formative measurement models answer different measurement questions, so they should not be validated with an interchangeable checklist.

References

Avkiran, N. K., & Ringle, C. M. (Eds.). (2018). Partial least squares structural equation modeling: Recent advances in banking and finance. Springer. https://doi.org/10.1007/978-3-319-71691-6

Hair, J. F., Jr., Hult, G. T. M., Ringle, C. M., & Sarstedt, M. (2017). A primer on partial least squares structural equation modeling (PLS-SEM) (2nd ed.). SAGE Publications.

Hair, J. F., Jr., Hult, G. T. M., Ringle, C. M., Sarstedt, M., Danks, N. P., & Ray, S. (2021). Partial least squares structural equation modeling (PLS-SEM) using R: A workbook. Springer. https://doi.org/10.1007/978-3-030-80519-7

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry