Resource

From 18 Survey Items to a Defensible Factor Structure: An EFA Case Study

An EFA should do more than produce a rotated loading matrix. This synthetic 18-item case study shows how construct definitions, correlations, factor retention, rotation, cross-loadings, and item content combine to produce a provisional factor structure—and why that structure still requires independent confirmation.

A questionnaire can begin with a plausible conceptual structure and still produce an uncertain empirical one. That uncertainty is exactly where exploratory factor analysis (EFA) is useful: the question is not simply whether items correlate, but what latent dimensions could plausibly account for their pattern of relationships. Tabachnick and Fidell (2013) distinguish this exploratory purpose from confirmatory factor analysis (CFA), where a proposed factor structure is tested against the data.

Synthetic-data disclosure: This case study follows a fictional 18-item organizational questionnaire from construct definition through exploratory item revision. All respondents, item responses, correlations, factor loadings, and numerical results are synthetic. They illustrate an analysis workflow; they are not empirical findings about employees or organizations.

The measurement problem

Suppose an organizational research team wants to measure how employees experience day-to-day participation in changing work environments. Interviews and prior conceptual work suggest three related domains:

Voice Safety

Feeling able to raise concerns, ask difficult questions, and disagree without expecting interpersonal penalty.

Role Clarity

Understanding priorities, responsibilities, decision authority, and what successful performance requires.

Adaptive Energy

Having the psychological energy and willingness to adjust effort and approach when work changes.

These definitions come before the factor analysis. Measurement requires translating an underlying construct into observable indicators, while construct validity concerns whether a measure adequately represents the hypothetical construct it is intended to measure (Adams & Lawrence, 2018). Content coverage also matters: item development should represent the relevant aspects of the construct rather than allowing a statistical solution alone to define what the construct means (Adams & Lawrence, 2018).

Factor analysis can then examine which questionnaire items behave as related sets and help identify meaningful subscales, but the researcher still has to interpret and name those dimensions substantively (Adams & Lawrence, 2018). Sekaran and Bougie (2016) similarly describe factor analysis as a tool for examining dimensions of an operationalized concept and identifying which items are most appropriate for those dimensions.

The initial 18-item questionnaire

Respondents answer each statement on a seven-category agreement scale.

Item Synthetic questionnaire statement Intended domain
I01 I can raise concerns about how our work is being done. Voice Safety
I02 I can question a team decision without worrying about how I will be treated. Voice Safety
I03 I feel able to admit when I need help. Voice Safety
I04 I can point out a potential mistake even when others disagree. Voice Safety
I05 Difficult work conversations can be handled openly in my team. Voice Safety
I06 I know when I should speak up rather than handle an issue myself. Voice Safety / Role Clarity
I07 I understand what is expected of me in my role. Role Clarity
I08 My main work priorities are clear. Role Clarity
I09 I know which decisions I am responsible for making. Role Clarity
I10 I understand how my work will be evaluated. Role Clarity
I11 When priorities compete, I know which should come first. Role Clarity
I12 When work changes, I know what I should focus my effort on. Role Clarity / Adaptive Energy
I13 I can stay engaged when work requires a new approach. Adaptive Energy
I14 I have enough energy to adjust when priorities change. Adaptive Energy
I15 I can regain momentum after an unexpected setback. Adaptive Energy
I16 I am willing to change my usual approach when circumstances require it. Adaptive Energy
I17 I have the confidence to suggest a different approach when the current one is not working. Adaptive Energy / Voice Safety
I18 I usually cope with whatever happens at work. Adaptive Energy

Three items are deliberately ambiguous. I06 mixes speaking up with knowing one's decision boundary; I12 mixes clarity with adaptation; and I17 combines adaptive behavior with interpersonal voice. I18 is deliberately broad.

Measurement point: That ambiguity is not a data error. It is part of the measurement problem.

Why EFA rather than jumping directly to CFA?

The proposed three-domain framework is a starting theory, not a sufficiently established measurement model. We do not yet know whether employees distinguish Voice Safety from Role Clarity, whether Role Clarity and Adaptive Energy collapse together, or whether the ambiguous items define an unexpected dimension.

EFA is therefore aligned with the research question: What underlying processes could plausibly account for the observed item correlations? In contrast, CFA specifies relationships between indicators and latent variables in advance and evaluates whether that proposed factorial structure is consistent with the data (Tabachnick & Fidell, 2013; Wang & Wang, 2012).

Wang and Wang (2012) describe CFA as a method for determining or confirming the factorial structure of an already developed measurement instrument. That distinction matters here. Treating an uncertain 18-item structure as confirmed simply because three constructs were written into the questionnaire blueprint would put the claim ahead of the evidence.

Synthetic data and screening

Dataset

360 cases

18 questionnaire items.

Item-level missingness

86 of 6,480 responses

About 1.3% of possible item responses.

Complete-case EFA

282 cases

Remaining after complete-case analysis for the illustrative EFA.

Before factor analysis, the workflow checked response ranges, missingness, item distributions, and unusual response patterns. Data screening matters because errors, missing-value handling, unusual observations, and distributional problems can affect multivariate analyses and their correlations (Tabachnick & Fidell, 2013).

No universal “respondents per item” rule was used to declare the sample adequate. Tabachnick and Fidell (2013) explicitly describe factor-analysis sample requirements as dependent on conditions such as the strength of the correlations, communalities, number of factors, and how well those factors are determined. A fixed ratio therefore would hide information that actually matters.

The defensible question is not: “Did we achieve an arbitrary ratio?”

It is: “Are the correlations sufficiently stable and structured to support an interpretable factor analysis?”

Start with the correlation structure

Factor analysis operates on relationships among observed variables. The latent-factor interpretation is useful when shared variation among observed indicators can plausibly be attributed to underlying variables that are not themselves directly observed (Lovric, 2011; Tabachnick & Fidell, 2013).

Median absolute inter-item correlation

.16

Largest absolute inter-item correlation

.51

The matrix showed recognizable within-domain clusters without approaching a situation in which all 18 questions behaved like interchangeable versions of one item.

That pattern is substantively useful. We want enough shared variation to make common dimensions plausible, but we also want the questionnaire to contain distinctions worth explaining.

How many factors should we investigate?

This is one of the most consequential EFA decisions.

The first six eigenvalues from the synthetic correlation matrix were:

Component/factor position Observed eigenvalue 95th-percentile random eigenvalue
1 4.34 1.56
2 2.12 1.44
3 2.01 1.36
4 1.00 1.29
5 0.82 1.23
6 0.81 1.18

A scree plot would suggest a marked flattening after the first three dimensions. Tabachnick and Fidell (2013) describe the scree test as a judgment-based examination of where the decline in eigenvalues changes slope, while warning that the location of the break is not exact.

Parallel analysis provides another perspective by comparing observed eigenvalues with eigenvalues obtained from random data of corresponding dimensions. Its value is that random data themselves can generate nontrivial eigenvalues, so an observed eigenvalue should not automatically be interpreted as evidence for a meaningful factor (Tabachnick & Fidell, 2013).

In this synthetic analysis, three observed eigenvalues exceeded the corresponding 95th-percentile random eigenvalues. The fourth did not.

Factor-retention decision: Investigate three factors as the primary solution, while treating neighboring solutions as sensitivity checks rather than pretending that factor retention is mechanically determined by one statistic. The three-factor solution also remained consistent with the questionnaire's conceptual starting point.

Extraction: analyze common variance, not just total variance

Principal components analysis and common factor analysis answer related but different questions. Tabachnick and Fidell (2013) explain that principal components analysis analyzes total observed variance, whereas factor analysis focuses on shared variance while attempting to separate unique and error variance.

Extraction decision: Because the substantive objective here is to identify latent constructs underlying the item responses—not merely compress 18 columns into fewer numerical summaries—a common-factor extraction is the appropriate conceptual choice (Tabachnick & Fidell, 2013).

Tabachnick and Fidell (2013) discuss several factor-extraction procedures, including principal factors and maximum-likelihood factoring, and emphasize that researchers may examine alternative extraction specifications when assessing the stability and interpretability of a solution. The choice should therefore follow the measurement question and data conditions rather than software defaults.

For this synthetic case, the primary interpretation is framed as a common-factor EFA. The numerical loading illustration below was generated reproducibly from the synthetic dataset to expose the intended measurement complications; it should not be treated as evidence that one extraction algorithm is universally preferable.

Rotation: related constructs should be allowed to relate

Rotation is used to make an extracted factor solution easier to interpret. Tabachnick and Fidell (2013) distinguish orthogonal rotation, which constrains factors to be uncorrelated, from oblique rotation, which permits factors to correlate.

Rotation decision: There is no strong substantive reason to insist that feeling safe to speak, understanding one's role, and adapting energetically at work must be statistically independent. The measurement rationale instead anticipates related organizational experiences. An oblique solution is therefore the more natural interpretive target.

With oblique rotation, interpretation should distinguish the pattern matrix, which represents the unique relationships between variables and correlated factors, from the structure matrix, which contains correlations between variables and factors. Tabachnick and Fidell (2013) recommend interpreting factor meaning from the pattern matrix following oblique rotation.

This distinction becomes especially important for cross-loading items: an item can correlate with more than one factor partly because the factors themselves correlate.

What did the factor loadings show?

The synthetic loading pattern below summarizes the substantive result. Signs are oriented so that larger positive values correspond to more of the named construct; factor signs themselves are arbitrary in factor analysis.

Item Voice Safety Role Clarity Adaptive Energy Interpretation
I01 .66 .06 .14 Clear Voice Safety item
I02 .62 .12 .14 Clear Voice Safety item
I03 .59 .09 .11 Clear Voice Safety item
I04 .60 .10 .13 Clear Voice Safety item
I05 .48 .09 .05 Weaker but interpretable Voice Safety item
I06 .34 .46 .02 Cross-domain ambiguity
I07 .08 .68 .06 Clear Role Clarity item
I08 .12 .62 .16 Clear Role Clarity item
I09 .12 .60 .12 Clear Role Clarity item
I10 .04 .59 .05 Clear Role Clarity item
I11 .13 .52 .07 Clear Role Clarity item
I12 .06 .35 .41 Cross-domain ambiguity
I13 .08 .10 .69 Clear Adaptive Energy item
I14 .06 .07 .70 Clear Adaptive Energy item
I15 .05 .14 .55 Clear Adaptive Energy item
I16 .15 .03 .69 Clear Adaptive Energy item
I17 .31 .07 .48 Adaptive item with Voice Safety overlap
I18 .01 .14 .29 Weak, broadly worded item

All values are synthetic and rounded. The table is intended to demonstrate interpretation rather than establish universal loading thresholds.

A factor loading expresses the relationship between an observed variable and a factor; following rotation, researchers examine which variables are most strongly associated with each factor to interpret the dimensions (Tabachnick & Fidell, 2013).

The important feature here is not whether every number passes a universal cutoff. It is the pattern.

I01–I05 cluster on Voice Safety. I07–I11 cluster on Role Clarity. I13–I16 cluster on Adaptive Energy. That is substantial evidence that the intended conceptual distinctions are present in the synthetic data.

But four items deserve closer scrutiny.

Problematic item 1: I06 mixes two decisions

I06: “I know when I should speak up rather than handle an issue myself.”

The item was intended as Voice Safety, but its loading pattern points more strongly toward Role Clarity while retaining a meaningful Voice Safety relationship.

The wording explains why. “Speak up” invokes interpersonal voice, whereas “know when” and “handle an issue myself” invoke role boundaries and decision authority.

The statistical result therefore reveals a content problem rather than merely a bad number. Factor analysis is most useful in scale development when loading evidence is interpreted together with what the item actually asks (Adams & Lawrence, 2018; Sekaran & Bougie, 2016).

Exploratory decision: revise rather than automatically delete. A later item could isolate either interpersonal safety (“I can raise concerns even when doing so is uncomfortable”) or decision-boundary clarity (“I know which issues I should escalate rather than resolve independently”).

Problematic item 2: I12 bridges clarity and adaptation

I12: “When work changes, I know what I should focus my effort on.”

I12 loads on both Role Clarity and Adaptive Energy.

Again, that is substantively plausible. “Know what I should focus on” represents clarity, while “when work changes” embeds the item in an adaptive context.

Removing it might produce a cleaner statistical pattern, but cleaner is not automatically better. If the intended construct of Role Clarity explicitly includes clarity under changing conditions, the item may cover important content.

Exploratory decision: flag I12 for rewriting and compare the revised version in new data. Do not erase substantive coverage solely to simplify a loading table.

Problematic item 3: I17 may represent a real construct boundary

I17: “I have the confidence to suggest a different approach when the current one is not working.”

Its strongest loading is on Adaptive Energy, but it also relates to Voice Safety.

That overlap makes conceptual sense: suggesting a different approach requires adaptation, but doing so publicly also involves speaking up.

A cross-loading is therefore not necessarily statistical contamination. It can indicate that an item genuinely spans two neighboring constructs. The researcher's task is to decide whether that breadth is desirable for the intended measurement claim.

Exploratory decision: retain I17 provisionally if adaptive initiative is central to Adaptive Energy, but test a narrower rewrite if separation from Voice Safety is important.

Problematic item 4: I18 is too generic

I18: “I usually cope with whatever happens at work.”

Unlike I06, I12, and I17, the problem is not primarily competing factors. I18 simply shows a weak relationship with the intended Adaptive Energy dimension.

The wording is correspondingly broad. “Whatever happens” could involve workload, conflict, uncertainty, technical difficulty, health, managerial behavior, or almost anything else.

Exploratory decision: remove I18 from the provisional version and replace it with an item tied more directly to adaptive energy. That preserves the construct definition rather than trying to rescue a vague indicator.

The provisional factor structure

After interpreting the loadings together with item content, the synthetic EFA supports this provisional structure:

Factor 1: Voice Safety

Strongest indicators: I01–I05.

The common theme is interpersonal freedom to raise concerns, admit uncertainty, question decisions, and discuss difficult issues.

I06 is not treated as a clean indicator because it mixes voice with role boundaries.

Factor 2: Role Clarity

Strongest indicators: I07–I11.

These items concern expectations, priorities, decision responsibility, and performance criteria.

I12 sits near the boundary between clarity and adaptation and should be revised rather than silently assigned to whichever column contains the larger loading.

Factor 3: Adaptive Energy

Strongest indicators: I13–I16.

These items describe sustained engagement, energy, recovery, and willingness to alter one's approach when conditions change.

I17 is provisionally retained as a broader adaptive-initiative item. I18 is removed from the provisional scale because its content and factor behavior are both weak.

Why the EFA does not “validate” the questionnaire

It would be tempting to report that the analysis “confirmed a three-factor structure.” That language is too strong.

The same synthetic observations were used to discover the factor count, inspect alternative solutions, identify problematic items, reinterpret item meaning, and make retention decisions. Those choices make the resulting structure partly adapted to this particular dataset.

EFA is associated with theory development, whereas CFA evaluates a hypothesized factorial structure specified through relationships between observed indicators and latent factors (Tabachnick & Fidell, 2013; Wang & Wang, 2012). Wang and Wang (2012) specifically describe CFA as testing whether a theoretically defined or hypothesized factorial structure is supported by the data.

The defensible conclusion is that the synthetic EFA produced an interpretable provisional three-factor structure.

It is not that the questionnaire now has a confirmed three-factor measurement model.

Why exploratory item revisions must stay exploratory

Suppose I18 is removed, I06 and I12 are rewritten, and I17 is retained after reviewing its conceptual role. If the factor analysis is simply rerun on the same cases, the revised solution may look cleaner.

That is useful for development—but it is still development.

The item decisions were informed by those cases. Reusing them to declare the revised model confirmed would blur the distinction between discovering a model and testing one.

The revised questionnaire should therefore be documented as an exploratory measurement specification generated by this EFA. The decision record should retain what changed, why it changed, and which evidence motivated each revision. That preserves the distinction between theoretically motivated refinement and post hoc statistical optimization.

What would be needed before claiming confirmation?

A stronger next stage would collect or reserve data that were not used to choose the factor structure or revise the items and specify the proposed measurement model before examining the confirmatory results.

CFA explicitly links observed indicators to hypothesized latent factors and evaluates whether that prespecified factorial structure is consistent with the data (Wang & Wang, 2012). Measurement models should be established before relying on a broader structural model because problems in the measurement portion can undermine interpretation of the structural relationships (Wang & Wang, 2012).

For this questionnaire, the confirmatory model would specify:

  • Voice Safety using the retained or rewritten Voice Safety indicators;
  • Role Clarity using its prespecified indicators;
  • Adaptive Energy using its prespecified indicators;
  • correlations among the three latent factors where theoretically justified; and
  • no new item deletion or factor reassignment simply because the confirmatory output suggests an easier-fitting alternative.

Any substantial modification made after inspecting the confirmatory data would need to be labeled accordingly rather than folded invisibly into a claim of confirmation.

Confirmation would also not eliminate the broader validity question. Construct validity concerns whether the measure represents the intended theoretical construct, and content validity asks whether its items adequately cover that construct (Adams & Lawrence, 2018). A statistically tidy factor model is therefore evidence about measurement structure, not a substitute for substantive construct definition.

What this EFA example demonstrates

The main decision in exploratory factor analysis is not “Which button produces a rotated matrix?” It is whether the proposed factor structure can be defended as a measurement interpretation.

In this synthetic case, several pieces of evidence converged on three provisional dimensions: the correlation pattern contained meaningful clusters, the first three observed eigenvalues exceeded the parallel-analysis comparison values, the loading pattern produced three interpretable groups, and those groups largely matched the constructs defined before analysis. Factor-retention methods should nevertheless be interpreted rather than treated as infallible automatic rules (Tabachnick & Fidell, 2013).

The difficult items were also informative. I06 exposed overlap between voice and role boundaries. I12 exposed overlap between clarity and adaptation. I17 sat at a theoretically plausible boundary between adaptive initiative and voice. I18 was simply too broad to function well as an indicator.

That is a more useful outcome than mechanically deleting every imperfect item.

The defensible endpoint of EFA is therefore not a declaration that the scale has been proved correct. It is a transparent, theoretically interpretable candidate measurement structure that is specific enough to be tested next.

Practical EFA decision checklist

  • Define the constructs before interpreting factors.
  • Check item coding, ranges, missingness, distributions, and unusual cases before modeling.
  • Inspect whether the correlation matrix contains meaningful shared structure.
  • Do not justify sample adequacy with an unsupported universal respondent-to-item ratio.
  • Investigate factor count using multiple pieces of evidence rather than one automatic rule.
  • Match the extraction approach to the question being asked.
  • Do not force orthogonal rotation when correlated constructs are substantively plausible.
  • With oblique rotation, distinguish pattern coefficients from simple item-factor correlations.
  • Review primary loadings and competing loadings together with item wording.
  • Treat cross-loadings as measurement evidence requiring interpretation, not automatic deletion commands.
  • Preserve content coverage when considering item removal.
  • Record every exploratory item revision.
  • Call a data-informed revision exploratory.
  • Test the resulting prespecified structure on independent or genuinely held-out data before describing it as confirmed.

Conclusion

Exploratory factor analysis is most defensible when it is treated as a measurement decision process rather than a search for a perfectly clean loading matrix.

For this synthetic 18-item questionnaire, the evidence supports investigating three related constructs: Voice Safety, Role Clarity, and Adaptive Energy. Most items behave coherently, but I06, I12, I17, and I18 reveal different kinds of measurement problems. The first three illustrate meaningful construct overlap; the fourth is simply too broad.

The result is therefore a provisional three-factor questionnaire requiring revision and independent confirmatory testing. That conclusion is less dramatic than claiming the scale has been validated, but it is much easier to defend.

References

Adams, K. A., & Lawrence, E. K. (2018). Research methods, statistics, and applications (2nd ed.). SAGE Publications.

Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2

Sekaran, U., & Bougie, R. (2016). Research methods for business: A skill-building approach (7th ed.). John Wiley & Sons.

Tabachnick, B. G., & Fidell, L. S. (2013). Using multivariate statistics (6th ed.). Pearson.

Wang, J., & Wang, X. (2012). Structural equation modeling: Applications using Mplus. John Wiley & Sons.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry