Resource

PCA or Factor Analysis? The Same Variables, Two Different Research Questions

Principal component analysis and factor analysis can produce similar-looking output while answering different research questions. This synthetic case study uses the same 12 correlated survey variables to show why PCA is a dimensionality-reduction method and common factor analysis is a latent-variable framework.

Principal component analysis (PCA) and factor analysis can begin with the same correlation matrix, appear beside one another in statistical software, and even produce superficially similar loading tables. That does not make them interchangeable.

The choice begins with the research question. PCA is appropriate when the objective is to summarize many observed variables with fewer components while retaining substantial variation in the observed data. Common factor analysis instead represents observed correlations in terms of a smaller number of underlying factors, separating common variance from variance unique to individual variables and error (Lovric, 2011; Tabachnick & Fidell, 2013).

This synthetic case study holds the variables and respondents constant. Only the stakeholder's research objective changes. The result is two different analytical frameworks—and two different kinds of conclusion.

Synthetic-data disclosure: All respondents, variables, correlations, and numerical results below are synthetic and were created solely for methodological illustration. They are not empirical findings about a real population.

The research situation

A research team administers a 12-item workplace survey to 480 synthetic respondents. Each item uses a seven-point response scale from 1 = strongly disagree to 7 = strongly agree.

The items deliberately form correlated clusters:

Synthetic survey items
Variable Synthetic survey item
Q1 I can decide how to organize my work.
Q2 I have discretion over how I complete important tasks.
Q3 I can adjust my work methods when circumstances change.
Q4 I have meaningful control over day-to-day work decisions.
Q5 Colleagues provide help when I encounter difficulties.
Q6 People in my team are willing to support one another.
Q7 I can obtain useful assistance from coworkers when needed.
Q8 Team members make time to help each other.
Q9 I feel mentally drained by the end of the workday.
Q10 My work leaves me emotionally depleted.
Q11 It is difficult to recover my energy after work.
Q12 Work demands consume much of my mental energy.

The labels are deliberately withheld from the analysis. The point is to ask what the statistical procedure itself is being asked to accomplish.

The stakeholders then present two goals.

Stakeholder A

“Twelve variables are cumbersome. Can we replace them with a few summary variables without throwing away most of the information?”

Objective: Reduce the observed variables to a smaller set of useful summaries.

Stakeholder B

“We think the responses reflect a smaller set of underlying psychological constructs. What latent dimensions could account for the correlations among the items?”

Objective: Investigate underlying dimensions that may account for correlations among the observed items.

Those questions sound related. Statistically, however, they are not the same.

The decision comes before the software

Tabachnick and Fidell (2013) make the distinction in terms of the goals and variance being analyzed. PCA analyzes the variance of the observed variables and seeks components that successively account for as much of that variance as possible. Common factor analysis focuses on the variance shared among variables and attempts to represent their correlations through underlying common factors.

Lovric (2011) similarly describes PCA as a dimension-reduction technique that forms linear combinations of the observed variables, whereas factor analysis belongs to the latent-variable framework in which unobservable variables are used to explain correlations among observed variables.

The decision can therefore be summarized as follows:

PCA and common factor analysis answer different research questions
Decision PCA Common factor analysis
Primary question Can many observed variables be summarized by fewer composites? What underlying factors may account for correlations among observed variables?
Main object Components Latent/common factors
Variance emphasis Total observed variance Shared/common variance
Component/factor meaning Weighted combination of observed variables Unobserved dimension represented through observed indicators
Typical use here Dimensionality reduction Investigation of latent structure
Appropriate conclusion “These variables can be summarized by these components.” “The correlation structure is consistent with these underlying factors.”

Core distinction: The same data do not remove the conceptual distinction. The appropriate method depends on what the researcher is asking the analysis to represent.


Goal A: Reduce 12 observed variables to a smaller set of composites

For Stakeholder A, the problem is operational: reduce dimensionality.

That objective points naturally toward PCA. PCA finds linear combinations of the observed variables called principal components. The first component is selected to have the largest possible variance; subsequent components successively maximize remaining variance while being uncorrelated with earlier components in the unrotated solution (Lovric, 2011; Tabachnick & Fidell, 2013).

What variance does PCA use?

With standardized variables, PCA begins with ones on the diagonal of the correlation matrix. In this formulation, each observed variable contributes its full unit variance. Common, unique, and error-related variation are therefore not separated before components are extracted (Tabachnick & Fidell, 2013).

This feature makes sense for Stakeholder A. The objective is not to construct a model of latent causes. It is to preserve as much information in the observed variables as possible while using fewer dimensions.

Synthetic PCA results

PCA of the 12 standardized synthetic variables produces the following first six eigenvalues:

First six PCA eigenvalues in the synthetic example
Component Eigenvalue Variance explained Cumulative variance
PC1 4.23 35.2% 35.2%
PC2 1.92 16.0% 51.2%
PC3 1.51 12.5% 63.7%
PC4 0.64 5.3% 69.0%
PC5 0.57 4.8% 73.8%
PC6 0.52 4.4% 78.2%

Observed variables

12 standardized synthetic survey variables.

First three components

PC1, PC2, and PC3 provide the proposed lower-dimensional representation in this illustration.

Cumulative variance

The first three components summarize approximately 63.7% of the total standardized-variable variance.

These values are properties of this synthetic demonstration, not recommendations for how many components should always be retained.

The first three components summarize approximately 63.7% of the total standardized-variable variance. The research team could therefore replace 12 observed variables with three component scores if that degree of compression is adequate for its intended use.

PCA retention should be treated as a substantive dimensionality-reduction decision rather than as a universal percentage rule. Lovric (2011) notes that decisions about how many components to retain depend on context and that several approaches are available.

What is a principal component?

A component is a weighted linear combination of the observed variables. Conceptually:

PC1 = a1Q1 + a2Q2 + ⋯ + a12Q12

The coefficients are chosen so that the first component captures maximum variance, with subsequent components accounting for additional variance subject to the PCA constraints (Lovric, 2011; Tabachnick & Fidell, 2013).

That direction matters.

The component is constructed from the observed variables. It should not automatically be reinterpreted as an unobserved psychological entity that generated those variables.

Rotation can make the PCA easier to interpret

An unrotated PCA is optimized for variance extraction, not necessarily for substantive simplicity. Rotation can be applied after retaining components to make their relationships with the observed variables easier to interpret. Rotation changes the orientation used to describe the retained multidimensional space rather than turning components into latent factors (Lovric, 2011; Tabachnick & Fidell, 2013).

For illustration, an orthogonal varimax rotation of the first three synthetic components produces a simple pattern. After reordering the rotated components for presentation, the approximate loadings are:

Approximate rotated PCA loadings
Variable Component 1 Component 2 Component 3
Q1.80.20.12
Q2.80.12.17
Q3.79.05.07
Q4.68.32.12
Q5.15.81.08
Q6.12.80.11
Q7.24.72.07
Q8.11.75.13
Q9.22.06.80
Q10.07.15.81
Q11.13.12.74
Q12.06.06.78

The clusters are substantively recognizable. Component 1 is dominated by Q1–Q4, Component 2 by Q5–Q8, and Component 3 by Q9–Q12.

The stakeholder could give those components descriptive labels such as work discretion, coworker support, and work depletion.

But the wording of the conclusion matters.

“The 12 observed survey variables can be represented compactly by three interpretable weighted composites in this synthetic dataset.”

A PCA alone does not justify:

“Three latent psychological constructs generated these responses.”

That second statement asks a factor-model question.


Goal B: Investigate underlying latent constructs

Stakeholder B has a different objective. The team is not merely trying to compress a spreadsheet. It proposes that correlations among the observed responses arise because respondents differ on a smaller number of unobserved constructs.

That objective points toward common factor analysis.

Factor analysis belongs to a latent-variable framework. A latent variable cannot be observed directly; observed variables serve as manifestations or indicators through which the underlying dimension is investigated (Lovric, 2011). Wang and Wang (2012) likewise describe latent factors as unobservable variables that must be estimated indirectly through observed indicators.

Shared variance is now central

The distinction becomes especially clear on the diagonal of the correlation matrix.

PCA analyzes the full standardized variance of each variable. Common factor analysis instead estimates the amount of variance each observed variable shares with the other observed variables—the variable's communality—and analyzes that common variance. Unique and error variance are not treated as common-factor variance (Tabachnick & Fidell, 2013).

Suppose Q1 has substantial variation that is specific to its wording or measurement. PCA still incorporates that observed variance into its component solution. A common-factor model attempts instead to isolate the part of Q1 that participates in the common covariance structure.

That is why changing the stakeholder's question changes the appropriate model even though Q1 through Q12 have not changed.

The extraction logic is different

Several common-factor extraction approaches exist, and they do not all optimize the same criterion. For example, principal-factor extraction begins with estimates of communalities rather than ones on the diagonal and seeks factors from common variance, whereas maximum-likelihood factor extraction estimates loadings under a likelihood framework. Other approaches seek to minimize discrepancies between observed and reproduced correlations (Tabachnick & Fidell, 2013).

The important point for this case is not that one extraction method is universally superior. It is that common-factor extraction is built around a factor model rather than the PCA objective of decomposing total observed variance.

For the synthetic illustration, a three-factor common-factor solution was estimated from the same 12 variables.

The estimated communalities were approximately:

Estimated communalities in the synthetic common-factor solution
Variable Estimated communality
Q1.60
Q2.58
Q3.45
Q4.47
Q5.59
Q6.54
Q7.45
Q8.45
Q9.59
Q10.59
Q11.44
Q12.46

These numbers illustrate the key distinction. The factor model is not treating every standardized variable as contributing a full variance of 1 to the common-factor solution. Instead, only part of each variable's variance is represented as common.

What is a latent factor?

The conceptual direction now reverses.

The observed variables are modeled as indicators of underlying factors. Wang and Wang (2012) describe factor loadings as links between latent factors and their observed indicators; the observed measure also contains a residual or measurement-error component.

A simplified representation is:

Qi = λi1F1 + λi2F2 + λi3F3 + ei

where the F's are latent factors, the λ's are factor loadings, and ei represents variance not accounted for by the common factors.

The research interpretation is therefore different from PCA. The question is whether a relatively small number of underlying dimensions can account for the observed correlation pattern.

Rotation is about interpretability, not changing the model into something else

Unrotated factor solutions are often difficult to interpret. Rotation is therefore commonly used to obtain a clearer structure. Tabachnick and Fidell (2013) emphasize that rotation is intended to improve interpretability and scientific usefulness rather than improve the mathematical fit of an orthogonally rotated solution.

A major decision is whether the factors should be constrained to remain uncorrelated.

Orthogonal rotation

Factors remain uncorrelated.

Oblique rotation

Factors are permitted to correlate.

Consider when: the underlying processes are theoretically expected to relate to one another.

If the underlying processes are theoretically expected to relate to one another, an oblique solution can be more appropriate than forcing independence merely for convenience (Tabachnick & Fidell, 2013).

For this survey, it is plausible that perceived discretion, coworker support, and depletion could be related. A researcher investigating latent constructs should therefore consider an oblique factor solution rather than automatically selecting varimax simply because it appears as a familiar software option.

When oblique rotation is used, interpretation also requires attention to the distinction between the pattern and structure matrices. The pattern matrix reflects the distinctive relationship of a factor with a variable after overlap among correlated factors is accounted for, whereas the structure matrix contains correlations between variables and factors that also incorporate factor overlap. Tabachnick and Fidell (2013) note that researchers commonly emphasize the pattern matrix while considering the factor correlations and other output.

Why the two analyses may look surprisingly similar

At first glance, the synthetic PCA and common-factor results both suggest three clusters:

  • Q1–Q4;
  • Q5–Q8;
  • Q9–Q12.

This resemblance is precisely where researchers can get into trouble.

Similar numerical patterns do not make PCA and factor analysis equivalent. Tabachnick and Fidell (2013) show that extraction methods can yield similar-looking loading patterns when the data have a strong structure even though the statistical models and variance being analyzed differ.

The distinction is therefore conceptual as well as computational.

PCA asks

Which weighted combinations summarize the observed variables efficiently?

Factor analysis asks

Which underlying common dimensions can account for the correlations among the observed variables?

Those questions can sometimes lead to visually similar loading tables. They still support different claims.

Why software menus make the distinction easy to miss

Statistical software contributes to the confusion.

PCA and common factor analysis may be offered within the same “factor analysis” procedure. Users may select the same variables, request an extraction, choose a number of dimensions, apply a rotation, and receive matrices of loadings that look almost identical in layout.

Lovric (2011) explicitly notes that confusion arises partly because widely used software packages may treat PCA as though it were a special case of factor analysis, even though PCA is not factor analysis.

Tabachnick and Fidell (2013) likewise discuss PCA and several factor-extraction methods within the same practical workflow and note that both IBM SPSS and SAS provide these procedures. The common interface is a software convenience; it does not erase the methodological distinction.

A menu sequence such as:

Analyze → Dimension Reduction → Factor → Extraction

can therefore hide the most important decision.

The researcher must still decide what is being modeled.

Principal Components

Selecting Principal Components asks the program for a component solution based on total observed variance.

Common-factor extraction

Selecting a common-factor extraction method changes the framework to one concerned with communalities and shared covariance.

Changing one dropdown can therefore change the interpretation of the entire analysis.


Same variables, different claims

Stakeholder A: defensible PCA interpretation

The team can say:

“The 12 survey variables contained substantial redundancy. Three principal components provided a lower-dimensional representation of the observed variables and accounted for approximately 63.7% of their total standardized variance in this synthetic dataset.”

The team can then use the retained component scores as compact observed-variable summaries when that serves the subsequent research purpose.

PCA is particularly well aligned with this goal because dimensionality reduction is one of its central uses (Lovric, 2011; Tabachnick & Fidell, 2013).

What Stakeholder A cannot claim from PCA alone

The team should not convert that result into a latent-variable claim.

PCA alone does not demonstrate that:

  • three unobservable constructs generated the responses;
  • the components are error-free psychological traits;
  • the component structure establishes construct validity;
  • the variables satisfy a latent-factor measurement model.

The components are weighted summaries of observed variables.

Stakeholder B: defensible factor-analysis interpretation

The second team can instead say:

“A three-factor common-factor model provided an interpretable representation of the shared covariance among the 12 synthetic survey items, with Q1–Q4, Q5–Q8, and Q9–Q12 primarily associated with different underlying dimensions.”

That interpretation is explicitly about common latent dimensions rather than merely compressing observed variables.

Factor analysis is designed for this type of latent-variable question: latent variables are introduced to account for associations among manifest variables (Lovric, 2011).

What Stakeholder B still cannot claim

Exploratory factor analysis does not make a proposed construct unquestionably real.

Wang and Wang (2012) distinguish exploratory factor analysis from confirmatory factor analysis: EFA is used when the factorial structure is not fully specified in advance, whereas CFA evaluates a theoretically specified measurement structure.

A sensible exploratory factor solution can therefore motivate a more explicit measurement model, but it is not equivalent to confirmatory evidence.

Nor does finding factors establish causal relationships among those factors or between the factors and other variables.

A practical decision framework

When PCA and factor analysis both appear plausible, begin with the research objective rather than the software menu.

Research-question guide for choosing the stronger starting framework
Ask this question If yes, the stronger starting point is
Do I mainly need fewer variables for description, visualization, prediction, or later analysis? PCA
Do I want weighted summaries that preserve substantial observed variance? PCA
Is the substantive question explicitly about unobserved constructs underlying observed indicators? Common factor analysis
Do I need to distinguish shared variance from variable-specific/error variance? Common factor analysis
Am I evaluating whether observed indicators behave as manifestations of latent dimensions? Factor-analytic framework
Am I tempted to call PCA components “latent factors” simply because software labels the procedure “Factor”? Reconsider the interpretation

This is not a contest in which one technique is universally better. Each method answers a different question.

The central lesson from the synthetic case

Nothing about the dataset changed.

There were still 480 synthetic respondents. There were still 12 correlated survey variables. The correlation matrix was the same.

What changed was the research question.

Stakeholder A

The observed variables themselves were the objects to be summarized.

Framework: PCA, producing weighted components intended to retain substantial total observed variance.

Stakeholder B

The observed variables were treated as manifestations of underlying dimensions.

Framework: Common factor analysis, focusing on shared variance and representing correlations through latent factors.

That distinction is more important than whether two loading matrices happen to look similar.

PCA is a dimensionality-reduction framework. Common factor analysis is a latent-variable framework. The same variables can legitimately be analyzed with either—but only when the research question justifies the corresponding interpretation. (Lovric, 2011; Tabachnick & Fidell, 2013)

References

Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2

Tabachnick, B. G., & Fidell, L. S. (2013). Using multivariate statistics (6th ed.). Pearson.

Wang, J., & Wang, X. (2012). Structural equation modeling: Applications using Mplus. John Wiley & Sons.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry