Resource

Main Effects vs Interaction Effects: How to Interpret a Factorial Model

Main effects summarize differences across a factor, while interactions ask whether those differences depend on another factor. Four synthetic 2 × 2 examples show why cell means, simple effects, and interaction plots matter alongside factorial ANOVA tests.

A factorial model asks more than whether one factor is associated with an outcome. It asks whether two or more factors matter separately and whether their effects depend on each other. In a two-factor design, this distinction is captured by the main effects of the factors and their interaction effect. Two-way ANOVA is one common way to analyze such a design when the factors are categorical and the outcome is quantitative (Moore et al., 2014, 2021; Sekaran & Bougie, 2016).

The central interpretive question is not simply, “Which effects are significant?”

What pattern of cell means generated the main effects and interaction, and which comparison answers the research question?

Synthetic examples: This Resource uses four synthetic 2 × 2 mean tables to show why that question matters.

What is a factorial design?

A factorial design includes two or more factors so that their separate and joint relationships with an outcome can be studied in the same design. With two factors, a two-way factorial analysis can assess the main effect of each factor and the interaction between them (Adams & Lawrence, 2018; Moore et al., 2014, 2021; Sekaran & Bougie, 2016).

Suppose a synthetic study examines employee performance after two types of training under two workload conditions:

  • Training: Standard or Interactive
  • Workload: Low or High
  • Outcome: Mean task-performance score, where higher scores indicate better performance

This is a 2 × 2 factorial design because there are two factors and each has two levels. The four combinations of training and workload form four cells.

Factorial designs are useful when the research question concerns joint effects. Sekaran and Bougie (2016) note that factorial designs allow multiple manipulations to be examined simultaneously, including their main and interaction effects, and can be more efficient than conducting several separate single-factor randomized designs.

Main effect vs interaction effect

A main effect describes differences associated with one factor when considered across the levels of the other factor. In a balanced 2 × 2 table, this can be seen by comparing the relevant marginal means—the means obtained after averaging across the other factor. Two-way ANOVA separates variation associated with each factor from variation associated with their interaction (Moore et al., 2014, 2021).

An interaction effect addresses a different question: does the effect or difference associated with one factor change depending on the level of the other factor? Adams and Lawrence (2018) define an interaction in terms of the effect of one variable across different levels of another, while Sekaran and Bougie (2016) similarly describe an interaction as occurring when one factor's effect depends on the level of the other factor.

Hayes (2022) expresses the same idea through moderation. In a two-factor model, an interaction is fundamentally a difference in differences: the effect of one factor at one level of the other factor differs from its effect at another level. A 2 × 2 factorial ANOVA can also be represented as a regression model containing the two constituent variables and their product term, although the interpretation of individual regression coefficients depends on how categorical variables are coded (Hayes, 2022).

That distinction gives three different research questions:

Training main effect

Does performance differ between Standard and Interactive training when averaged across workload?

Workload main effect

Does performance differ between Low and High workload when averaged across training?

Training × Workload interaction

Does the training difference depend on workload?

The third question cannot be answered from the first two.

Synthetic pattern 1: Two main effects, no interaction

Consider these illustrative means:

Synthetic Pattern 1: cell means and marginal means
Training Low workload High workload Training marginal mean
Standard 60 70 65
Interactive 70 80 75
Workload marginal mean 65 75

The Interactive-minus-Standard difference is:

  • Low workload: 70 − 60 = 10
  • High workload: 80 − 70 = 10

The High-minus-Low workload difference is also 10 under both training conditions.

There are therefore two clear descriptive main-effect patterns. Interactive training is associated with scores 10 points higher on average than Standard training, and High workload is associated with scores 10 points higher on average than Low workload.

But there is no interaction in these means because the training difference is identical at both workload levels. Equivalently, the workload difference is identical under both training conditions. The effects are additive rather than conditional (Hayes, 2022; Moore et al., 2021).

What would the interaction plot look like?

Plot workload on the horizontal axis and performance on the vertical axis, with one line for each training condition. Both lines rise by 10 points and remain parallel.

Parallel lines are the characteristic graphical pattern of no interaction in a simple 2 × 2 mean plot: the difference represented by the gap between the lines remains constant (Adams & Lawrence, 2018; Hayes, 2022).

The practical interpretation is therefore straightforward: the descriptive training difference does not change with workload.

Synthetic pattern 2: Interaction, no main effects

Now change the cell means:

Synthetic Pattern 2: cell means and marginal means
Training Low workload High workload Training marginal mean
Standard 60 80 70
Interactive 80 60 70
Workload marginal mean 70 70

The marginal means suggest:

  • Standard training mean = 70
  • Interactive training mean = 70
  • Low-workload mean = 70
  • High-workload mean = 70

There is therefore no descriptive main-effect difference for either factor.

Yet the cell means contain a large interaction.

The Interactive-minus-Standard difference is:

  • Low workload: 80 − 60 = +20
  • High workload: 60 − 80 = −20

The difference not only changes in magnitude; it reverses direction.

“No main effect” does not mean “nothing is happening.” Marginal averaging has completely hidden the conditional pattern. At low workload, Interactive training has the higher mean; at high workload, Standard training has the higher mean.

The interaction is the difference between those differences:

(+20) − (−20) = 40 points.

Hayes (2022) emphasizes that an interaction or moderation effect concerns whether conditional effects differ. The relevant question is not whether one conditional comparison is statistically significant while another is not; the interaction itself tests whether the effects differ from each other.

What would the interaction plot look like?

The two lines cross. One rises from 60 to 80, while the other falls from 80 to 60.

That graph communicates something that the marginal means cannot: which training condition looks better depends entirely on workload.

Synthetic pattern 3: Main effects and an interaction

A factorial model can contain both marginal differences and an interaction.

Synthetic Pattern 3: cell means and marginal means
Training Low workload High workload Training marginal mean
Standard 60 65 62.5
Interactive 70 90 80
Workload marginal mean 65 77.5

The marginal means show substantial descriptive differences:

  • Interactive versus Standard: 80 − 62.5 = 17.5
  • High versus Low workload: 77.5 − 65 = 12.5

But those averages do not tell the full story.

The Interactive-minus-Standard difference is:

  • Low workload: 70 − 60 = 10
  • High workload: 90 − 65 = 25

The training difference is therefore much larger under High workload. The difference in differences is:

25 − 10 = 15 points.

This is an interaction even though Interactive training has the higher mean at both workload levels.

An interaction does not require the lines to cross. Nonparallel lines are enough to indicate an interaction pattern in the cell means because the size of one factor's difference changes across levels of the other factor (Adams & Lawrence, 2018; Hayes, 2022).

The marginal statement—“Interactive training has a higher overall mean”—remains descriptively true. But it is incomplete. If the research question asks how much training matters under each workload condition, the conditional comparisons of 10 and 25 points are more informative than the single marginal difference of 17.5 points.

Synthetic pattern 4: A crossover interaction

Consider a second crossing pattern:

Synthetic Pattern 4: cell means and marginal means
Training Low workload High workload Training marginal mean
Standard 60 90 75
Interactive 85 65 75
Workload marginal mean 72.5 77.5

There is no overall training difference: both training marginal means equal 75.

There is a modest descriptive High-versus-Low workload difference of 5 points after averaging across training.

But the conditional training comparisons tell a radically different story:

  • Low workload: Interactive − Standard = 85 − 60 = +25
  • High workload: Interactive − Standard = 65 − 90 = −25

The lines cross because the direction of the training difference reverses.

Dmitrienko and Koch (2017), discussing treatment-by-stratum interactions, distinguish quantitative interactions, in which an effect changes in magnitude without changing direction, from qualitative interactions, in which the direction of the effect changes across strata. They note that qualitative interactions are also called crossover interactions.

Pattern 3: Quantitative-type pattern

Interactive training remains higher under both workloads, but the size of the difference changes.

Pattern 4: Qualitative or crossover pattern

Which training condition has the higher mean reverses across workload.

This distinction is substantive, not merely graphical. A change from a 10-point advantage to a 25-point advantage is different from a change from a 25-point advantage to a 25-point disadvantage.

What does a significant interaction mean?

In a factorial ANOVA, a statistically significant interaction provides evidence that the effect associated with one factor differs across levels of the other factor. In a 2 × 2 setting, it is a statistical test of a difference in simple effects, or equivalently a difference in differences (Hayes, 2022).

It does not, by itself, tell you:

  • which conditional effect is substantively important;
  • whether each individual simple effect is statistically distinguishable from zero;
  • whether the pattern is practically meaningful;
  • whether the interaction is a magnitude change or a direction reversal; or
  • whether a causal interpretation is warranted by the study design.

Those questions require examination of the estimated means, conditional effects, uncertainty, study design, and research question rather than the interaction p-value alone (Hayes, 2022).

A difference in statistical significance is not itself evidence of a statistically significant difference. A common mistake is to reason that because one simple effect is significant and another is not, the two simple effects must differ. Hayes (2022) explicitly cautions against this logic. If the research question concerns interaction, the interaction must be tested directly.

Why a significant interaction changes the interpretation of main effects

Main effects in a factorial analysis are marginal summaries. They describe what happens after combining information across levels of another factor. An interaction says that the underlying conditional differences are not constant across those levels (Moore et al., 2014, 2021).

That creates an interpretive tension.

Consider Pattern 3. Saying that Interactive training is 17.5 points higher “on average” is mathematically informative, but it combines a 10-point difference under Low workload with a 25-point difference under High workload.

Pattern 4 is more striking. The overall training difference is zero even though there are large conditional differences in opposite directions.

Therefore, a significant interaction does not mechanically make main effects invalid or require researchers to stop reporting them. Instead, it changes what those main effects can reasonably communicate. When conditional effects differ meaningfully, a marginal main effect may be too coarse to answer the substantive question (Hayes, 2022).

Interpretive sequence: Interaction → inspect cell means and plot → estimate relevant conditional/simple effects → interpret marginal main effects in light of that pattern.

Conditional effects and simple effects

A conditional effect describes the effect of one variable at a specified value or level of another variable. In a categorical 2 × 2 factorial model, these comparisons are commonly described as simple effects (Hayes, 2022).

For Pattern 3, the simple effect of training is:

  • 10 points under Low workload
  • 25 points under High workload

The interaction asks whether those two effects differ. The simple effects then describe where and how the interaction operates.

Testing an interaction and probing an interaction are therefore different tasks. The interaction test evaluates whether the conditional effects differ; probing describes and, where appropriate, tests particular conditional effects after that dependency has been considered (Hayes, 2022).

This distinction helps prevent a common reporting error:

“Training worked under High workload because p < .05 but did not work under Low workload because p > .05.”

That conclusion does not establish interaction. The appropriate interaction question is whether the training effect under High workload differs from the training effect under Low workload (Hayes, 2022).

How to interpret an interaction plot

An interaction plot should be read as a visualization of the cell means, not as a substitute for statistical inference.

For a 2 × 2 design, ask three questions.

  1. Are the lines approximately parallel? Parallel lines indicate that the difference associated with one factor remains constant across levels of the other factor, corresponding to no interaction in the plotted means. Nonparallel lines indicate an interaction pattern (Adams & Lawrence, 2018; Hayes, 2022).

  2. How large is the vertical gap between the lines at each level? That gap represents a conditional or simple effect when the lines represent levels of the focal factor. Compare the gaps rather than merely asking whether the lines cross (Hayes, 2022).

  3. Does the direction reverse? Crossing lines can indicate a crossover pattern in which the sign of the conditional difference changes. But crossing is not required for interaction; changing magnitude without reversal can also produce an interaction (Dmitrienko & Koch, 2017; Hayes, 2022).

Adams and Lawrence (2018) also show how 2 × 2 plots help distinguish main effects from interactions: the slope of lines can reflect differences associated with the factor on the horizontal axis, while separation between lines reflects differences associated with the factor represented by those lines. Parallelism versus nonparallelism then helps reveal whether the factors interact.

Main effects are averages; interactions are dependencies

A compact way to distinguish the concepts is:

Research questions and the corresponding factorial effects
Question Effect being examined
Does Training differ overall across workload conditions? Training main effect
Does Workload differ overall across training conditions? Workload main effect
Does the Training difference change with Workload? Training × Workload interaction
How large is the Training difference under Low workload? Conditional/simple effect
How large is the Training difference under High workload? Conditional/simple effect

This is why factorial analysis should begin with the research question rather than with whichever p-value appears first in an ANOVA table. Factorial designs were developed precisely to represent joint effects that cannot necessarily be understood by considering each factor separately (Lovric, 2011; Moore et al., 2021; Sekaran & Bougie, 2016).

Do not interpret the factorial ANOVA from p-values alone

A p-value addresses sampling uncertainty under a specified statistical model. It does not communicate the magnitude or substantive pattern of an interaction by itself.

For each factorial effect, examine at least:

  • the four cell means;
  • relevant marginal means;
  • the direction and magnitude of simple differences;
  • the difference in differences;
  • an interaction plot;
  • estimates of uncertainty, such as confidence intervals where appropriate; and
  • the practical importance of the observed differences.

Interpretive limitation of the synthetic examples: The four synthetic tables above deliberately contain no sampling variability, standard errors, confidence intervals, or sample sizes. They therefore illustrate patterns of means, not evidence that any population effect is statistically significant. Real data require inference and appropriate model checks before sample patterns are generalized to a population (Moore et al., 2014, 2021).

Research-question dependence: which effect actually matters?

There is no universal rule that the interaction is always more important than the main effects. Importance depends on the question the study was designed to answer.

Overall performance question

Question: “Which training format performs better on average across the workload conditions represented in this study?”

Relevant effect: A training main effect may be directly relevant.

Dependency question

Question: “Does the effectiveness of training depend on workload?”

Relevant effect: The interaction is the focal effect.

Specific-condition question

Question: “Which training format performs better specifically under high workload?”

Relevant effect: The conditional or simple training effect at High workload.

Hayes (2022) frames moderation around questions of when or under what circumstances an effect occurs. Factorial models therefore become especially valuable when theory predicts that a relationship is contingent rather than uniform.

Statistical interpretation should follow that logic:

research question → factorial design → cell means → interaction → conditional effects → marginal effects → substantive conclusion.

Checks before interpreting a factorial model

Interpretation also depends on whether the statistical model is appropriate for the design and data. Conventional ANOVA models make assumptions about the error structure, including independence and a common error variance, with Normal-error modeling underlying standard ANOVA inference (Moore et al., 2014, 2021).

Researchers should therefore examine the design and model rather than treating the ANOVA table as self-validating. Randomization and the structure of the experiment affect what causal conclusions are defensible, while graphical and residual checks help assess whether the fitted model provides an adequate description of the data (Moore et al., 2014, 2021).

The substantive interaction should also be specified in relation to the research problem rather than discovered by indiscriminately searching many possible effects. Adams and Lawrence (2018) emphasize connecting factorial hypotheses and analyses to the intended design and research questions.

Common interpretation mistakes

Mistake 1: “There is no main effect, so the factor does not matter.”

Pattern 2 shows why this is unsafe. Large conditional effects can cancel when averaged.

Mistake 2: “The main effect is significant, so the effect is the same everywhere.”

A main effect is an average. Pattern 3 shows that a factor can have a clear marginal difference while its conditional effect still varies substantially.

Mistake 3: “The lines do not cross, so there is no interaction.”

Nonparallel lines can represent an interaction even without reversal. Crossing is a special and potentially important interaction pattern, not a requirement for interaction (Adams & Lawrence, 2018; Hayes, 2022).

Mistake 4: “One simple effect is significant and the other is not, so the interaction is significant.”

Different significance decisions do not establish that the effects differ. Test the interaction directly (Hayes, 2022).

Mistake 5: “A significant interaction means ignore all main effects.”

A significant interaction means that marginal averages must be interpreted in light of heterogeneous conditional effects. Whether the marginal effect remains substantively useful depends on the research question.

Mistake 6: Reporting only the ANOVA table

The ANOVA test does not reveal the full pattern. Report or visualize cell means and describe the conditional differences that produce the interaction.

A practical interpretation workflow

When interpreting a factorial ANOVA interaction, use this sequence:

  1. State the research question. Decide whether the main interest is an overall factor difference, conditional difference, or interaction.
  2. Inspect the cell means. Understand the actual pattern before interpreting inferential tests.
  3. Calculate marginal means. These show the descriptive main-effect patterns.
  4. Calculate simple differences. Compare one factor within each level of the other.
  5. Compare those simple differences. Their difference is the core 2 × 2 interaction contrast (Hayes, 2022).
  6. Plot the interaction. Look for parallelism, changing separation, and possible reversal.
  7. Evaluate statistical uncertainty. Use the factorial model to determine whether the observed interaction is distinguishable from sampling variation.
  8. Probe a supported interaction where the research question requires it. Estimate and interpret the relevant conditional/simple effects rather than relying only on the interaction p-value (Hayes, 2022).
  9. Return to the main effects. Decide whether the marginal averages remain useful given the conditional pattern.
  10. Interpret magnitude and practical relevance. Statistical significance alone does not describe whether the observed differences matter for the research problem.

The key distinction

A main effect summarizes how the outcome differs across levels of one factor after averaging across the other factor.

An interaction effect asks whether that difference itself changes across levels of the other factor (Moore et al., 2014, 2021; Sekaran & Bougie, 2016).

When an interaction is present, the marginal means can remain mathematically correct while becoming substantively incomplete. That is why understanding how to interpret an interaction requires looking beyond the factorial ANOVA p-value to the cell means, interaction plot, simple effects, effect magnitudes, uncertainty, and—above all—the research question.

References

Adams, K. A., & Lawrence, E. K. (2018). Research methods, statistics, and applications (2nd ed.). SAGE Publications.

Dmitrienko, A., & Koch, G. G. (Eds.). (2017). Analysis of clinical trials using SAS: A practical guide (2nd ed.). SAS Institute.

Hayes, A. F. (2022). Introduction to mediation, moderation, and conditional process analysis: A regression-based approach (3rd ed.). The Guilford Press.

Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2

Moore, D. S., McCabe, G. P., & Craig, B. A. (2014). Introduction to the practice of statistics (8th ed.). W. H. Freeman.

Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). Macmillan Learning.

Sekaran, U., & Bougie, R. (2016). Research methods for business: A skill-building approach (7th ed.). John Wiley & Sons.

Need help with a similar research question?

Share a short, non-confidential summary of your study and the decision you need to make.

Send an enquiry