How to Analyze Likert Scale Data: From Single Items to Composite Scores
Learn how to analyze Likert scale data by distinguishing single ordered items from multi-item composite scores, then matching the measurement interpretation, estimand, study design, and statistical method to the research question.
Researchers often ask a deceptively simple question after collecting a survey: “What statistical test should I use for my Likert data?”
The answer depends on what the data actually represent. A single response such as strongly disagree → strongly agree is not analytically identical to a score created by combining several items. Nor does assigning the numbers 1 through 5 automatically establish equal quantitative distances between response categories. The appropriate analysis therefore starts with the measurement structure and research question, not with a blanket rule about “Likert data.” (Sekaran & Bougie, 2016; Adams & Lawrence, 2018; Meier et al., 2014).
This distinction explains why the question “Is a Likert scale ordinal or interval?” does not have one universally applicable answer in the approved sources. Sekaran and Bougie (2016) explicitly describe the ordinal-versus-interval status of Likert scales as debated: adjacent response categories cannot automatically be assumed equidistant, although Likert scales are commonly treated as interval scales in applied research. Adams and Lawrence (2018) similarly state that treating a Likert-type response as interval requires an equal-interval assumption and note that many, but not all, researchers make that assumption. Lovric (2011), by contrast, includes a more conservative treatment of Likert responses as ordered categorical ratings and emphasizes nonparametric methods for ratings and opinions.
Practical implication: Do not choose a side by slogan. Identify what was measured, what score will actually enter the analysis, and what quantity the analysis is supposed to estimate.
Start With the Data Structure
The first decision is whether you have a single Likert-type item or several items intended to measure a common construct.
A Likert item asks respondents to indicate the strength of agreement or disagreement with a statement using ordered response categories. Sekaran and Bougie (2016), for example, describe the familiar five-point structure ranging from strongly disagree to strongly agree. The categories clearly have an order, but their numerical labels do not by themselves demonstrate that the psychological distance from 1 to 2 is identical to the distance from 4 to 5. (Sekaran & Bougie, 2016).
Several items can instead be combined into a summated or composite score when they are intended to measure the relevant concept and the scoring procedure supports combining them. Sekaran and Bougie (2016) explicitly distinguish item-by-item analysis from summing responses across items and describe the Likert scale as a summated scale in this context. Adams and Lawrence (2018) likewise describe computing total scale scores from multiple items after any necessary recoding.
Single Likert-type item
The observed response remains an ordered category unless an additional quantitative interpretation is deliberately imposed and defended.
(Sekaran & Bougie, 2016; Lovric, 2011).
Multiple-item scale
The analyst may construct a total or other prescribed composite after ensuring that items are scored in a common direction and that combining them is consistent with the measurement procedure.
(Sekaran & Bougie, 2016; Adams & Lawrence, 2018).
A composite is therefore not merely “a Likert item with more numbers.” It is a new score produced from several item responses, and its interpretation depends on how those items were designed, coded, and combined. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018).
Nominal, Ordinal, and Interval-Type Reasoning
The measurement level determines which numerical operations have a clear interpretation.
- Nominal
- Classifies observations into categories without ordering them.
- Ordinal
- Adds meaningful order but does not establish how far apart adjacent categories are.
- Interval
- Adds interpretable numerical differences between values.
(Meier et al., 2014; Sekaran & Bougie, 2016).
Meier et al. (2014) emphasize this distinction for attitude and opinion measurements: ordinal responses can establish that one observation represents more or less of a characteristic, but not necessarily how much more or less. They therefore treat attitude and opinion responses conservatively as ordinal measurements.
Sekaran and Bougie (2016), however, explicitly acknowledge common applied practice in which Likert scales are analyzed as if they were interval measures, permitting averages, standard deviations, and more advanced statistical analyses. They also acknowledge the objection: equal distances between adjacent response levels cannot simply be assumed.
Adams and Lawrence (2018) take a similarly conditional position. They describe a Likert-type response as interval when equal intervals between successive numerical values are assumed, while explicitly noting that this assumption is not accepted by all researchers.
Ask the more useful question: “For this particular item or composite, what interpretation of numerical distance am I willing and able to defend?”
Likert Scale Mean or Median?
The answer again depends on what is being summarized.
For a single ordinal Likert-type item
Frequencies and percentages preserve the response categories directly and show how respondents are distributed across them. Ordinal data can also be summarized by a median because the categories can be ranked. Meier et al. (2014) explicitly demonstrate frequency distributions and medians for ordered attitude responses while rejecting the mean when the measurement is treated strictly as ordinal.
For example, reporting that 18% strongly disagreed, 24% disagreed, 20% were neutral, 27% agreed, and 11% strongly agreed retains information that a single central statistic can conceal.
Under strict ordinal reasoning, therefore, frequencies and the median are natural descriptive summaries. (Meier et al., 2014).
When interval-type reasoning is adopted
A mean answers a different question: it uses the numerical values assigned to categories and therefore treats their differences as quantitatively meaningful. Sekaran and Bougie (2016) note that treating Likert scales as interval is common precisely because it permits calculation of averages and standard deviations. Adams and Lawrence (2018) similarly connect mathematical operations on Likert-type responses to the equal-interval assumption.
Consequently, the choice between a Likert scale mean or median should not be framed as “mean is always wrong” versus “mean is always acceptable.” It depends on the measurement interpretation being used and, importantly, whether the analyst is describing an individual ordered item or a constructed scale score. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018; Meier et al., 2014).
Interpretation note: When interpretation is potentially contentious, reporting the category distribution alongside any numerical summary makes the measurement structure visible rather than hiding it behind a single statistic.
Several Likert Items: Build the Score Before Choosing the Test
If several items are intended to form a scale, analysis should not jump immediately to a t test, ANOVA, or regression.
First determine how the scale itself is scored.
Sekaran and Bougie (2016) explain that responses to multiple Likert items can be examined individually or summed into a total score. They also show why negatively oriented items may need to be reverse-scored before summation so that high values have a consistent substantive meaning across the scale.
Adams and Lawrence (2018) make the same point operationally: before computing a total scale score, identify items whose direction conflicts with the intended interpretation, recode those items, and then combine the appropriately oriented responses. They further distinguish computing a total score from assessing internal consistency; adding items produces the score, whereas an internal-consistency analysis evaluates relationships among the constituent items. (Adams & Lawrence, 2018).
Composite-score workflow
- Check item direction.
- Recode where required.
- Review measurement and reliability.
- Calculate the prescribed composite.
- Analyze the resulting score.
In compact form: item direction → recoding where required → measurement/reliability review → prescribed composite calculation → statistical analysis of the resulting score.
A multi-item composite should not be created merely because several questions happen to share the same 1–5 response options. The substantive and measurement rationale for combining them comes first. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018).
Decision Table: What Analysis Matches the Question?
| Data structure | Research question | Key considerations | Candidate analysis | Main caution |
|---|---|---|---|---|
| One Likert-type item | What responses were most common? | Ordered categories; distances need not be equal. | Frequencies, percentages, mode; median if ordinal ordering is used. | A numerical code does not itself establish interval measurement. (Meier et al., 2014; Sekaran & Bougie, 2016). |
| One Likert-type item | Where is the typical response? | Median requires ordering; mean additionally requires quantitative-distance reasoning. | Median under ordinal reasoning; mean only under a defended interval-type interpretation. | Mean and median embody different measurement assumptions. (Meier et al., 2014; Adams & Lawrence, 2018). |
| One ordinal item, two independent groups | Do response rankings/distributions differ between groups? | Independence and ordered outcome; target is rank/distribution comparison rather than a mean difference. | Mann–Whitney/Wilcoxon rank-sum framework. | Do not describe a rank test as a test of means. (Adams & Lawrence, 2018; Moore et al., 2021). |
| One ordinal item, three or more independent groups | Do response rankings/distributions differ across groups? | Ordered outcome and independent groups. | Kruskal–Wallis. | It is a rank-based procedure, not simply “nonparametric ANOVA” with an identical estimand. (Adams & Lawrence, 2018; Moore et al., 2021). |
| Ordered Likert-type outcome with predictors | How does an ordered response vary with predictors? | Outcome remains categorical and ordered. | An ordinal-response model such as a proportional-odds model may match the outcome structure. | Model interpretation concerns ordered response categories rather than an ordinary mean difference. (Lovric, 2011). |
| Defensible multi-item composite | Do two independent groups differ in mean scale score? | Composite construction, measurement interpretation, distribution, unusual observations, and assumptions. | Independent-samples t procedure when the target is a mean difference and quantitative treatment is justified. | Do not select the test solely because the source items were Likert items. (Sekaran & Bougie, 2016; Moore et al., 2021). |
| Defensible multi-item composite | Do three or more groups differ in mean scale score? | Same issues, plus the multiple-group design. | ANOVA when the estimand is a difference among means and quantitative treatment is justified. | ANOVA answers a mean-based question; Kruskal–Wallis answers a rank-based question. (Sekaran & Bougie, 2016; Moore et al., 2021). |
| Defensible multi-item composite | How is the scale score associated with one or more predictors? | Whether the composite is being treated quantitatively; regression assumptions and diagnostics. | Linear regression when a quantitative conditional-mean model matches the question. | Treating a composite as quantitative is an analytic decision that should be defensible. (Sekaran & Bougie, 2016). |
When Do Rank and Nonparametric Methods Become Relevant?
Nonparametric methods become especially relevant when the measurement itself supports ordering but not meaningful quantitative distances. Adams and Lawrence (2018) explicitly place rank-based procedures in their treatment of ordinal data: Mann–Whitney-type procedures address two independent groups, Wilcoxon procedures address relevant dependent-group settings, and Kruskal–Wallis addresses three or more independent groups.
Moore et al. (2021) likewise distinguish rank procedures from t procedures and emphasize that they do not simply reproduce the same analysis after replacing the word “parametric” with “nonparametric.” The hypotheses and interpretation of the rank procedures need to be considered.
Design matters as well as measurement level. Two independent groups, paired responses, three independent groups, and an ordered outcome with several predictors are different statistical problems even if every outcome was collected using the same five response categories. (Adams & Lawrence, 2018; Moore et al., 2021).
Does Likert Data Mean You Must Use Nonparametric Tests?
No universal rule of that form is supported across the approved sources.
There is genuine methodological tension in the sources. Lovric (2011) presents a strong case for nonparametric reasoning with ratings and Likert data, emphasizing their ordered categorical character and the weaker distributional requirements of nonparametric methods. Meier et al. (2014) likewise treat attitude and opinion responses conservatively as ordinal measurements.
Sekaran and Bougie (2016), however, explicitly state that although the ordinal-versus-interval classification of Likert scales is debated, Likert scales are commonly treated as interval measures in applied analysis. Adams and Lawrence (2018) similarly state that many—but not all—researchers assume equal intervals for Likert-type scales.
The statement “Likert data are always nonparametric” is therefore too simplistic as a summary of these sources. Equally, the opposite statement—“Likert responses are interval, so parametric methods are always fine”—would also erase the measurement issue the sources explicitly identify. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018; Meier et al., 2014; Lovric, 2011).
A defensible decision identifies whether the analysis concerns a single ordered response or a constructed composite, states the measurement interpretation being adopted, and then chooses a procedure whose estimand matches the research question.
When Do t Tests and ANOVA Make Sense?
A t test is fundamentally a procedure for a question about means, not a generic procedure for “survey data.” Sekaran and Bougie (2016) describe independent-samples t testing in terms of comparing two groups on an interval- or ratio-scaled dependent variable. They similarly describe ANOVA as a method for comparing mean values across more than two groups when the dependent variable is treated as interval or ratio scaled.
That makes the estimand crucial.
Mean-based question
“Do the groups differ in their mean composite satisfaction score?”
A t test or ANOVA may match the question when the composite has been defensibly constructed and treated quantitatively and the relevant model conditions are credible.
(Sekaran & Bougie, 2016; Moore et al., 2021).
Ordered-response question
“Do the groups differ in how their ordered response categories are distributed or ranked?”
A rank-based or ordinal-response analysis may better match the outcome and estimand.
(Adams & Lawrence, 2018; Lovric, 2011).
Switching from a t test to Mann–Whitney, or from ANOVA to Kruskal–Wallis, is therefore not merely a technical change designed to make an assumption problem disappear. It can change the statistical question being answered. (Moore et al., 2021).
What About Regression?
The same principle applies to regression: the outcome helps determine the model.
Sekaran and Bougie (2016) place ordinary regression among methods involving metric dependent variables and distinguish it from models used for nonmetric dependent variables. Thus, treating a multi-item composite as a quantitative outcome can support a linear-regression formulation when that interpretation and the model assumptions are appropriate. (Sekaran & Bougie, 2016).
A single ordered response creates a different problem. Lovric (2011) notes extensions of logistic modeling to categorical responses, including proportional-odds models. Such an ordinal-response model preserves the ordering of the categories without requiring an ordinary linear model of their numerical codes.
Linear regression
Targets conditional mean differences on a quantitative outcome scale.
Ordinal-response model
Addresses probabilities or odds across ordered outcome categories.
These are not interchangeable interpretations. (Sekaran & Bougie, 2016; Lovric, 2011).
Missing Item Responses: Decide Before Computing the Composite
Missing responses require an explicit rule because they can change both the scale score and the sample entering subsequent analyses.
Sekaran and Bougie (2016) warn that simply ignoring missing responses can reduce the available sample and can bias results when missingness is not completely random. They describe multiple possible approaches to blank responses and explicitly note that each has advantages and disadvantages rather than presenting one universally correct solution.
Adams and Lawrence (2018) likewise emphasize that the analyst must first establish how missing values are represented and handled before computing scale scores.
No universal 80% rule is established by the approved sources. The sources do not establish a general rule such as “calculate the composite whenever 80% of items are answered.” Such a threshold should therefore not be manufactured as a general Likert-analysis rule.
Instead:
- Follow the validated instrument's scoring instructions when they specify how incomplete responses are handled.
- Identify how many responses are missing and whether missingness changes the usable sample.
- Reverse-score applicable items before computing any composite.
- State explicitly how incomplete item sets were treated.
- Avoid silently converting a missing response into a substantive response category.
The key principle is transparency: the composite-scoring rule should be defined before inferential results are interpreted, because different missing-item decisions can change who receives a score and what that score represents. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018).
Common Likert Analysis Mistakes
1. Treating every Likert variable as the same kind of data
A single ordered item and a multi-item composite are different analytical objects. Sekaran and Bougie (2016) explicitly distinguish item-level analysis from summated scoring.
2. Assuming that numeric codes prove equal intervals
Coding strongly disagree as 1 and strongly agree as 5 establishes an ordering convention; it does not by itself demonstrate equal psychological distances between adjacent categories. (Sekaran & Bougie, 2016; Meier et al., 2014).
3. Saying “Likert data are ordinal, therefore always use nonparametric tests”
The approved sources do not support this as a universal rule. They contain both conservative ordinal treatments and explicit descriptions of common interval-type treatment of Likert scales. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018; Lovric, 2011).
4. Saying “everyone treats Likert scales as interval, so it does not matter”
That is also too strong. Sekaran and Bougie (2016) explicitly call the classification debated, and Adams and Lawrence (2018) state that not all researchers accept the equal-interval assumption.
5. Averaging items before checking their direction
A negatively worded item can reverse the substantive meaning of the composite unless it is appropriately recoded before summation. Both Sekaran and Bougie (2016) and Adams and Lawrence (2018) explicitly demonstrate this issue.
6. Choosing Mann–Whitney or Kruskal–Wallis as automatic substitutes
Rank procedures address rank-based hypotheses and must still match the independent/dependent-group structure of the study. They should not be selected merely because the word “Likert” appears in the questionnaire. (Adams & Lawrence, 2018; Moore et al., 2021).
7. Choosing t tests or ANOVA without identifying the estimand
These procedures address mean comparisons. If the scientific question is about ordered response probabilities or ranks instead, another analysis may better match the intended inference. (Sekaran & Bougie, 2016; Lovric, 2011).
8. Creating a composite simply because several questions use the same response options
Common response coding does not itself establish that the items constitute one coherent measure. Composite construction should follow the intended measurement structure, correct item direction, and relevant reliability and validity evidence. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018).
9. Hiding the response distribution behind a mean
For a single ordinal item, frequencies and percentages directly display how respondents used the available categories. A central summary alone can remove that information. (Meier et al., 2014).
10. Handling missing items silently
Deleting incomplete cases, ignoring missing cells, or substituting values can change the analysis. The scoring and missing-data rule should therefore be stated rather than left to an undocumented software default. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018).
A Practical Decision Sequence
When deciding how to analyze Likert scale data, work through the problem in this order.
-
Identify the unit being analyzed.
Is it one Likert-type item or a score constructed from several items? (Sekaran & Bougie, 2016).
-
Identify the measurement claim.
For an individual item, the ordering is explicit but equal distances are debatable. For a composite, determine whether quantitative interpretation of the resulting score is defensible for the intended analysis. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018; Meier et al., 2014).
-
Prepare the score correctly.
Reverse-code items where necessary, establish how missing items will be handled, and compute the composite according to the intended scoring procedure. (Sekaran & Bougie, 2016; Adams & Lawrence, 2018).
-
State the estimand.
Are you interested in category frequencies, a median response, a mean score, rank differences, or ordered-category probabilities? Different analyses answer these different questions. (Meier et al., 2014; Adams & Lawrence, 2018; Moore et al., 2021; Lovric, 2011).
-
Match the design.
Independent groups, paired observations, several independent groups, and multivariable prediction require different procedures even when the outcome measurement is otherwise identical. (Adams & Lawrence, 2018; Moore et al., 2021).
-
Check the assumptions and interpretation of the selected method.
Calling a method “parametric” or “nonparametric” is not enough. The method must answer the intended research question for the data structure actually analyzed. (Moore et al., 2021).
The Bottom Line
There is no single statistical test for Likert scale data.
Single Likert-type item
Treating responses as ordered categories is the conservative starting point: frequencies, percentages, and ordinal summaries preserve the observed measurement structure. Rank-based procedures or ordinal-response models become candidates when inferential questions concern ordered outcomes.
(Meier et al., 2014; Adams & Lawrence, 2018; Lovric, 2011).
Several items intended to form a scale
First establish the scoring direction and composite construction. A properly constructed total or composite score creates a different analytical object from any individual response item. When that score is defensibly treated quantitatively and the research question concerns means or conditional means, t tests, ANOVA, or linear regression can become relevant candidates.
(Sekaran & Bougie, 2016; Adams & Lawrence, 2018).
Most defensible workflow: question → item or composite → measurement interpretation → scoring and missing-data rule → estimand → design → statistical method → diagnostics → interpretation.
That approach avoids both extremes: treating every Likert response as automatically continuous and treating every analysis involving Likert items as automatically nonparametric.
References
Adams, K. A., & Lawrence, E. K. (2018). Research methods, statistics, and applications (2nd ed.). SAGE Publications.
Lovric, M. (Ed.). (2011). International encyclopedia of statistical science. Springer. https://doi.org/10.1007/978-3-642-04898-2
Meier, K. J., Brudney, J. L., & Bohte, J. (2014). Applied statistics for public and nonprofit administration (9th ed.). Cengage Learning.
Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). Macmillan Learning.
Sekaran, U., & Bougie, R. (2016). Research methods for business: A skill-building approach (7th ed.). John Wiley & Sons.
Need help with a similar research question?
Share a short, non-confidential summary of your study and the decision you need to make.