FESK.COMYour global study desk
Email us
Statistics & Data

Analysing Likert Data: When Averages Work and When They Mislead

The ordinal-versus-interval argument about Likert data has run for decades. The practical answer lies in separating a single item from a multi-item scale, and in presenting results so readers see the real distribution.

Analysing Likert Data: When Averages Work and When They Mislead

A survey report stated: “Average satisfaction with the library was 3.02 out of 5, a moderate level.” The frequency table showed 45% choosing “very dissatisfied” and 43% “very satisfied”, with almost nobody in between. The figure 3.02 describes a student who does not exist. The real story was two groups with completely different library experiences, and that was the finding worth writing up.

Likert data appear in almost every survey thesis, and the argument over whether they can be averaged has lasted for decades. Rather than picking a side, identify what kind of data you actually have and ask the right question of it.

A Likert item and a Likert scale are different things

This is the most important distinction, and the most often ignored.

Single Likert itemLikert scale
What it isOne statement with 5 or 7 response levelsThe sum or mean of several items measuring one construct
NatureOrdinal, few valuesCloser to continuous, many possible values
Suitable descriptionFrequencies, percentages per level, medianMean and standard deviation, if the distribution is reasonable
Suitable analysisRank tests, ordinal regression, dichotomise then logistict-tests, ANOVA, linear regression

A six-item scale scored 1–5 yields a mean with 25 possible values from 1.00 to 5.00. For a reliable scale, methodological research suggests parametric methods are fairly robust. For single items, that argument is much weaker.

When the mean of a single item is still acceptable

In practice, item means are widely reported, especially for quick comparisons across many items. They are defensible if:

  • The distribution is not piled at both ends or polarised, as in the opening example.
  • You show the full distribution alongside the mean.
  • Your conclusions do not depend on exact distances between levels (do not claim “a 0.4-point rise in satisfaction” as if it were a meaningful unit).

Avoid attaching interpretive labels to ranges of means, such as “1.00–1.80 is very low, 1.81–2.60 is low”. These bands are common in theses but have no measurement basis, and they hide the real distribution.

Show the distribution: diverging stacked bars

The best display for several Likert items is usually a diverging stacked bar chart: one bar per item, disagreement extending left, agreement right, the neutral category centred. Readers immediately see which items divide opinion and which attract consensus. Order the items by level of agreement rather than questionnaire order.

In tables, report the percentage at each level, or at least the percentage agreeing (top two levels combined) with the number who answered that item.

Midpoints, “don’t know” and “not applicable”

  • A neutral midpoint may mean no opinion, ambivalence or not understanding the question. Do not read it as “half satisfied”.
  • “Don’t know” and “not applicable” are not points on the scale. If they were coded as 6, recode them to missing before calculating anything.
  • A very high neutral rate on one item tells you something about the item: it may be ambiguous, or respondents may lack experience.

Response styles: comparing across cultures

Respondents differ not only in opinions but in how they use scales. Some groups tend to agree with almost everything (acquiescence), some avoid the extremes, others favour them. When comparing students in two countries or regions, a 0.3-point difference may reflect response style rather than a real difference. To reduce the risk, use scales with both positively and negatively worded items, test measurement invariance across groups before comparing means, and interpret small differences cautiously.

Analysing single items: your options

  1. Two-group comparison: Mann–Whitney with a rank-based effect size. Remember it compares tendency to be larger, not medians.
  2. Dichotomise then logistic regression: “agree” (4–5) versus the rest. Easy to explain but loses information; set the cut-off in advance.
  3. Ordinal regression (ordinal logistic, proportional odds): keeps the ordering, adjusts for several variables and produces interpretable odds ratios. Check the proportional odds assumption.
  4. Rank correlation (Spearman or Kendall) between two items.

Analysing multi-item scales: conditions before summing

Before summing or averaging items you need evidence that they measure one construct: internal consistency, an appropriate factor structure, reversed items correctly scored. If the scale was translated or used in a new setting, check these properties in your own sample. After that, scale scores can be analysed with parametric methods, but still look at the distribution for ceiling and floor effects.

In practice: a Likert analysis checklist

  1. Decide whether each variable is a single item or a multi-item scale score.
  2. Recode “don’t know” and “not applicable” to missing; reverse-score negative items.
  3. For single items: frequency tables and diverging stacked bars; rank tests or ordinal regression.
  4. For scales: check reliability and structure, then describe with mean, standard deviation and distribution.
  5. Do not make “low, moderate, high” mean bands your main conclusion.
  6. If a distribution is polarised, report that as a finding rather than letting the mean hide it.

Next step: pick the three most important Likert items in your survey, draw a diverging stacked bar chart for them and compare it with the story your table of means is telling. If the two stories differ, trust the chart.

Câu hỏi thường gặp

Can you calculate the mean of Likert scale data?

For multi-item scale scores with good reliability, usually yes. For single items, prefer frequencies, medians and ordinal methods, or at least show the distribution alongside the mean.

What is the difference between a Likert item and a Likert scale?

A Likert item is one question with a few response levels; a Likert scale is the sum or mean of several items measuring the same construct and behaves more like a continuous variable.

What chart is best for Likert data?

A diverging stacked bar chart is usually clearest: disagreement to the left, agreement to the right, with items sorted by level of agreement.

When should I use ordinal regression?

When the outcome is a single Likert item or other ordinal variable and you need to adjust for several predictors. It respects the ordering and yields interpretable odds ratios.

Should I classify mean scores into low, moderate and high bands?

Not as a main conclusion. Such bands have no measurement basis and can hide the actual distribution of responses.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com