By the Wavefield Research team · Published Aug 26, 2026
A Likert scale asks people to rate agreement with a statement across ordered, labeled options — classically five, from Strongly disagree to Strongly agree. It is the workhorse of attitude measurement: easy to answer, comparable across respondents, and countable. The craft is in the labels, the number of points, and knowing what statistics the answers can honestly support.
A single rated statement is technically a Likert-type item; the true Likert scale, as Rensis Likert designed it in 1932, is the summed score across several such items. Most modern usage calls the item format a Likert scale, and this guide follows that convention. The distinction still matters in one place: reliability statistics like Cronbach's alpha describe the summed multi-item scale, not a single rated statement.
A Likert scale question presents statements, not questions, each rated on the same labeled scale — usually as a grid. Grids are efficient and comparable, and they carry two design duties: rotate the rows (so order effects average out across respondents) and keep the grid short (past six or seven rows, attention collapses and straightlining begins).
Where the format sits among the other question types — single choice, ranking, matrix, open ends — is covered in survey question types, with examples.
How much do you agree or disagree with each statement?
| Strongly disagree | Disagree | Neither | Agree | Strongly agree | |
|---|---|---|---|---|---|
| The checkout process was easy to complete | |||||
| The delivery arrived when promised | |||||
| Customer support resolved my issue quickly |
Rows rotate to fight order bias; identical answers down a column on every row is what straightline detection flags.
The agreement version is the famous Likert scale example, but forcing every question onto agreement labels adds a translation step for respondents. Use the label family that matches what you are measuring:
| Construct | 5-point labels |
|---|---|
| Agreement | Strongly disagree · Disagree · Neither agree nor disagree · Agree · Strongly agree |
| Satisfaction | Very dissatisfied · Dissatisfied · Neither · Satisfied · Very satisfied |
| Frequency | Never · Rarely · Sometimes · Often · Always |
| Importance | Not at all important · Slightly important · Moderately important · Very important · Extremely important |
| Likelihood | Very unlikely · Unlikely · Neither · Likely · Very likely |
| Quality | Very poor · Poor · Fair · Good · Excellent |
Seven-point versions insert “Somewhat” steps (Somewhat agree / Somewhat disagree); ten- and eleven-point scales (like NPS) leave the middle unlabeled and belong to a different family — rating scales — with their own rules.
A 3-point Likert scale (Disagree / Neutral / Agree) collapses mild and strong views into one bucket. You lose the discrimination that makes trend lines and segment comparisons interesting. Defensible only where space is brutal — SMS surveys, kiosk screens.
A 5-point Likert scale has enough range to separate strong from mild views, and few enough points that every one earns a label. Most benchmarked instruments (CSAT among them) live here, which makes your numbers comparable to published ones.
A 7-point Likert scale adds a Somewhat step on each side. When answers cluster at one end — satisfaction usually skews positive — those extra steps spread the pile-up so movement stays visible. Beyond 7, added precision is imaginary: respondents cannot reliably tell an 8 from a 9.
The midpoint question divides practitioners more than the point count. Removing it (a 4-point forced choice) makes sense when a lean matters more than precision — a screener that must sort people into camps. For measurement, keep it: genuinely neutral respondents exist, and a scale that denies them an honest answer converts neutrality into noise, not signal. What outranks both decisions: never change the scale mid-tracker. A wave that moves from 5 to 7 points breaks every trend it touches, which is why wave-over-wave platforms enforce identical instruments (ours does this structurally).
“The product is affordable and reliable” is two questions wearing one row — a respondent who finds it affordable but flaky has no honest answer. Split double-barreled items, always.
Fully labeled scales are answered more consistently than numeric scales with labeled endpoints, because respondents stop inventing their own meaning for the middle points. If you must go numeric (7+ points), anchor both ends clearly.
Reversed items (“I would not recommend this”) were once standard advice for catching inattention. In practice they mostly catch honest respondents misreading the reversal. Keep statements pointing the same way and use a dedicated attention check instead.
Asking about frequency on an agreement scale (“I often contact support — agree or disagree?”) forces a translation step that adds noise. Ask frequency questions on frequency labels.
Two negative options, a midpoint, and two positive options. A scale with three shades of positive and one negative is a push poll with extra steps, and any reviewer will spot it.
The short answer: yes, you can use means on Likert scale data for comparison, and no, you should not use them as absolute claims. Likert responses are ordinal — coded 1 to 5, ordered but not guaranteed evenly spaced — so a mean of the codes is, strictly, an average of labels. In practice, means on 5-point Likert items track the technically proper alternatives closely, which makes them serviceable for comparing segments or tracking waves, provided differences are significance-tested. What a Likert mean cannot honestly do is stand alone as a quantity: “satisfaction is 3.8 out of 5” invites ratio thinking the scale cannot support (a 4 is not twice a 2), and a polarized split can hide behind a middling average. Report the distribution and the top-2-box share alongside any Likert mean, and treat the mean as a comparison device, not a measurement.
Look this up and you will find the top sources flatly contradicting each other: statistics texts say means are inappropriate for ordinal data, use the median; practitioner blogs say use the mean, everyone does; vendor guides duck the fight and recommend top-2-box percentages. Each is right about something, and none tells you when the other is.
The technical objection is real: Likert codes are ordered, not evenly spaced. Averaging 1-to-5 codes assumes the psychological distance from “Agree” to “Strongly agree” equals the distance from “Neither” to “Agree” — an assumption nobody has verified for your respondents. And a mean of 3.2 can hide wildly different realities: everyone near the middle, or a polarized split with almost nobody at 3.
The practitioner defense is also real: for comparing groups or tracking change, means on 5-point items behave almost identically to the technically proper alternatives in simulation after simulation. The mean is not lying about direction — a segment at 4.1 really does agree more than a segment at 3.4.
The honest synthesis is about use, not dogma. Means are defensible for comparison — between segments, between waves, tested for significance. Means are misleading as absolute claims: reporting “satisfaction is 3.8 out of 5” as if 3.8 were a quantity invites exactly the ratio thinking the data cannot support (a 4 is not “twice” a 2). And a mean without its distribution can actively deceive when responses polarize.
That is why serious reporting shows three things together: the full distribution (the percentage at each point), the top-2-box share (a percentage, always defensible, resistant to the spacing objection), and the mean for compact comparison — with any difference you plan to celebrate tested for significance first. That combination is what our reports compute by default: distributions and means on every topline and crosstab, with significance letters on honest effective bases.
A detail every guide skips: if your sample skews (young, online, engaged), every Likert distribution inherits that skew. Weighting to population targets corrects the shares — and the mean — the same way it corrects a vote share. Unweighted Likert reporting from a skewed sample is a precise summary of the wrong population. How raking works: survey weighting, explained.
The Likert grid's efficiency is also its failure mode: a bored respondent picks one column and rides it down every row. Quality checking should flag straight-line patterns on multi-row grids (ours marks them automatically, alongside speeders and duplicates — flagged for review, never silently deleted). A grid full of straightliners produces beautiful, worthless means.
The classic agreement version: Strongly disagree, Disagree, Neither agree nor disagree, Agree, Strongly agree. The same structure adapts to other constructs — Very dissatisfied to Very satisfied for satisfaction, Never to Always for frequency, Not at all important to Extremely important for importance. The rule is two negative points, a genuine midpoint, and two positive points, all labeled.
A scale with the midpoint removed — Strongly disagree, Disagree, Agree, Strongly agree — forcing respondents to lean one way. Use it when fence-sitting would dodge the question entirely (screening for a preference you must know). The cost: people with genuinely neutral views answer essentially at random, which adds noise exactly where you removed the signal. For most measurement, keep the midpoint; neutrality is data.
Quantitative — it produces ordinal data: the categories have a defined order, but the gaps between them are not guaranteed equal (the distance from Agree to Strongly agree may not equal the distance from Neither to Agree). Percentages and counts per category are always defensible; treating the codes as evenly spaced numbers for means is the contested part, which is why serious reporting shows the distribution alongside any average.
Five and seven points are the evidence-backed sweet spot: three points loses discrimination (mild and strong views collapse together), while beyond seven, extra points add precision respondents cannot actually deliver. Five fully labeled points is the safest default; seven helps when you expect answers to cluster at one end. What matters more than the choice: keep it identical across waves, because a scale change breaks every trend line it touches.
Related: survey question types · how to analyze survey results · open-ended survey questions
Brief the agent and it programs the grid: balanced labels, rotated rows, straightline flags, significance-tested reporting. From $99 per study.
Start a study