A doctoral student collected 420 questionnaires, ran a factor analysis and discovered that three reverse-worded items loaded on the wrong factor. On inspection, all three used double negatives: “I do not believe my job lacks meaning.” Respondents skimming on a phone read them the wrong way round. No statistical technique can rescue that data.
A questionnaire is a measuring instrument. A miscalibrated scale gives the wrong weight however many times you use it. Most questionnaire faults can be caught by reading each item aloud and asking: what would an ordinary person, in a hurry, reading this on a six-inch screen, think it means?
The nine most common item errors
| Error | Example | Fix |
|---|---|---|
| Double-barrelled | My lecturer is enthusiastic and clear | Split into two items |
| Leading | Do you agree that tuition fees are a burden? | Ask neutrally about degree |
| Double negative | I do not oppose not taking attendance | Rewrite in positive form |
| Overlapping options | 1–3 hours; 3–5 hours | Under 3 hours; 3 to under 5 hours |
| Missing option | Yes/No only, where some cannot answer | Add “Not applicable” |
| Vague time frame | Do you often read books? | In the past 7 days, on how many days did you read a book? |
| Jargon | Rate your learner autonomy | Describe concrete behaviours |
| Built-in assumption | Which study app do you use? | First ask whether they use one |
| Blunt sensitive item | Have you ever cheated in an exam? | Normalise, assure anonymity, place late |
The vague time frame is the most dangerous because it is silent. “Often” means daily to one person and monthly to another. Two identical ticks on a form may describe behaviour that differs thirtyfold.
Likert scales: five or seven points, midpoint or not
There is no single right answer, but a few practical rules hold:
- Five points are easier on mobile and for respondents unused to surveys. Seven points offer slightly finer resolution with highly educated respondents.
- If you use a published scale, keep the original number of points so results remain comparable.
- A neutral midpoint makes sense where genuine neutrality exists. Removing it forces a side and can increase item non-response.
- Label every point in words, not just the ends, so the steps are understood consistently.
- Keep direction consistent throughout: 1 always means “strongly disagree”.
On reverse-worded items: they are meant to catch straight-lining, but in translation they often produce awkward negatives. If you use them, reverse the meaning (“My work bores me”) rather than inserting “not”.
Adapting a scale from another language: back-translation is not enough
The familiar approach is forward translation followed by back-translation into the source language for comparison. It catches errors of meaning but not stiff, unnatural target-language wording. A more robust process:
- Two independent forward translators, one with subject expertise and one without.
- Reconcile the versions, recording each disagreement and its resolution.
- Back-translation by someone who has not seen the original.
- Expert review for semantic and cultural equivalence.
- Cognitive interviews with five to ten members of the target population.
Check the licence before translating; some instruments are copyrighted and require permission or a fee.
Cognitive interviewing: the step almost everyone skips
Sit beside a respondent while they complete the questionnaire and either ask them to think aloud or probe after each item: “What is this question asking you?”, “Why did you choose 4?”, “Was there any item where you had to guess?”. Five to ten such sessions usually uncover more than a 50-person pilot, because you hear how people interpret items rather than only seeing the numbers they pick.
A real-world example: “I receive adequate support from my university.” Cognitive interviews revealed that half the participants thought of financial aid and half of academic advising. The item measured two different things depending on the reader.
Length, order and attention checks
- In voluntary online surveys, anything much beyond fifteen minutes on a phone tends to see break-off climb in the second half. Time it on a phone, not a laptop.
- Open with easy, non-sensitive items; put demographics and sensitive questions at the end.
- Where possible, place the outcome measure before predictors so earlier answers do not prime later ones.
- Include one or two attention checks (“For this item, please select 2”), but not so many that respondents feel policed.
- Record completion time; responses faster than a third of the median deserve scrutiny.
What you do not need
You do not need to ask about variables you have no plan to analyse “in case they are useful later”. Every unnecessary item lowers answer quality on the items that matter. You do not need to write a new scale when a validated one exists for your construct. And you do not need an open-ended box after every section unless you have a plan to code the responses.
A pre-launch checklist
- Every item maps to a variable in your model or a specific research question.
- You have read everything aloud and removed double-barrelled, double-negative and leading items.
- Every behavioural item has an explicit time frame.
- Response options are mutually exclusive and exhaustive.
- At least five cognitive interviews with the target population are complete.
- A pilot of 30–50 respondents has checked timing and preliminary reliability.
- Exclusion rules are written down before you look at the data.
Do this today: print your questionnaire, hand it to someone from your target population who does not study your subject, sit beside them and note every place they pause for more than three seconds. Each pause marks an item to rewrite.
Câu hỏi thường gặp
Should I use a 5-point or 7-point Likert scale?
Both are acceptable. Five points are easier on mobile devices, while seven offer slightly more discrimination. If you use a published scale, keep the original number of points.
What is a double-barrelled question?
It is an item that asks about two things at once, such as “my lecturer is enthusiastic and clear”. Respondents who agree with only one part cannot answer accurately, so split it into two items.
How do you translate a validated questionnaire?
Use two independent forward translations, reconcile them, back-translate, have experts review equivalence and run cognitive interviews with the target group. Back-translation alone misses awkward wording.
How long should an online survey be?
For voluntary online surveys, aim to keep completion under roughly fifteen minutes on a phone. Time it realistically and cut items that do not serve a research question.
How many people do I need to pilot a questionnaire?
A pilot of about 30–50 respondents checks timing, missing data and preliminary reliability. Add five to ten cognitive interviews to check how items are actually understood.