A doctoral student posts her survey link in three large social media groups, promising a small gift voucher to the first 50 respondents. Overnight, responses pass 1,200. When she opens the data, more than 300 people have completed a 60-item questionnaire in under two minutes, dozens of open-ended answers are identical word for word, and a cluster of responses from the same network address arrived within 15 minutes. It is hard to fully trust the rest.
Online surveys are cheap, fast and far-reaching. Those same features attract poor-quality responses. Data quality has to be designed in before the link goes out; it can never be fully rescued afterwards.
Four kinds of problem response
| Type | Warning signs | Typical cause |
|---|---|---|
| Careless responding | Very fast completion, same option across a whole page, failed attention checks | Real people not reading, often for the reward or from fatigue |
| Duplicates | Same demographics and device, submitted minutes apart | One person submitting several times to improve their odds |
| Bots and automated responses | Nonsense or copied open answers, oddly uniform timings, bursts of submissions | Public links with cash rewards |
| Ineligible respondents | Do not meet criteria but answer anyway; contradictions between items | Screening criteria that are too obvious |
Prevention at the design and distribution stage
- Do not advertise a generous reward on an open link. If you offer one, use a prize draw after quality checks and say clearly that invalid responses are not eligible.
- Distribute through controlled channels where possible: departmental mailing lists or personalised single-use links.
- Hide your screening criteria. Instead of “Are you a primary school teacher?”, ask “What is your current job?” with several options.
- Switch on platform features: block repeat submissions from the same browser, add bot verification such as CAPTCHA, record page timings. These reduce the problem but do not remove it.
- Keep it short. Careless responding rises with length, and the last items of a long questionnaire are usually the worst quality.
Attention checks, used properly
Attention checks come in two forms. Instructed items: “For this item, please select Strongly disagree.” Bogus items: “I have never used a phone in my life.” Careful readers answer correctly; random clickers are likely to fail.
Some practical points:
- Use two or three, spread out, and exclude only respondents who fail two or more, so a single slip does not cost you an honest participant.
- Avoid trick questions or wordplay; the aim is to detect people who are not reading, not to test intelligence.
- Place them inside long scale blocks, where careless answering is most likely.
- With professional or expert samples, checks can feel patronising; consider relying on timing and response patterns instead.
Indicators after collection
- Completion time: set a plausible minimum, for instance about two seconds per short scale item. Someone finishing 60 items in 90 seconds almost certainly did not read them.
- Straightlining: the same option for every item in a block, especially suspicious when the block contains reverse-worded items.
- Internal inconsistency: strongly agreeing with an item and its opposite; age 19 with 15 years of teaching experience.
- Open-ended answers: nonsense, off-topic text or answers identical to other responses. A short open question in the middle of the survey is a surprisingly effective bot detector.
- Technical traces: several responses from the same network address in a short window, identical device profiles. Dormitories and offices can share addresses, so never exclude on this sign alone.
Set exclusion rules before you see the results
This is the step people skip. If you run the analysis, dislike the result and then decide which exclusion rule to apply, you have introduced analytic flexibility that can distort your findings. Write the rules into your analysis plan before opening the data. For example: “Exclude responses meeting any of three conditions: failing two or more attention checks; completion time under 40 per cent of the median; open-ended answers judged meaningless by two independent raters.”
Then report fully: how many responses were received, how many were excluded under each rule (one response can meet several), and whether the main results change when excluded responses are kept. That sensitivity analysis reassures readers that cleaning did not drive the conclusions.
Deployment checklist
- The shortest questionnaire possible; median completion time estimated by piloting with 5 to 10 people.
- Two or three attention checks and one short open question midway.
- Screening questions that do not reveal the criteria.
- Duplicate blocking, bot verification and timing switched on.
- Rewards through a post-check prize draw, never first-come-first-served.
- Exclusion rules written into the analysis plan.
- Hourly monitoring of submissions in the first days to catch unusual bursts.
Before launch, take your own survey once as fast as someone who only wants the reward. Note your time and the items you could answer without reading. Those are the spots that need an attention check or a cut.
Câu hỏi thường gặp
How do you detect careless responses in an online survey?
Combine several indicators: very short completion times, the same answer across a whole block, failed attention checks, contradictory answers and meaningless open responses. Avoid excluding anyone on a single sign alone.
What is an attention check question?
It is an item inserted to see whether respondents are reading, for example one that asks them to choose a specific option. Use two or three and exclude only those who fail at least two.
Should I offer incentives for completing a survey?
Incentives raise response rates but also attract low-quality responses when the link is public. A safer approach is a prize draw held after quality checks, with a clear statement that invalid responses are not eligible.
Can I decide exclusion criteria after looking at my results?
You should not, because choosing criteria after seeing results introduces flexibility that can bias conclusions. Write the rules in advance and report a sensitivity analysis that keeps the excluded responses.
How can I stop bots from answering my survey?
Use the platform’s bot verification, avoid public links with cash rewards, include a short open question to catch nonsense answers and watch for sudden bursts of submissions. No single measure is enough, so combine them.