A team pilots an intervention with 15 people per arm, observes an effect of d = 0.8, and uses it to size the main trial: 26 per arm. The main trial finds nothing. Meanwhile, the pilot data held a detail nobody acted on: 6 of the 15 people in the intervention arm skipped session three because it clashed with their shifts.
That is the classic misreading of a pilot study. People stare at the effect estimate and ignore what the pilot can actually tell them — whether the main study will run at all.
The right question: “can we do this?”, not “does it work?”
A pilot is a small-scale rehearsal of the main study, designed to find what will break before you commit the full budget and timeline. Questions it can answer:
- How many people can be recruited per month, through which channels, and what share of eligible people agree?
- How many stay until the final measurement?
- Do participants adhere to the intervention, and do staff deliver it as intended?
- Are the measures understandable, and can they be completed in the expected time?
- Do data entry, storage and analysis pipelines work end to end?
Terminology helps here. A feasibility study is the broad umbrella: should and can this be done? A pilot study is a type of feasibility study that runs all or part of the planned protocol. A pretest checks only the instrument. CONSORT has a dedicated extension for pilot and feasibility trials, worth following when you write up.
Set progression criteria before you start
A pilot without pre-set success criteria will always be declared “feasible”. Traffic-light criteria keep you honest:
| Indicator | Green: proceed | Amber: modify, then proceed | Red: stop and redesign |
|---|---|---|---|
| Recruits per month | ≥ 12 | 8–11 | < 8 |
| Retention to final follow-up | ≥ 80% | 65–79% | < 65% |
| Attending ≥ 75% of sessions | ≥ 70% of participants | 50–69% | < 50% |
| Primary outcome data complete | ≥ 90% | 75–89% | < 75% |
These thresholds are illustrations; yours should be worked backwards from the main study’s target sample and timeline. If you need 200 participants in 12 months, 12 recruits a month at one site tells you that you need at least two sites.
The biggest mistake: powering the main study on the pilot effect
With 15 per group, an observed d = 0.5 has a 95% confidence interval running from roughly −0.2 to 1.2. The data are compatible with a mildly harmful intervention and with a very large benefit. Sizing a trial on the midpoint of that interval is a gamble.
It gets worse through selection. Teams whose pilots show small effects tend to drop the idea; teams with large effects move forward. The main trials that do happen are therefore built on inflated estimates, and end up underpowered.
What to do instead:
- Base the sample size on the smallest effect of practical importance — the difference practitioners would consider worth changing what they do.
- Use the pilot to estimate the outcome’s standard deviation, and plan with an upper confidence limit (such as the 80% upper limit) rather than the point estimate, so you don’t underestimate it.
- Use the pilot to estimate attrition, and inflate the sample size to match.
- For cluster designs, a pilot gives a rough feel for the intra-cluster correlation, but estimates from a handful of clusters are unstable; check against published values.
And don’t run hypothesis tests on effectiveness with pilot data. If you report outcomes, report them descriptively with confidence intervals and say plainly that the pilot was not powered to test them.
How big should a pilot be?
Pilot sample size is driven by the precision you need for feasibility indicators, not by power. Two quick calculations:
- Estimating a proportion: if true retention is about 80%, a sample of 50 gives a 95% confidence interval of roughly ±11 percentage points; a sample of 20 gives about ±18. Ask whether that margin is narrow enough to tell green from red.
- Detecting a problem: if a problem (misreading an item, a device failure) affects 10% of participants, you need about 29 people for a 95% chance of seeing it at least once; for a 5% problem, about 59.
The methods literature offers rules of thumb — around 12 per group to estimate a standard deviation, or about 30 overall. Treat them as starting points, not substitutes for a calculation tied to the indicators you actually need.
Pretesting a questionnaire with cognitive interviews
Before sending a questionnaire to 50 people, sit with 5 to 8 and find out how they answer. Every response passes through four stages: understanding the question, retrieving information, forming a judgement, and mapping it onto the response options. Things can go wrong at any stage.
| Stage | Example probe | Problems it reveals |
|---|---|---|
| Comprehension | “In your own words, what is this question asking?” | Terms read differently, double-barrelled items |
| Retrieval | “What time period were you thinking about?” | Ignored reference periods, impossible recall |
| Judgement | “How did you arrive at that number?” | Guessing, rounding, socially desirable answers |
| Response | “Why did you choose 4 rather than 5?” | Missing options, scale points that blur together |
Two techniques: think-aloud, where people verbalise while answering, and verbal probing, asked after each item or at the end. Many participants find thinking aloud awkward, so targeted probes often yield more. Work in rounds: 5 to 8 people, revise, then a fresh round with new people, stopping when a round turns up nothing substantial. Log findings in a matrix of item × participant × problem type so repeat offenders stand out. Deliberately include the hardest cases: older adults, less confident readers, people answering on phones.
Dry-run the whole data pipeline
After cognitive interviews, field the revised questionnaire to 30 to 50 people exactly as you will in the main study, and measure:
- Median completion time, on phones if that is how people will respond.
- Drop-off by page — where people quit is where the problem is.
- Item distributions: an item where more than 80% pile up at one end barely discriminates between anyone.
- Rates of “don’t know” and skipped answers per item.
- Whether skip logic routes the right people to the right questions.
The most skipped step: export the pilot data and run your entire analysis script on it. This is where you find reverse-coded items, scales exported as text instead of numbers, and IDs that don’t match across waves.
What not to do
- Merge pilot data into the main study when the instrument or protocol has changed.
- Present a pilot as evidence of effectiveness.
- Relabel a small, non-significant study as a “pilot” after the fact.
- Skip ethics review: a pilot with human participants needs approval just like the main study.
A checklist for a pilot that earns its keep
- Write 4 to 6 specific feasibility questions, each tied to a measurable indicator.
- Set green, amber and red thresholds, worked back from the main study plan.
- Size the pilot for precision, not power.
- Run cognitive interviews in rounds before a wider field test.
- Have staff log every deviation from protocol.
- Run the full analysis script on pilot data.
- Write a short report: feasibility indicators with confidence intervals, the decision, and the list of changes.
If you are drafting a proposal, add a section on the pilot with a progression-criteria table. Funders and review panels tend to trust a plan that already knows where it might break and how it will find out.
Câu hỏi thường gặp
How many participants do you need for a pilot study?
There is no universal number. Size the pilot for the precision you need on feasibility indicators such as retention or recruitment rate. Rules of thumb like 12 per group or 30 overall are starting points, not answers.
Can I use pilot study results for a power calculation?
Avoid using the pilot’s observed effect size, because small samples give very wide intervals and selected pilots tend to overestimate effects. Base power on the smallest effect of practical importance, and use the pilot to estimate the standard deviation and attrition.
What is the difference between a pilot study and a feasibility study?
Feasibility study is the umbrella term for work asking whether a study should and can be done. A pilot study is a type of feasibility study that rehearses all or part of the planned protocol on a small scale.
What is cognitive interviewing in questionnaire design?
It is a pretesting method where respondents think aloud or answer probes so you can see how they understand and answer each item. It is usually done in rounds of 5 to 8 people, revising between rounds until no major problems remain.
Can pilot data be included in the main study analysis?
Only if the instrument and protocol stayed identical and this was planned in advance. If anything changed after the pilot, keep the data separate, because the two phases no longer measure the same thing in the same way.