It is one of the most common applied projects in teacher education worldwide: a teacher uses a new method with their class for eight weeks, the pre-test average is 58%, the post-test average is 73%, and the conclusion reads “the method was clearly effective”. It may well be. But the design cannot distinguish the method from eight weeks of ordinary teaching, from familiarity with the test format, or from the teacher’s own enthusiasm for trying something new.
The experiment is the one design that supports causal claims with reasonable confidence, but only when it genuinely rules out competing explanations. That is where most of the design effort goes.
Five rival explanations for a rising score
| Threat | Meaning | Classroom example |
|---|---|---|
| Maturation | Learners improve over time anyway | Eight weeks of the normal curriculum also raises scores |
| Testing effect | Taking a test improves later performance | Pre- and post-test share a format |
| Regression to the mean | Groups chosen for low scores drift upwards | Picking the weakest class for the trial |
| History | Something else happens at the same time | The school runs exam revision sessions |
| Novelty and expectancy | Everyone tries harder because it is a trial | The teacher prepares far more than usual |
Regression to the mean deserves special attention. A class chosen because of an unusually low pre-test is almost guaranteed to score higher next time, simply because part of that low score was bad luck. Without a comparison group selected the same way, you cannot separate this from any real effect.
The control group is not optional
Adding a comparison group taught as usual, tested at the same times with the same instruments, neutralises most of the threats above because they act on both groups. The question becomes: how much more did the intervention group improve?
There are three levels of rigour:
- Randomised experiment: individuals are randomly assigned to groups. The strongest option, but often impractical in schools where students learn in intact classes.
- Cluster randomisation: whole classes or schools are randomly assigned. More feasible, but you need many clusters; randomising two classes is barely different from choosing them by hand.
- Non-equivalent groups quasi-experiment: existing classes, no randomisation. The most common in theses; it needs a pre-test to check and adjust for baseline differences.
Redesigning a real-looking project
Topic: the effect of station-rotation teaching in Year 10 biology. Resources: one teacher with four parallel classes of about 30 students each.
- Randomly assign two classes to the intervention and two to comparison, rather than hand-picking a “suitable” class.
- Pre-test all four classes with the same test; post-test with a parallel form of equal difficulty, marked by someone who does not know which class is which.
- Hold teacher, content and time constant across groups; only the method differs.
- Keep an implementation log in intervention classes (fidelity) and record how comparison classes are taught.
- Monitor contamination: are students sharing materials across groups?
- Analyse with ANCOVA or regression, post-test as outcome and pre-test as covariate, accounting for students nested within classes.
With only four classes, the design remains weak at the class level, and the limitations section must say so. But it is far stronger than one class measured before and after.
Analysis: gain scores or baseline adjustment
Two approaches are common: compare gain scores (post minus pre) between groups, or run ANCOVA with the pre-test as a covariate. In randomised trials ANCOVA generally has more power. With non-equivalent groups the two approaches can disagree; reporting both and explaining the difference is the transparent route. Always report an effect size such as Hedges’ g with a confidence interval, not just a p-value.
Ethics and the comparison group
A frequent worry: is it unfair to leave the comparison group with the usual method? If it is genuinely unknown whether the new method is better — which is why you are studying it — nobody is knowingly disadvantaged. A common solution is a wait-list control, where the comparison group receives the new method after data collection ends. Obtain school permission, inform parents as your institution requires and seek ethics approval where applicable.
Reporting to a standard
For randomised trials, the CONSORT statement (and its extension for cluster trials) lists what to report: how the random sequence was generated, who knew the allocation, a flow diagram from recruitment to analysis, and attrition in each arm. Even for quasi-experiments, using CONSORT as a checklist makes a paper markedly more complete.
What you do not need
You do not need a year-long intervention for a master’s thesis; a tightly designed six-to-ten-week intervention is more convincing than a long, loosely controlled one. You do not need ten outcome measures; choose one primary outcome, declare it in advance, and treat the rest as secondary.
An experimental design checklist
- A comparison group measured at the same times.
- Random assignment at the most feasible level, or a stated reason why not.
- A pre-test and a check of baseline equivalence.
- Markers blind to group allocation.
- One primary outcome declared in advance.
- Monitoring of implementation fidelity and contamination.
- An analysis plan that accounts for nesting within classes.
Next step: copy the five threats from the table above next to your current design and write one line for each saying how your design rules it out. Any empty line is something to fix before the intervention starts.
Câu hỏi thường gặp
Is a one-group pretest-posttest design valid?
It is useful for piloting or describing change, but it cannot show that the intervention caused the change. Maturation, testing effects and regression to the mean can all explain a rising score.
What is the difference between an experiment and a quasi-experiment?
An experiment randomly assigns participants to groups, while a quasi-experiment uses existing groups such as intact classes. Quasi-experiments need baseline measurement and adjustment, and more cautious conclusions.
How should I analyse pre-test and post-test data?
A common approach is ANCOVA or regression with the post-test as outcome and the pre-test as covariate. If students are nested in classes, model that structure, and always report an effect size.
Is it unethical to have a control group in education research?
Not if it is genuinely unknown whether the new method is better. A wait-list design, where the control group receives the intervention afterwards, is a common solution.
What is the CONSORT statement?
CONSORT is a reporting guideline for randomised controlled trials, with a checklist and participant flow diagram. Extensions exist for cluster trials, and it is a useful checklist for quasi-experiments too.