FESK.COMYour global study desk
Email us
Statistics & Data

ANOVA and post hoc tests: choosing Tukey, Games-Howell or Bonferroni

A significant F is only half an answer. Checking assumptions, when to use Welch’s ANOVA, choosing the right post hoc test, reading interactions in two-way ANOVA and handling repeated-measures designs.

ANOVA and post hoc tests: choosing Tukey, Games-Howell or Bonferroni

A study compares three methods of teaching vocabulary, 40 students per group. ANOVA gives F(2, 117) = 6.1, p = 0.003. The authors write “the three methods differ significantly” and stop there. Readers still do not know the one thing they care about: which method beats which, and by how much. It may be that only method C differs from the other two, while A and B are practically identical.

ANOVA is a chain of decisions: check assumptions, choose the right variant, decide how to compare pairs, report effect sizes. Each step has a sensible default and some familiar traps.

Before you run it: assumptions and variants

One-way ANOVA assumes independent observations, approximately normal residuals within groups and similar variances across groups. In practice:

  • Independence comes from the design; if the same person appears in several conditions, you have a repeated-measures design.
  • Normality: with a few dozen per group, ANOVA is fairly robust. Look at a Q-Q plot of residuals rather than relying on a normality test.
  • Equal variances matter more, especially with unequal group sizes. If the smallest group has the largest variance, classical ANOVA is prone to false positives.

The safe choice when variances may differ is Welch’s ANOVA, the multi-group counterpart of Welch’s t-test. When variances really are equal, it gives almost the same answer as classical ANOVA. Many packages report both; say which you used.

Planned or exploratory?

There are two routes from F to specific comparisons:

  1. Planned contrasts: a few comparisons specified in advance on theoretical grounds, such as “the two new methods versus the traditional one”. Fewer comparisons, more power.
  2. Post hoc comparisons: all pairs, with adjustment to control the type I error rate.

If your hypotheses were clear from the start, planned contrasts are usually more honest and more powerful. Write them into the analysis plan before you see the data.

Choosing a post hoc test

TestWhen to useNotes
Tukey HSDAll pairwise comparisons, similar variancesThe common default; good balance of error control and power
Games-HowellAll pairwise comparisons, unequal variances or group sizesNatural partner for Welch’s ANOVA
BonferroniA small number of pre-selected comparisonsSimple but conservative when comparisons are many
HolmAs for BonferroniAlways at least as powerful as Bonferroni while still controlling overall error
DunnettEvery group against one controlDoes not compare treatment groups with each other

With three groups and similar variances, Tukey is almost always a reasonable choice. Avoid running a batch of unadjusted t-tests and calling them “post hoc”: five groups mean ten pairs, and the chance of at least one false positive far exceeds 5%.

Two-way ANOVA: read the interaction first

With two factors, say teaching method (A, B) and prior attainment (low, high), ANOVA gives three results: a main effect of method, a main effect of attainment and an interaction. The rule: read the interaction first.

If the interaction is significant, main effects can mislead. Suppose method A raises low-attaining students’ scores by 8 points but lowers high-attaining students’ scores by 2 points compared with B. The “average” main effect of A may be positive and significant, yet saying “A is better than B” is wrong for half the students. Analyse simple effects instead: compare A and B separately within each attainment group.

Always plot the means by both factors. Parallel lines suggest no interaction; crossing lines suggest a strong one, most striking when the direction of the effect reverses.

Repeated measures and pre-post designs

When the same people are measured at several time points, use repeated-measures ANOVA. Its particular assumption is sphericity; when violated, use the Greenhouse-Geisser correction most packages report automatically. With missing time points or irregular spacing, a mixed model is more flexible and keeps more participants in the analysis.

For a two-group design with pre- and post-test, many people run a 2 × 2 mixed ANOVA or compare change scores. In randomised trials the approach usually recommended is ANCOVA: post-test as the outcome, pre-test as a covariate. It tends to give more precise estimates and handles regression to the mean well.

Report effect sizes

Alongside F and p, report an overall effect size (η² or ω²; ω² is less biased in small samples) and, more usefully, mean differences with confidence intervals for the specific comparisons. “Method C scored 5.2 points higher than A (Tukey-adjusted 95% CI 1.4 to 9.0)” helps readers far more than any F statistic.

In practice: a workflow and a model sentence

  1. Plot the outcome by group with boxplots or dot plots.
  2. Check group SDs and sizes; if the largest SD is about twice the smallest and group sizes differ, prefer Welch.
  3. Run ANOVA (classical or Welch) and record F, degrees of freedom, p and ω².
  4. Run planned contrasts or an appropriate post hoc test.
  5. Report adjusted mean differences and confidence intervals for the key pairs.
  6. With two factors: test the interaction, plot it and analyse simple effects where needed.

Model sentence: “Vocabulary scores differed across the three methods, Welch F(2, 76.4) = 6.3, p = 0.003, ω² = 0.08. Games-Howell comparisons showed that method C scored 5.2 points higher than A (95% CI 1.4 to 9.0) and 4.6 points higher than B (0.7 to 8.5); A and B did not clearly differ (0.6 points; −3.1 to 4.3).” Fractional degrees of freedom are normal for Welch’s test.

Next step: for your next ANOVA, write down in advance whether you will use planned contrasts or post hoc tests, which ones and why. One line in the analysis plan heads off a reviewer’s hardest question: “were these comparisons chosen before or after looking at the data?”.

Câu hỏi thường gặp

When should I use Welch’s ANOVA?

When group variances may differ, particularly with unequal group sizes. When variances are equal it gives almost the same result as classical ANOVA, so it is a safe default.

What is the difference between Tukey and Bonferroni?

Tukey is designed for all pairwise comparisons and balances error control with power. Bonferroni is simple and suits a few pre-selected comparisons, but becomes very conservative as comparisons multiply.

What if ANOVA is significant but no post hoc pair is?

It can happen because the tests answer different questions. Report it honestly, with mean differences and confidence intervals for each pair.

How do I interpret main effects when the interaction is significant?

Cautiously. A main effect is an average that can hide effects differing between groups. Analyse simple effects within each level of the other factor.

How should I analyse a two-group pre-post design?

For randomised trials, ANCOVA with the post-test as outcome and the pre-test as covariate is usually recommended, as it gives more precise estimates than analysing change scores.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com