FESK.COMYour global study desk
Email us
Statistics & Data

Choosing a statistical test: five questions that lead to the right one

Running an independent t-test on before-and-after data from the same people is an error reviewers spot at once. Five questions about outcome, groups and independence point to the right test, plus a quick-reference table.

Choosing a statistical test: five questions that lead to the right one

A group of teachers measures 40 pupils before and after a term using a new method, then compares the two columns of scores with an independent-samples t-test. Result: p = 0.12, “not significant”. Reanalysed with a paired t-test, which matches the design because the same pupils were measured twice, p falls below 0.01. Same data, different test, opposite conclusion.

Choosing a test is not a matter of looking something up in a chart. It starts with the study design: what was measured, on whom and how many times. Answer the five questions below correctly and the test more or less chooses itself.

Question 1: what type of outcome?

  • Continuous: test scores, blood pressure, time, income.
  • Binary: pass or fail, disease or not.
  • Nominal with several categories: field of study, type of school.
  • Ordinal: satisfaction from 1 to 5, disease stage.
  • Counts or time to event: these need dedicated models such as Poisson or Cox regression and sit outside the basic chart.

Question 2: how many groups?

Two, or three or more. With three or more, do not run a t-test for every pair: three groups means three comparisons, five groups means ten, and the chance of at least one “significant” result by luck alone rises quickly. Use ANOVA or its equivalent first, then adjusted post hoc comparisons.

Question 3: independent or paired observations?

This is the most important question and the one most often skipped. Data are paired when:

  • The same person is measured more than once (before and after, several time points).
  • Individuals are matched by design (twins, cases and controls matched on age and sex).
  • Two parts of the same unit are measured (left and right eye, two sides of the jaw).

Pupils in the same class or patients in the same hospital are not fully independent either. Where data are clustered like this you need multilevel models or cluster-robust standard errors, not just a simple test.

Question 4: are the distributional assumptions reasonable?

Parametric tests such as the t-test assume that means are approximately normally distributed. Two common misunderstandings:

  1. The assumption concerns the distribution of the mean (or the residuals), not perfectly normal raw data. With a few dozen observations per group, the t-test is fairly robust to moderate skew.
  2. Normality tests such as Shapiro-Wilk are poor referees. In large samples they flag trivial departures; in small samples they lack the power to detect serious ones. Look at histograms and Q-Q plots instead.

When data are heavily skewed, samples small, or the outcome ordinal, use a non-parametric test. Note that the Mann-Whitney test does not strictly compare medians; it tests whether values in one group tend to be larger than in the other, and should be described that way.

Question 5: what do you want to know?

A difference between groups, the strength of association between two variables, or a prediction? For two continuous variables measured on the same people, use correlation: Pearson for linear relationships, Spearman for monotonic relationships or ordinal data. If you need to control for confounders, you need regression, not a two-variable test.

Quick-reference table

OutcomeDesignParametricNon-parametric / alternative
Continuous2 independent groupsWelch t-testMann-Whitney U
Continuous2 paired measurementsPaired t-testWilcoxon signed-rank
Continuous3+ independent groupsOne-way ANOVA (Welch if variances differ)Kruskal-Wallis
Continuous3+ measurements on same peopleRepeated-measures ANOVA or mixed modelFriedman
Binary / nominalIndependent groupsChi-squareFisher’s exact when expected counts are small
BinaryPairedMcNemarExact McNemar
Two continuous variablesAssociationPearson correlationSpearman correlation

Why “Welch t-test” rather than Student’s? Welch does not assume equal variances, and when variances really are equal it gives almost identical results. Many statisticians recommend it as the default, and R’s t.test function uses Welch unless told otherwise.

For chi-square, a common rule is to switch to Fisher’s exact test when more than 20 per cent of cells have expected counts below 5. Expected counts, not observed ones.

What not to do

  • Try several tests and report the one with the smallest p value. That is a form of p-hacking; choose the test in your analysis plan before you see results.
  • Switch to a one-sided test because the two-sided result missed the threshold. One-sided tests are justified only when specified in advance and when an effect in the opposite direction would genuinely be irrelevant.
  • Run a normality test and swap tests mechanically while ignoring the design.

In practice: one worked example

Question: does a new revision programme raise pass rates? 120 students, 60 randomised to the new programme and 60 to the old. Outcome: pass or fail. The five answers: binary, two groups, independent, no normality assumption needed, comparing proportions. Test: chi-square, or Fisher’s exact if expected counts are small. Report the pass rate in each group, the difference in proportions with a 95% confidence interval, and the p value. If students come from several classes, consider logistic regression with class as a factor.

Next step: write your five answers as five lines in your analysis plan before you open any software. If line three contains the word “paired” or “clustered”, stop and double-check the test you intended to use.

Câu hỏi thường gặp

When should I use a paired t-test instead of an independent t-test?

When the two measurements come from the same people or from pairs matched by design, such as scores before and after an intervention in the same pupils.

Can I use a t-test if my data are not normal?

With a few dozen observations per group, the t-test is fairly robust to moderate skew. With small samples and heavily skewed or ordinal data, use Mann-Whitney or Wilcoxon instead.

When should I use Fisher’s exact test instead of chi-square?

When more than about 20 per cent of cells have expected counts below 5, which is common with small samples. The rule concerns expected, not observed, counts.

Can I run multiple t-tests to compare three groups?

No. Multiple comparisons inflate the false positive rate. Use ANOVA or Kruskal-Wallis first, then adjusted pairwise comparisons.

Why use Welch’s t-test?

Welch’s test does not assume equal variances and gives nearly identical results to Student’s t-test when variances are equal, so it is a safer default.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com