FESK.COMYour global study desk
Email us
Statistics & Data

P values and confidence intervals: six misreadings and how to report results

p = 0.04 does not mean your hypothesis is 96% likely to be true, and p = 0.07 does not mean there is no effect. Six common misreadings, why confidence intervals say more, and wording for reporting results.

P values and confidence intervals: six misreadings and how to report results

Two studies evaluate the same programme. Study A reports an effect of 2 points, 95% confidence interval 0.1 to 3.9, p = 0.04, and is written up as “effective”. Study B reports an effect of 5 points, confidence interval −0.5 to 10.5, p ≈ 0.08, and is written up as “ineffective”. Read carefully, Study B hints at a possibly much larger effect, estimated imprecisely because its sample was small. The two opposite conclusions come down mostly to the number 0.05.

The p value is among the most misunderstood ideas in science, to the point that the American Statistical Association issued a formal statement in 2016 restating what it does and does not mean.

What a p value actually says

A p value is the probability of observing data at least as extreme as yours, if the null hypothesis (H0) were true and all the model’s assumptions held. It measures how incompatible the data are with H0. That is all.

Note the “if”. The p value is computed in an imagined world where H0 is true, so it cannot tell you the probability that H0 is true in the real one.

Six common misreadings

MisreadingWhy it is wrong
“p = 0.03 means a 3% chance H0 is true”p is calculated assuming H0 is true; it is not the probability of H0
“p = 0.03 means a 97% chance the finding is real”The same error in reverse
“p > 0.05 proves there is no difference”Failing to reject H0 is not proving it; the sample may simply be too small
“The smaller the p, the bigger the effect”p also depends on sample size; a tiny effect in a huge sample gives a tiny p
“Group A was significant and B was not, so A and B differ”You need to test the difference between A and B directly
“p < 0.05 means the result matters”Statistical significance is not practical significance

The fifth error is endemic in subgroup analyses: “the intervention worked in women (p = 0.02) but not in men (p = 0.20), so the effect differs by sex.” That claim needs a test of interaction, and such tests are very often non-significant.

Confidence intervals tell you more

A 95% confidence interval gives a range of effect sizes compatible with the data. It carries three pieces of information at once:

  • Direction and size of the estimate: the midpoint.
  • Precision: how wide or narrow the interval is.
  • The test result: if the 95% interval excludes the null value (0 for a difference, 1 for a ratio), then p < 0.05 for the corresponding test.

The strict interpretation: if the study were repeated many times with intervals computed the same way, about 95% of those intervals would contain the true value. In practice, reading the interval as the range of values the data do not rule out is useful and widely used.

Back to the opening example: Study B’s interval from −0.5 to 10.5 says the data are compatible with anything from essentially no effect to a very large one. The honest conclusion is “too imprecise to tell”, not “no effect”.

Overlapping intervals do not mean no difference

A very common chart-reading error: seeing that two groups’ 95% confidence intervals overlap and concluding there is no significant difference. In fact two intervals can overlap somewhat while the difference is still statistically significant. To know, compute a confidence interval for the difference itself.

Reporting: drop “NS” and asterisks

  • Report exact p values (p = 0.07) rather than “NS” or “p > 0.05”. For very small values, write “p < 0.001”.
  • Always pair the p value with an effect estimate and confidence interval. A p value alone is nearly useless.
  • Avoid phrases like “approaching significance” or “a trend towards significance” for p = 0.06. Describe the estimate and its uncertainty instead.
  • If you ran many tests, say whether and how you adjusted for multiple comparisons.

Where 0.05 comes from

The 0.05 threshold is a convenient convention popularised in the first half of the last century and associated with the statistician Ronald Fisher, who treated it as a signal worth following up rather than a line between true and false. Some fields use far stricter thresholds: particle physics requires the equivalent of “five sigma” before announcing a discovery, and genome-wide association studies use tiny thresholds because they run hundreds of thousands of tests. There is nothing magical about 0.05.

In practice: rewriting results sentences

Weak: “The intervention group scored higher than the control group (p < 0.05).”

Better: “The intervention group scored on average 4.2 points higher than the control group (95% CI 1.1 to 7.3; p = 0.008), a standardised effect of d = 0.45.”

Weak: “There was no difference in dropout between groups (p = 0.21).”

Better: “Dropout was 3 percentage points lower in the intervention group (95% CI from 8 points lower to 2 points higher; p = 0.21); the data are too imprecise to rule out a practically important effect in either direction.”

The general template: direction + size + units + confidence interval + p + cautious interpretation.

Do this now: open the results section of your manuscript, find every “p < 0.05”, “NS” or asterisk, and rewrite each sentence using the template. You will often find a conclusion that needs softening, or a “non-significant” result that deserves more discussion than it got.

Câu hỏi thường gặp

What does a p value of 0.05 mean?

If the null hypothesis and the model assumptions were true, there would be a 5% probability of observing data at least this extreme. It is not the probability that the null hypothesis is true.

Does p greater than 0.05 mean there is no effect?

No. It means the data are not strong enough to reject the null hypothesis. Look at the confidence interval to see how large an effect the data are still compatible with.

How do I interpret a 95% confidence interval?

If the study were repeated many times with intervals computed the same way, about 95% of them would contain the true value. In practice it shows the range of effect sizes compatible with your data.

Should I write NS instead of the p value?

No. Report the exact p value along with the effect estimate and confidence interval, and write p less than 0.001 only for very small values.

If two confidence intervals overlap, is the difference not significant?

Not necessarily. Intervals can overlap partly while the difference is still significant; compute a confidence interval for the difference itself.

Need specific advice for your case?

We will contact you within 24 hours.

Request consultation now

Related articles

🧭
Bạn đang ở chặng nào của đường học vị?
Nhập chỗ bạn đang đứng và đích bạn nhắm — công cụ trả về số năm, chi phí và việc phải làm từng chặng.
Show my pathway →
Miễn phí, không cần tài khoản. Xem tất cả công cụ

Need advice? Talk to us

Leave your details and our team will contact you within 24 hours. The first consultation is completely free.

or
info@fesk.com