What does a p-value actually mean? p = 0.04 is not 96% sure
What a p-value measures, what it never tells you, why 0.05 is a convention, the peeking habit that manufactures winners, and how to report an A/B test honestly: one page.
Get the free PDF
One page, print-ready, free to share. No signup needed.
p = 0.04 is not 96% sure. A p-value is the probability of data this extreme if nothing changed, and nearly every A/B readout flips it into the probability that the idea works. One page on what p measures, what it never tells you, and the peeking habit that manufactures winners. The print-ready A4 PDF is at the bottom.
What p says
- p is P(data | H0): the probability of seeing a result this extreme if nothing changed.
- p = 0.04 means data this extreme shows up 4 times in 100 by luck alone.
- It is not P(H0 | data). The flip is the trap.
What it is not
- Not "96% sure": see the first line. p says nothing about how likely your idea is.
- Not effect size: tiny lifts pass too, given enough users.
- Not importance: p ignores money.
The threshold
- alpha = 0.05 is a convention, not physics.
- 0.049 versus 0.051 is the same evidence; only the label changed.
- At 0.05 you accept 1 false alarm in 20 tests where nothing is going on.
Read it right
- Report the confidence interval: a range beats a verdict.
- The honest trio: lift + CI + n.
- Fix n up front, then stop looking.
The peeking trap
- Check daily and p dips under 0.05 by luck at some point.
- Stop at a win and you have just manufactured a false positive.
- If you must peek, use a sequential test built for repeated looks.
The false-positive machine, in 9 lines
rng = np.random.default_rng(7)
a = rng.binomial(1, 0.10, 20_000) # same rate
b = rng.binomial(1, 0.10, 20_000) # same rate
wins = 0
for n in range(500, 20_000, 500): # peek every 500
_, p = ztest(a[:n], b[:n])
wins += p < 0.05
# wins = 6 of 39 peeks flagged "significant"
A and B are the same coin. Peeking every 500 users and stopping the moment p dips under 0.05 turns a 5% false-alarm rate into a near certainty: 6 of the 39 peeks flagged a difference that does not exist.
Gotchas
- 20 metrics on one test: one wins by chance.
- p = 0.2 is no evidence at this n, not no effect.
- Huge n: everything is significant, so read the lift instead.
The trap: what the sentence actually claims
| You say | It means | Say instead |
|---|---|---|
| p = 0.04, so 96% sure | nothing about your idea | unlikely under no effect |
| p = 0.2, so no effect | no evidence at this n | inconclusive, need more |
| significant | probably not pure luck | lift of X, CI [a, b] |
Interview phrasing worth memorizing: a p-value is the probability of the data given no effect, never the probability of the effect given the data.
Frequently asked questions
What does p = 0.04 mean in an A/B test?
Why is a p-value not the probability that the hypothesis is true?
Why does checking an A/B test every day produce false positives?
How should you report an A/B test result?
Get the free PDF
One page, print-ready, free to share. No signup needed.