← All cheatsheets
Data Analyst · #061 · September 30, 2026 · 2 min read

What does a p-value actually mean? p = 0.04 is not 96% sure

What a p-value measures, what it never tells you, why 0.05 is a convention, the peeking habit that manufactures winners, and how to report an A/B test honestly: one page.

Get the free PDF

One page, print-ready, free to share. No signup needed.

Download the PDF

p = 0.04 is not 96% sure. A p-value is the probability of data this extreme if nothing changed, and nearly every A/B readout flips it into the probability that the idea works. One page on what p measures, what it never tells you, and the peeking habit that manufactures winners. The print-ready A4 PDF is at the bottom.

What p says

  • p is P(data | H0): the probability of seeing a result this extreme if nothing changed.
  • p = 0.04 means data this extreme shows up 4 times in 100 by luck alone.
  • It is not P(H0 | data). The flip is the trap.

What it is not

  • Not "96% sure": see the first line. p says nothing about how likely your idea is.
  • Not effect size: tiny lifts pass too, given enough users.
  • Not importance: p ignores money.

The threshold

  • alpha = 0.05 is a convention, not physics.
  • 0.049 versus 0.051 is the same evidence; only the label changed.
  • At 0.05 you accept 1 false alarm in 20 tests where nothing is going on.

Read it right

  • Report the confidence interval: a range beats a verdict.
  • The honest trio: lift + CI + n.
  • Fix n up front, then stop looking.

The peeking trap

  • Check daily and p dips under 0.05 by luck at some point.
  • Stop at a win and you have just manufactured a false positive.
  • If you must peek, use a sequential test built for repeated looks.

The false-positive machine, in 9 lines

rng = np.random.default_rng(7)
a = rng.binomial(1, 0.10, 20_000)  # same rate
b = rng.binomial(1, 0.10, 20_000)  # same rate

wins = 0
for n in range(500, 20_000, 500):  # peek every 500
    _, p = ztest(a[:n], b[:n])
    wins += p < 0.05
# wins = 6 of 39 peeks flagged "significant"

A and B are the same coin. Peeking every 500 users and stopping the moment p dips under 0.05 turns a 5% false-alarm rate into a near certainty: 6 of the 39 peeks flagged a difference that does not exist.

Gotchas

  • 20 metrics on one test: one wins by chance.
  • p = 0.2 is no evidence at this n, not no effect.
  • Huge n: everything is significant, so read the lift instead.

The trap: what the sentence actually claims

You sayIt meansSay instead
p = 0.04, so 96% surenothing about your ideaunlikely under no effect
p = 0.2, so no effectno evidence at this ninconclusive, need more
significantprobably not pure lucklift of X, CI [a, b]

Interview phrasing worth memorizing: a p-value is the probability of the data given no effect, never the probability of the effect given the data.

Frequently asked questions

What does p = 0.04 mean in an A/B test?
It means that if nothing had changed between the variants, data this extreme would show up about 4 times in 100 by luck alone. Formally it is P(data | H0), the probability of the observed result given no effect. It is not the probability that your variant works, and it is not 96% certainty of anything.
Why is a p-value not the probability that the hypothesis is true?
Because it conditions the wrong way round. A p-value is the probability of the data given no effect; the probability of the effect given the data would need a prior and is a different quantity. Flipping the two is the single most common misreading in slide decks, and it turns 'unlikely under no effect' into 'we are 96% sure'.
Why does checking an A/B test every day produce false positives?
Because with a 0.05 threshold, p dips under 0.05 by chance at some point in almost any long test, even when the variants are identical. If you stop the moment you see a win, you lock in that lucky dip. Fix the sample size up front and only read the result once, or use a sequential test designed for repeated looks.
How should you report an A/B test result?
Report the lift, its confidence interval, and n, not a verdict. A range tells the reader both the direction and how uncertain the estimate is, while 'significant' only says the result is probably not pure luck. A tiny lift can be significant with enough users and still be worthless, and p = 0.2 means inconclusive at this sample size, not 'no effect'.

Get the free PDF

One page, print-ready, free to share. No signup needed.

Download the PDF

More cheatsheets