What Daily Peeking Does to Your 5% False-Positive Rate
You launch a 30-day A/B test at the standard 5% significance level. Every morning you check the dashboard, and you'll ship the variant the first day it shows p < 0.05.
If the variant actually does nothing, what's the real probability you'll declare a winner by day 30?
Explain why it isn't 5%, and name a testing approach that makes peeking safe.
C) ~25% — peeking quietly multiplies your false-positive rate ~5×.
The 5% guarantee holds for exactly one look at a pre-committed sample size. Each daily peek is another chance for noise to momentarily cross the p < 0.05 line — and "stop at the first crossing" harvests exactly those moments. Simulations of this setup (the classic "how not to run an A/B test" analysis) show that ~10 peeks inflate the false-positive rate to ~20%, and daily peeks over a month push it to roughly 25–30%.
A quarter of your null A/A-quality tests would "win." The dashboards look rigorous; the error rate is closer to a coin flip than to 5%.
The rule of thumb to say out loud: p-values are valid for one look — every extra peek-and-stop inflates false positives, and daily peeking roughly quintuples them.
How to peek safely:
- Sequential testing — mSPRT / always-valid p-values, or group-sequential designs with alpha-spending (O'Brien–Fleming bounds): look as often as you like, the math prices in every look.
- Simplest fix — commit to sample size and duration up front, and read the p-value once, at the end.
Share this question