Can You Trust the Averages From Your Bandit?
You run a Thompson Sampling bandit across 3 subject lines for 4 weeks. As expected, traffic shifts increasingly toward whichever arm looks best so far. At the end, you compute each arm's plain observed conversion rate (successes / impressions for that arm) and report "Arm B beat Arm A by 2.1 points."
A colleague who only knows fixed 50/50 A/B tests asks: is that 2.1 point gap, computed this way, a trustworthy, unbiased estimate of the true difference between the arms?
No — the naive per-arm average is a biased estimator here, and the bias comes from the adaptivity itself.
In a fixed-allocation A/B test, each arm's assignment probability is a constant, decided in advance and independent of any outcome — so the simple sample average is unbiased for the arm's true mean. A bandit breaks that assumption on purpose: the probability of being assigned to an arm at time t depends on the outcomes observed before t. An arm that got lucky early gets more traffic while it looks good; if it regresses later, by then it's already accumulated the majority of its data during its "hot streak." The result is an estimator that's systematically optimistic for whichever arm the algorithm favored — structurally similar to a winner's-curse effect, except here it's baked into every arm's average, not just the ex-post "winner."
Concretely: data collected while Thompson Sampling was still uncertain (early weeks, near-equal traffic) behaves like a normal random sample, but data collected once the algorithm became confident is conditioned on the arm having looked good relative to the others at that moment — that's not the same as an independent draw from the arm's true conversion distribution. Averaging both regimes together without adjustment inherits that conditioning.
The fix is not "don't use bandits" — it's using adaptive-inference corrected estimators built for exactly this setting (e.g. AIPW-style estimators that weight each observation by the inverse of the time-varying assignment probability it actually had, replaying the sequential allocation rule rather than pretending it was fixed), which restore (asymptotically) valid, less-biased effect estimates from adaptively collected data. If you need a textbook-unbiased number with no correction required, that's what a fixed-allocation A/B test is for — it's a real trade-off against the bandit's lower regret during the experiment itself.
Share this question