Choosing Between a Bandit and an A/B Test for Subject Line Selection
Your team currently runs a 5-way, equal-split A/B test for two weeks every time it wants to compare push-notification subject lines, and you suspect this wastes opens. A stakeholder asks: "why not just always use a bandit instead of A/B testing, for everything?"
- Explain concretely, with the two-week/5-way test as your example, why fixed-split A/B testing costs opens relative to a bandit — be specific about when in the test the cost is incurred.
- Give a scenario elsewhere in the same company where a fixed-split A/B test is still the better tool than a bandit, and explain why.
- Propose a hybrid: how would you use a bandit for ongoing subject line selection while still being able to answer "was our new subject-line-generation approach better than the old one, with statistical confidence" for a quarterly review?
1. The cost of fixed-split testing
In a 5-way, equal-split, two-week test, every arm receives 20% of traffic for the entire two weeks by design — that is what makes the statistics clean. If one subject line is clearly worse by, say, day 2 (which is often visible with enough volume), the test design nonetheless keeps sending it to 20% of users for the remaining 12 days, because reallocating traffic mid-test based on interim results is exactly the kind of peeking that invalidates the significance calculation. The lost opens are the difference between what that 20% of traffic would have earned under the best-known arm and what it actually earned under the underperforming one, accumulated over those 12 "wasted" days. A bandit (UCB1 or Thompson sampling) would have started shrinking that arm's allocation within hours as evidence accumulated, capturing most of that lost value while still keeping a small, shrinking trickle of traffic on it in case it recovers.
2. Where A/B testing remains the better tool
A one-time, high-stakes launch decision — e.g., "should we replace the entire checkout flow with a redesigned version" — is a better fit for a fixed-split A/B test. The decision is made once, it's expensive to reverse, and the business needs a defensible, unbiased estimate of the effect size with a confidence interval to justify the change to stakeholders (and possibly to detect harm reliably before a full rollout). A bandit optimizing in-flight reward during that test could converge to one variant early based on noisy initial data and starve the other of the sample size needed to actually distinguish them with confidence — exactly the rigor an A/B test is designed to preserve.
3. A hybrid approach
Run the bandit for day-to-day subject line selection among the currently live candidates — this captures the ongoing traffic efficiency gains. Separately, when the team wants a rigorous generation-approach comparison for the quarterly review, run a proper randomized, fixed-split holdout: reserve a small, fixed percentage of traffic (say 5%), randomly and evenly split between "old approach's best subject line" and "new approach's best subject line," untouched by the bandit's adaptive allocation, for a pre-specified duration. This gives a clean, bandit-independent A/B comparison for the strategic question while the bulk of traffic still benefits from adaptive allocation among live candidates — a common pattern where bandits handle continuous optimization and a protected, randomized slice handles periodic, high-stakes evaluation (this also mirrors how a randomized exploration slice is used to keep unbiased data flowing in ML monitoring more broadly).
Share this question