Practice — Multi-Armed Bandits & Exploration (6 questions)
Choosing Between a Bandit and an A/B Test for Subject Line Selection Permalink →
Your team currently runs a 5-way, equal-split A/B test for two weeks every time it wants to compare push-notification subject lines, and you suspect this wastes opens. A stakeholder asks: "why not just always use a bandit instead of A/B testing, for everything?"
- Explain concretely, with the two-week/5-way test as your example, why fixed-split A/B testing costs opens relative to a bandit — be specific about when in the test the cost is incurred.
- Give a scenario elsewhere in the same company where a fixed-split A/B test is still the better tool than a bandit, and explain why.
- Propose a hybrid: how would you use a bandit for ongoing subject line selection while still being able to answer "was our new subject-line-generation approach better than the old one, with statistical confidence" for a quarterly review?
Share this question
Run Thompson Sampling by Hand on a Beta-Bernoulli Bandit
Unlock this question →Interpret and Compare Regret Curves for Two Deployed Policies
Unlock this question →Can You Trust the Averages From Your Bandit? Permalink →
You run a Thompson Sampling bandit across 3 subject lines for 4 weeks. As expected, traffic shifts increasingly toward whichever arm looks best so far. At the end, you compute each arm's plain observed conversion rate (successes / impressions for that arm) and report "Arm B beat Arm A by 2.1 points."
A colleague who only knows fixed 50/50 A/B tests asks: is that 2.1 point gap, computed this way, a trustworthy, unbiased estimate of the true difference between the arms?
Share this question