Intermediate
Open
Pro
Diagnose a Stuck ε-Greedy Policy
A team deployed ε-greedy with a fixed ε = 0.05 to choose among 8 notification send-time buckets. Six months and 50 million sends later, one bucket that was slightly ahead in the first few thousand sends (due to noise) is still receiving 95%+ of traffic, and a teammate discovers — by manually forcing traffic to another bucket for a week — that a different bucket is actually about 15% better.
- Explain, mechanistically, how ε-greedy with fixed ε = 0.05 allowed this to happen despite sending 5% of 50 million (2.5 million) sends to non-leading buckets.
- Why didn't 2.5 million exploratory sends across 7 other buckets (roughly 357,000 each) surface the better bucket?
- Propose two concrete fixes and explain the mechanism by which each one would have caught this faster.
Share this question