Advanced
Open
Pro
Diagnose and Fix Position Bias in a Bandit-Ranked Carousel
A content app uses a contextual bandit to rank 5 content items in a home-screen carousel. Six weeks after launch, an analyst notices that whatever item the bandit puts in slot 1 has, on average, a much higher observed click-through rate than the same items get when placed in slot 3 — even comparing the same item across different impressions.
- Explain the mechanism by which this creates a self-reinforcing problem for the bandit's reward estimates specifically (not just a general observation about position bias).
- Propose two concrete fixes, drawing directly on techniques named in this subject, and explain how each interrupts the loop.
- Explain precisely why the bandit's own UCB or Thompson-sampling exploration mechanism did not already prevent this.
Share this question