Intermediate
Open
Pro
Run Thompson Sampling by Hand on a Beta-Bernoulli Bandit
Two homepage banners, Sale and NewArrivals, start with a
Beta(1,1) prior on click-through rate. Over one day of traffic:
Sale gets 12 clicks out of 60 impressions; NewArrivals gets 3
clicks out of 10 impressions.
- Compute the posterior distribution for each banner after this day's data, and report the posterior mean for each.
- Without drawing an actual random sample, explain qualitatively why
NewArrivals, despite fewer total clicks, has a real chance of being selected on the next round even though its posterior mean is lower thanSale's (assume it currently is — check this first). - A colleague suggests initializing both arms with a strong prior, Beta(50, 50), "to avoid wild early swings." What is the tradeoff of doing this?
Share this question