Advanced
Open
Pro
Interpret and Compare Regret Curves for Two Deployed Policies
Your company ran two live bandit deployments for the same subject-line-selection problem (5 arms) over 6 months and logged cumulative regret weekly. Policy 1 (fixed ε = 0.2 ε-greedy) shows regret climbing at a roughly constant rate every week for all 26 weeks. Policy 2 (Thompson sampling) shows regret climbing quickly in the first few weeks, then visibly flattening out by week 10 onward.
- Explain what each regret curve shape implies about each policy's long-run behavior, using the linear vs. sublinear regret distinction.
- A stakeholder looks at total cumulative regret at week 26 and notes Policy 2's is lower, and asks: "does that mean Thompson sampling is strictly better in every week?" Answer precisely.
- Estimate, qualitatively, what would happen to each policy's regret curve shape if the best arm's true mean started drifting slowly starting in week 15 (a novelty-decay effect on the previously-best subject line).
Share this question