Advanced
Open
Pro
When to Graduate from a Contextual Bandit to Full RL
Your team's contextual bandit for notification content selection has been live for three months. It treats each day's notification decision independently: given today's state, pick the action (send with content X / send with content Y / suppress) that maximizes today's shaped reward. It has improved the shaped-reward metric by 9% over the rules baseline, and the team is debating whether to invest in a full multi-step RL system next, or to keep improving the bandit's features and model instead.
- Describe a concrete scenario where the bandit's myopia (treating each day independently) leaves value on the table that only a multi-step model could capture.
- What evidence, specifically, would you look for in the bandit's own logs to decide whether that scenario is actually costing the business anything, before committing to the more expensive RL build?
- If you do graduate to full RL, what should carry over unchanged from the bandit stage, and why does reusing it (rather than rebuilding) matter for how quickly you can trust the new system?
Share this question