Multi-Armed Bandits & Exploration
The explore/exploit tradeoff, ε-greedy, UCB1, Thompson sampling, and regret — the simplest RL you will actually ship
Learn the multi-armed bandit framework as the simplest form of reinforcement learning used in production: the explore/exploit tradeoff, ε-greedy, UCB1 with a derivation sketch, Thompson sampling with a worked Beta-Bernoulli example, and regret as the metric that makes 'how good is this exploration strategy' precise — contrasted directly with fixed-split A/B testing.
Practice questions (6)
-
View →
Choosing Between a Bandit and an A/B Test for Subject Line Selection
Intermediate · Free -
View →
Diagnose a Stuck ε-Greedy Policy
Intermediate -
View →
Compute and Interpret UCB1 Scores
Intermediate -
View →
Run Thompson Sampling by Hand on a Beta-Bernoulli Bandit
Intermediate -
View →
Interpret and Compare Regret Curves for Two Deployed Policies
Advanced