Match a job Paths Subjects Questions Quizzes Pricing

Multi-Armed Bandits & Exploration

The explore/exploit tradeoff, ε-greedy, UCB1, Thompson sampling, and regret — the simplest RL you will actually ship

Overview Read

Multi-Armed Bandits & Exploration

Most of what gets called "reinforcement learning" in industry is not the deep, multi-step planning you saw in RL Foundations: MDPs & Value Functions — it is a much simpler object called a multi-armed bandit, and it is quietly running behind a huge share of the decisions a consumer app makes every second: which of five push-notification subject lines to send, which homepage hero image to show, which of three pricing tiers to default to, which ad creative to serve, which onboarding flow variant to route a new user into. If you ship one RL-flavored system in your career, it is more likely to be a bandit than a deep policy network — which is exactly why interviewers probe it so directly, and why it is the natural place to start turning MDP theory into something you actually deploy.

A bandit is an MDP with the sequential structure stripped out: there is exactly one state (or the state resets every round), a fixed set of arms (actions) to choose from, and each choice yields an immediate reward — no notion of the choice changing tomorrow's situation. If the notification agent in the foundations subject had to reason about fatigue building up over a user relationship, a bandit choosing which subject line to send this instant has no such carry-over: each impression is (approximately) an independent trial. That simplification is not a limitation to apologize for — it is what makes bandits fast to reason about, cheap to implement, and, critically, the subject of clean theoretical guarantees that most of deep RL cannot offer. Contextual Bandits for Personalization, next in this track, adds back a notion of "state" in the form of per-round context features; this subject is the context-free foundation those algorithms build on.


The Explore/Exploit Tradeoff

Picture five candidate subject lines for a re-engagement notification, call them A through E, each with some true (unknown) open rate. You get to send one subject line per user, observe whether they opened it, and you want to maximize total opens over, say, the next million sends. The problem: you do not know the true open rates, and the only way to learn them is to send each one enough times to measure it — but every send you spend on a low-performing subject line to learn that it's low-performing is a send you didn't spend on the best one.

This is the explore/exploit tradeoff in its purest form: exploit means send whatever currently looks best, based on the data so far; explore means send something else, purely to reduce uncertainty about arms you haven't tried enough. Pure exploitation risks getting permanently stuck on an arm that looked good by chance in the first few trials (a subject line that got 3 opens out of 5 by luck looks amazing, but 5 samples is nothing). Pure exploration never capitalizes on what you've learned. Every bandit algorithm in this subject is a different mathematically principled answer to "how do I balance the two."

graph TD
    EXPLOIT["pure exploitation<br/>(greedy, can get stuck)"] --> GOOD["the good policies live here<br/>ε-greedy, UCB1, Thompson sampling"]
    EXPLORE["pure exploration<br/>(random, never capitalizes)"] --> GOOD

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.