Justify a Contextual Bandit for Send-Time Personalization
Your team currently uses a plain (context-free) multi-armed bandit with 4 arms — send now, in 2h, in 6h, tomorrow morning — for notification send-time selection. A analysis shows the best arm differs substantially between "night owl" users (peak engagement 11pm–1am) and "early bird" users (peak engagement 6–8am), and these two segments are roughly equal in size.
- Explain concretely what a plain bandit converges to in this situation, and quantify (qualitatively) how much value is left on the table.
- Describe how you would restructure this as a contextual bandit, including a specific context feature vector.
- A colleague argues "we could just run two separate plain bandits, one for night owls and one for early birds, and get the same benefit without learning a contextual model." Evaluate this proposal, including what breaks as you add more segments.
1. What the plain bandit converges to
A plain bandit has no way to distinguish night owls from early birds — both populations produce rewards that get pooled into the same per-arm estimate. It will converge to whichever single arm has the best population-average reward across both segments combined. If, say, "in 6h" happens to average decently for both groups without being the true best for either, the bandit locks onto "in 6h" for everyone — even though the actual best arm for night owls might be "send now" (during their peak) and for early birds might be "tomorrow morning" (waiting for their peak). The value left on the table is, roughly, the gap between each segment's true best-arm reward and the reward of whatever compromise arm the plain bandit settles on — potentially a large fraction of the segment-specific upside, since the compromise arm is optimized for neither segment specifically.
2. Restructuring as a contextual bandit
Keep the same 4 arms, but add a context vector capturing whatever
distinguishes the segments — at minimum something like
x = [hour_of_day_sin, hour_of_day_cos, days_since_last_open, historical_peak_hour_estimate, 1] (sine/cosine encoding of
hour-of-day avoids the discontinuity of raw hour values, and a bias
term of 1). A LinUCB-style model then learns a separate (or shared +
per-arm, in the hybrid formulation) linear reward model \theta_a
per arm as a function of this context, so it can learn, e.g., that
"send now" has meaningfully higher predicted reward specifically when
the context indicates late evening, without needing an explicit
"segment" label — the continuous hour/recency features let the model
discover the two behavior patterns (and any others present) directly
from data.
3. Evaluating the "two separate plain bandits" proposal
Running two separate plain bandits, keyed by a hand-labeled segment, would in fact capture most of the benefit for exactly two, cleanly-defined segments — it's a reasonable quick fix. But it doesn't scale: (a) it requires manually defining and maintaining segment boundaries rather than letting the model learn continuous structure (what about users who are neither clearly night owls nor early birds?); (b) each additional segment (device type, day of week, lifecycle stage, geography) multiplies the number of separate bandits needed, and each one starts cold and needs its own sufficient traffic to converge — a combinatorial explosion as more segmenting dimensions are added, with less and less data per bucket; and (c) it shares no information across segments — a night-owl bandit and an early-bird bandit learn from scratch independently even though the send-time buckets and much of the reward structure are the same object being viewed from two contexts. A contextual bandit handles this uniformly: it can incorporate many context dimensions simultaneously (continuous, not manually bucketed), pool information across similar contexts via its shared parameterization, and needs no manual segment definition or maintenance as more context signals are identified.
Share this question