Match a job Paths Subjects Questions Quizzes Pricing

Case Study: Notification Timing & Content Optimization

Walk through a full ML system design interview answer for an RL-based send/delay/suppress policy

Overview Read

Case Study: Notification Timing & Content Optimization

"Design our notification system" sounds like a small feature question until you notice what it is actually asking: decide, for every user, every day, whether to interrupt them, with what, and when — using a decision that only pays off (or doesn't) days or weeks later, in a currency you can't cheaply measure (did this user stay?). It is one of the cleanest real-world settings for reinforcement learning outside of games, which is exactly why it shows up as a system design prompt at consumer companies with a notification surface: messaging apps, social feeds, delivery and ride-hailing apps, subscription products fighting churn. And it is a trap for candidates who reach for the first metric they can measure — open rate — because a policy that optimizes open rate is trivially easy to build and reliably makes the product worse.

This subject is a model interview answer, structured with the same step progression used throughout this track's system design material — clarify requirements → translate the business goal into an ML objective → state/action/reward design → data and model progression → offline evaluation gate → production guardrails → experiment design — the same shape as the ML System Design Framework subject's approach, specialized here for a decision-and-control problem rather than a pure prediction problem. It applies ideas taught elsewhere in this track rather than re-deriving them: state/action/reward formalism from the RL Foundations: MDPs & Value Functions subject, reward shaping and credit assignment from Reward Design & Delayed Credit Assignment, the bandit-to-RL progression from Multi-Armed Bandits & Exploration and Contextual Bandits for Personalization, the mandatory pre-launch check from Off-Policy Evaluation & Offline RL, and the serving discipline — propensity logging, guardrails, exploration budgets — from RL in Production: Safe Exploration & Serving. Read those subjects first if any term here is unfamiliar; this case study's job is to show how they fit together in one concrete system, not to teach them from scratch.

Read it as the answer you would give in 40 minutes, then use the follow-up questions near the end to pressure-test yourself.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.