Advanced
Open
Pro
Design and Debug a Replay Buffer for a Recommendation Agent
A team is building a DQN agent that decides which of 8 content "slots" to fill next in a feed, treating each user session as a sequence of state transitions with reward tied to engagement. They report that after adding a replay buffer and target network, training became stable but the agent's policy plateaus at a noticeably suboptimal level, and offline evaluation shows it rarely recommends 2 of the 8 slot types at all, even though spot-checks suggest those slot types would perform well in specific contexts.
- Give two plausible replay-buffer or exploration-related causes for an agent that trains stably but converges to a narrow, suboptimal policy, and explain the mechanism for each.
- The team's replay buffer is a fixed-size ring buffer of the most recent 50,000 transitions, and their ε-greedy schedule anneals from 1.0 to 0.02 over the first 20,000 steps of a training run that lasts 2,000,000 steps. What is the problem with this combination specifically?
- Propose a concrete fix, including one alternative to plain uniform replay sampling, and explain what trade-off it introduces.
Share this question