Advanced
Open
Pro
Diagnose a Diverging Training Run
A team trains a DQN-style agent for a warehouse-robot task. They implement the semi-gradient Q-learning update directly — no replay buffer, no target network — training online on each transition as it arrives, using an ε-greedy behavior policy. After a few thousand steps the loss explodes and the Q-values diverge toward large magnitudes.
- Identify all three legs of the deadly triad present in this setup, mapping each one to a specific piece of their implementation.
- If they fixed only bootstrapping (e.g. by switching to a Monte-Carlo target — the full observed return — while keeping the neural network and the ε-greedy off-policy behavior), would you expect the instability to go away? What would they give up?
- Propose the minimal two changes that would let them keep Q- learning's bootstrapped, off-policy update while regaining practical stability, and explain the mechanism by which each one helps.
Share this question