Value-Based Methods: Q-Learning to DQN
From tabular TD learning to Deep Q-Networks — and why the combination almost doesn't work
Learn temporal-difference learning, tabular Q-learning and its off-policy Bellman update, why tables collapse at scale, function approximation with neural networks, and the experience replay, target network and double-Q fixes that make Deep Q-Networks trainable — framed around the deadly triad of function approximation, bootstrapping and off-policy learning.
Practice questions (6)
-
View →
Trace a Manual Q-Learning Update for a Cloud Autoscaler
Advanced · Free -
View →
Migrate a Tabular Agent to Function Approximation
Advanced -
View →
Diagnose a Diverging Training Run
Advanced -
View →
Quantify and Fix Overestimation Bias in a Trading Agent
Advanced -
View →
Design and Debug a Replay Buffer for a Recommendation Agent
Advanced