Match a job Paths Subjects Questions Quizzes Pricing

Value-Based Methods: Q-Learning to DQN

From tabular TD learning to Deep Q-Networks — and why the combination almost doesn't work

Overview Read

Value-Based Methods: Q-Learning to DQN

A video-transcoding platform runs an autoscaling agent in front of its GPU worker fleet. Every few seconds it observes the job queue depth and the current fleet size, and picks one of three actions: add a worker, remove a worker, or hold. Adding workers costs money immediately; removing them risks a queue backlog and SLA penalties a few minutes later; holding is free but does nothing about a queue that is already growing. A supervised model that predicts "will the queue be long in 5 minutes" cannot make this decision on its own — it has no notion of which action, taken now, leads to the best sequence of outcomes over the next hour, including the downstream effect that scaling up now changes the queue state the agent will face at the next step. This is a reinforcement learning problem: an agent choosing actions in a Markov decision process (MDP), trying to maximize cumulative future reward, defined in full in the RL Foundations: MDPs & Value Functions subject — states, actions, rewards, transition dynamics, discount factor \gamma, and the value functions V^\pi(s) and Q^\pi(s,a).

This subject is about learning those value functions from experience, without ever being handed the transition dynamics, and then acting greedily with respect to them. That family — temporal-difference learning, Q-learning, and its deep-network descendant DQN — is the value-based branch of RL: it learns "how good is each action here" and derives a policy by picking the best one. It sits in contrast to the policy-based branch covered in the Policy Gradients & Actor-Critic subject, which skips the value function and directly optimizes the action-choosing policy; that subject opens with the same comparison from the other side. The reason both branches exist in every RL curriculum, and in every interview rubric, is that value-based methods are sample-efficient and stable in small, discrete-action problems but become awkward — and in continuous action spaces, intractable — exactly where policy-based methods are natural, and vice versa.

By the end of this subject you should be able to derive the Q-learning update on a whiteboard, explain in one sentence why it is off-policy, explain why a table of |S| \times |A| numbers cannot survive contact with a real state space, and — the part interviewers spend the most time on — name the specific failure mode that appears when you plug a neural network into that update, and the three engineering fixes (experience replay, target networks, double Q-learning) that make Deep Q-Networks trainable in practice.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.