RL Foundations: MDPs & Value Functions
The formal language of sequential decisions — states, rewards, returns, and the Bellman equations that make them computable
Learn how to model a real decision problem — a messaging agent choosing whether to send a notification — as a Markov Decision Process, define returns and discounting with worked numbers, derive the Bellman expectation and optimality equations from first principles, and use policy iteration and value iteration to compute optimal behavior.
Practice questions (5)
-
View →
Design the MDP for a Notification Send/Wait Agent
Intermediate · Free -
View →
Choosing and Justifying the Discount Factor
Intermediate -
View →
Solve a Small Bellman System by Hand
Intermediate -
View →
Policy Iteration vs. Value Iteration for a Send-Time Optimizer
Intermediate -
View →
Diagnose a Non-Markov State in a Live Notification System
Advanced