Match a job Paths Subjects Questions Quizzes Pricing
Machine Learning Intermediate Pro

RL Foundations: MDPs & Value Functions

The formal language of sequential decisions — states, rewards, returns, and the Bellman equations that make them computable

30 min read 20 views

Learn how to model a real decision problem — a messaging agent choosing whether to send a notification — as a Markov Decision Process, define returns and discounting with worked numbers, derive the Bellman expectation and optimality equations from first principles, and use policy iteration and value iteration to compute optimal behavior.

Practice questions (5)

  • Design the MDP for a Notification Send/Wait Agent

    Intermediate · Free
    View →
  • Choosing and Justifying the Discount Factor

    Intermediate
    View →
  • Solve a Small Bellman System by Hand

    Intermediate
    View →
  • Policy Iteration vs. Value Iteration for a Send-Time Optimizer

    Intermediate
    View →
  • Diagnose a Non-Markov State in a Live Notification System

    Advanced
    View →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.