Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Pro

Choosing and Justifying the Discount Factor

Your team is debating the discount factor for the notification agent. One engineer argues for \gamma = 0.5 ("we mostly care about the next few notifications"); another argues for \gamma = 0.99 ("churn is a long-term problem, and we should account for it fully").

  1. Using the effective-horizon interpretation of \gamma, quantify roughly how many future notifications each choice weighs meaningfully.
  2. Given a trajectory where sending now yields r_{t+1} = +1 but raises mute probability such that the expected reward 4 steps later is r_{t+5} = -15 (conditional on reaching that step), compute the discounted contribution of that -15 under both \gamma values and explain what this implies about each policy's behavior.
  3. What downside is there to always choosing the largest possible \gamma close to 1?

Share this question

← Back to RL Foundations: MDPs & Value Functions practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.