Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Apply Per-Decision Importance Sampling to a Multi-Step Session

A 3-step notification-sequencing agent logs one user's session: at each of 3 decision points the agent chose whether/what to send, with behavior-policy propensities \pi_b = [0.5, 0.4, 0.6] and the candidate target policy's propensities for those same logged actions \pi_e = [0.6, 0.2, 0.6] at steps 1, 2, 3 respectively. Rewards observed at each step were r = [0, 1, 2], with \gamma = 0.9.

  1. Compute the full-trajectory importance weight (the product of all three per-step ratios) and the resulting full-trajectory IS contribution for this single episode.
  2. Compute the per-decision importance sampling (PDIS) contribution for this same episode, reweighting each reward only by the product of ratios up to and including that step. Compare the two results and explain why they differ.
  3. Explain, using this example, why evaluating a 20-step session with full-trajectory IS would be expected to be dramatically worse (in variance terms) than evaluating this 3-step one, even if the per-step ratios were of similar typical magnitude throughout.

Share this question

← Back to Off-Policy Evaluation & Offline RL practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.