Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Diagnose a Planning Failure Caused by Compounding Model Error

A team trains a learned dynamics model for a robot arm and uses it to plan 30-step action sequences via trajectory optimization (choose the sequence the model predicts leads to the highest cumulative reward), executing the full 30-step plan open-loop before re-planning. In simulation-against-the-model, the plans look excellent — high predicted reward. When executed on the real arm, performance is mediocre and highly inconsistent.

  1. Explain the most likely mechanism behind the gap between predicted-in-model performance and real performance.
  2. Propose two independent mitigations, and explain the mechanism by which each one addresses the problem you identified in part 1.
  3. The team could also invest in a much larger, more accurate model instead of changing the planning procedure. Explain why this alone may not fully solve the problem, using the idea of distribution shift between the policy's visited states and the model's training data.

Share this question

← Back to Model-Based RL & Planning practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.