Advanced
Open
Pro
Design a Dyna-Style Architecture for a Costly Real-World Environment
A warehouse-robotics team is training a Q-learning-based policy where each real robot-arm trial costs roughly $2 in wear, electricity, and supervised operator time, and takes 8 seconds. They currently need about 4,000 real trials to reach an acceptable success rate using pure model-free Q-learning — about $8,000 and 9 hours of real robot time per training run, which they run often during development.
- Design a Dyna-style architecture for this team: describe the three update steps that happen per real interaction and what each one does.
- If they add Dyna-style planning with k = 15 simulated updates per real step, and the model becomes reasonably accurate after about the first 300 real steps, roughly characterize the real-step savings you'd expect and why the savings isn't a clean 16x reduction.
- Their model is noticeably less accurate for arm configurations near the edge of the workspace (rarely visited during early training). What risk does this create if k is set high, and how would you mitigate it?
Share this question