Match a job Paths Subjects Questions Quizzes Pricing
Advanced Pro

Training Reasoning Models: STaR, RLVR, and Reward Models Quiz 3

A third, scenario-driven pass through reasoning-model training: computing a concrete GRPO advantage from a mixed group, the ORM blind spot for lucky-but-unsound traces, what Math-Shepherd rollout labels do and don't guarantee, telling reward hacking apart from a verifier coverage gap, scoring timeouts and crashes explicitly in a code reward, when an external self-refine loop becomes redundant, how distillation shifts the best-of-N trade-off, R1-Zero's emergent backtracking versus deliberately supervised Stream-of-Search, why RFT deduplicates near-identical traces, and what KL-anchoring actually buys against a proxy reward.

10 questions 15 min Pass: 70%

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.