Match a job Paths Subjects Questions Quizzes Pricing
Advanced Pro

Training Reasoning Models: STaR, RLVR, and Reward Models Quiz 2

A second pass through how reasoning models are trained: rejection-sampling fine-tuning vs. STaR, the plateau that motivates moving to RL, the full DeepSeek-R1 pipeline's stages, GRPO's literal advantage formula, how ORM training data is produced, why distillation skips RL entirely, what a Stream-of-Search training example actually contains, why RLVR concentrates in math and code, length-inflation mitigations, and what a passing test suite doesn't prove.

10 questions 15 min Pass: 70%

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.