Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Training Reasoning Models: STaR, RLVR, and Reward Models (15 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

STaR's Rationalization vs. Plain Rejection Sampling on a Hard Dataset Permalink →

Your team is bootstrapping reasoning-trace training data for a model on a dataset of word problems where 70% are "easy" (the base model already gets these right most of the time when sampled) and 30% are "hard" (the base model almost never reaches the correct answer unassisted, even across many samples). Two options: (a) plain rejection-sampling fine-tuning (RFT) — sample many attempts per question, keep only ones with a verified-correct final answer; or (b) STaR with rationalization — same as (a), but for questions where no sampled attempt succeeds, show the model the correct answer and ask it to construct a plausible reasoning trace toward it.

  1. Predict what happens to the training set's composition under plain RFT (option a) specifically for the hard 30%, and why.
  2. Explain what rationalization changes about that outcome, and name the concrete risk it introduces that plain RFT does not have.
  3. Would you use rationalized traces at the same weight/frequency in training as genuinely-discovered correct traces? Justify your answer.

Share this question

Advanced Open Pro

Would GRPO's Group-Relative Baseline Work for a Robotics RL Task?

Unlock this question →
Advanced Open Pro

ORM or PRM for a Multi-Step Geometry Proof Assistant?

Unlock this question →
Advanced Open Pro

Diagnosing Length Inflation vs. Reward Hacking in a Training Run

Unlock this question →
Advanced Open Pro

Auditing a Reward Design for a Code-Reasoning Model

Unlock this question →
Advanced Open Pro

GRPO's Advantage Formula When an Entire Sampled Group Gets the Same Reward

Unlock this question →
Advanced Open Pro

Shipping R1-Zero-Style Pure RLVR Directly to a Customer-Facing Chatbot

Unlock this question →
Advanced Open Pro

Distilling a Reasoning Teacher into a 3B On-Device Triage Model

Unlock this question →
Advanced Open Pro

Choosing How to Build Self-Refinement Training Data When Correct Unassisted Traces Are Rare

Unlock this question →
Advanced Open Pro

Designing Staged Partial Credit for a Long, Multi-Function Code-Reasoning Task

Unlock this question →
Advanced Open Pro

GRPO Advantages Under Partial Credit: When 11/12 Tests Passed Gets Pushed Down as Hard as 0/12

Unlock this question →
Advanced Open Pro

A Math Policy Learns to Emit Two Boxed Answers: Exploiting a Loose Answer Extractor

Unlock this question →
Advanced Open Pro

Iterated Rejection-Sampling Fine-Tuning Drifting Toward the Easiest Solution Style

Unlock this question →
Advanced Open Pro

Choosing a Search-Trace Source for Stream-of-Search Training Without Baking In Overthinking

Unlock this question →
Advanced Open Pro

Swapping a Rule-Based Verifier for a Dense PRM Reward — and Watching Accuracy Fall

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.