Paths Subjects Questions Quizzes Pricing Search
Intermediate Open Pro

Sizing Self-Consistency Under a Fixed Compute Budget

You're building a math-word-problem tutor that grades a student's final numeric answer against the model's own computed answer for the same problem. You have a fixed budget of 8x the tokens of a single CoT call per graded problem, and two options: (a) self-consistency with N=8 single CoT chains, majority vote; (b) N=4 chains at double the reasoning-token allowance each (i.e., longer individual chains, still totaling ~8x tokens).

  1. Explain the mechanism each option is betting on, and why they are not interchangeable even though they cost the same.
  2. Describe the empirical test that would tell you which option is better for this specific problem set, rather than picking by intuition.
  3. A colleague suggests skipping the whole comparison and just using a reasoning model with an extended-thinking budget instead of either option. Is that a reasonable alternative, and what would you still want to verify before shipping it?

Share this question

← Back to Advanced Prompting Techniques practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.