Reasoning Models and Inference-Time Scaling
o-series, DeepSeek-R1, and extended thinking, and the family of techniques — parallel sampling, sequential revision, tree search, search against a verifier — that spend more inference compute for a better answer
A precise treatment of reasoning models for AI engineering interviews: what actually makes a 'thinking' model different from a standard model given a CoT prompt (visible vs. hidden traces, trained thinking-token budgets); the inference-time-scaling technique family building on the chain-of-thought baseline — best-of-N and self-consistency, iterative self-refinement, tree-of-thought search, and search against a trained verifier; a worked, real-numbers comparison of one large-model call against many small-model samples plus a verifier at equal compute; and the senior-engineer judgment call of when paying for a reasoning model or extra inference-time compute is worth it versus when it actively hurts.
Practice questions (5)
-
View →
Is This Actually a Reasoning Model, or Just a Good CoT Prompt?
Advanced · Free -
View →
Best-of-N With a Verifier That Isn't as Good as You Think
Advanced -
View →
Self-Refinement Loop or Tree Search for a Contract-Clause Rewriter?
Advanced -
View →
Sizing a Compute-Matched Comparison for a Code-Generation Feature
Advanced -
View →
Vetoing a Reasoning Model for a Live Autocomplete Feature
Advanced