Reasoning Models and Inference-Time Scaling
Every model provider now ships a "thinking" tier — OpenAI's o-series and GPT-5 thinking, DeepSeek-R1, Claude's extended thinking — and every one of them is priced and latency-budgeted differently from the standard model sitting next to it in the same API. Interviewers ask about reasoning models for exactly that reason: it is a live, current architecture-and-cost decision, not a history lesson. The candidate who says "reasoning models think more before answering" is repeating a marketing sentence. The candidate who can say what changed mechanically between a prompted model and a trained one, name the specific family of techniques that let you buy more accuracy with more inference compute, and — critically — say when that trade is a bad one, is giving the answer a senior AI engineer gives.
This subject assumes you already have the CoT baseline from prompt-engineering — the idea that generating intermediate tokens before an answer gives the model more computation to work with — and the technique-level detail in advanced-prompting-techniques — self-consistency's majority vote, ReAct's grounded loop, tree-of-thought's search shape, and the closing section on how reasoning models change the calculus for manually-triggered CoT. Neither of those subjects re-derives here. What this subject adds is the piece neither of them owns: a precise account of what a reasoning model actually is under the hood, and the full inference-time-scaling family — parallel sampling, sequential revision, tree search, and search against a trained verifier — treated as one continuum of "spend more compute at answer time" techniques, with the arithmetic to reason about when each is worth its cost. The sibling subject, training-reasoning-models-star-rlvr-and-reward-models, covers how these models and their verifiers are actually trained (STaR, RLVR, ORM/PRM, distillation) — this subject is deliberately scoped to what happens at inference time, with training treated as a black box that hands you a model and, sometimes, a verifier.