Match a job Paths Subjects Questions Quizzes Pricing

How LLMs Are Built: Pretraining to Chatbot

The full pipeline from raw web data to a deployed assistant, and how to reason about the growing menu of open-weight and closed model families

Overview Read

How LLMs Are Built: Pretraining to Chatbot

Most AI Engineer candidates can name the stages — "pretraining, then fine-tuning, then RLHF" — the way most people can name the planets in order. Fewer can hold the whole pipeline in one piece: why deduplication matters enough to be a dedicated engineering effort, what "compute-optimal" actually trades off against what, and why a model that costs $20 in API credits and a model you can download and run in your own VPC are not just priced differently but built differently in ways that predict where each one breaks. Interviewers ask about this pipeline not to test trivia recall, but to see whether you can reason about a new model release the day it drops — is it dense or MoE, what's the licensing story, does "open weight" mean what the press release implies — instead of waiting for a benchmark leaderboard to tell you what to think.

This subject tells that pipeline as one story, end to end, and then does something the pipeline story alone doesn't cover: comparing model families — closed vs. open-weight, dense vs. mixture-of-experts — as an engineering decision with real tradeoffs, not a popularity contest. Two things this subject deliberately does not re-explain: the training-algorithm mechanics of SFT, LoRA/QLoRA, RLHF and DPO, which fine-tuning-sft-lora-rlhf-dpo covers in full depth — this subject tells you where each stage sits in the bigger pipeline and cross-links out for how each stage works internally; and subword tokenizer mechanics (BPE merges), which tokenization-and-context-windows owns — this subject touches tokenizer training only long enough to place it correctly in the pipeline. Exact parameter counts and benchmark scores are treated as background color here, not facts to memorize — they date within a model generation or two, while the pipeline shape and the architectural tradeoffs stay true much longer.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.