How LLMs Are Built: Pretraining to Chatbot
The full pipeline from raw web data to a deployed assistant, and how to reason about the growing menu of open-weight and closed model families
The pipeline story an AI Engineer interview expects you to hold in one piece: web-scale data collection and dedup, tokenizer training, the pretraining objective and scaling laws (a worked compute-optimal tradeoff), then base model to SFT to preference tuning as the stages that turn raw next-token prediction into a deployed chatbot. Plus a model-family comparison — closed vs. open-weight, dense vs. mixture-of-experts, licensing, context and pricing tradeoffs — and a worked decision scenario for choosing a model family under real constraints.
Practice questions (5)
-
View →
Why Dedup Instead of Just Training for Fewer Epochs?
Advanced · Free -
View →
Sizing a Model Under a Fixed Training-Compute Budget
Advanced -
View →
Why Not Skip SFT and Go Straight From Base Model to DPO?
Advanced -
View →
Does Mixture-of-Experts Actually Save You Memory?
Advanced -
View →
A Data-Residency Constraint That Collapses the Decision Before Capability Even Matters
Advanced