Match a job Paths Subjects Questions Quizzes Pricing
← All paths

AI Engineer Interview

What it takes to design and ship production LLM applications — and to explain the reasoning in an interview. Starts with the model fundamentals (transformers, tokenization, the full pretraining-to-chatbot pipeline, prompting, fine-tuning) and reasoning models (inference-time scaling, training reasoning models with RLVR and reward models), then context engineering and multimodal models, retrieval-augmented generation end to end (architecture, chunking and embeddings, vector search, evaluation), agents and tool use (workflow design patterns, the agentic loop, memory, multi-agent orchestration, MCP, agent evaluation), and production concerns (guardrails and prompt-injection defense, cost and latency engineering, observability, local and open-weight deployment). Surveys the framework landscape, then finishes with five full worked designs: a customer-support assistant, a web-search agent, a deep-research agent, a Claude Code-shaped coding agent, and an RLHF data and training platform.

1 of 36 subjects free

Who it's for

Built for anyone who wants a structured, ordered path through AI Engineer Interview — 36 subjects, free to start, at your own pace.

0 of 36 subjects complete 0%
Start: Transformers for AI Engineers →

Sign up free to save your progress through this path.

What you'll learn

  1. 1

    Transformers for AI Engineers

    Covers what an AI Engineer interview actually probes about transformer internals: scaled dot-product and multi-head attention, why decoder-only architectures won, why the KV cache exists and how its memory footprint is computed, what parameter count does and doesn't predict, and positional encoding — with worked numbers for KV cache memory and prefill-vs-decode cost.

    Free Start →
  2. 2

    Tokenization and Context Windows

    A deep, numbers-first look at subword tokenization (BPE), why tokens are not words or characters, the 'lost in the middle' effective-context problem, worked token-budget arithmetic for a real prompt, and why long context windows do not make RAG obsolete. Written as an AI Engineer interview reference with concrete worked examples and comparison tables.

    Pro Start →
  3. 3

    How LLMs Are Built: Pretraining to Chatbot

    The pipeline story an AI Engineer interview expects you to hold in one piece: web-scale data collection and dedup, tokenizer training, the pretraining objective and scaling laws (a worked compute-optimal tradeoff), then base model to SFT to preference tuning as the stages that turn raw next-token prediction into a deployed chatbot. Plus a model-family comparison — closed vs. open-weight, dense vs. mixture-of-experts, licensing, context and pricing tradeoffs — and a worked decision scenario for choosing a model family under real constraints.

    Pro Start →
  4. 4

    Prompt Engineering

    A practitioner's tour of prompt engineering as an AI Engineer interview topic: what belongs in the system prompt vs the user prompt and why, when few-shot examples help and when they stop paying off, what chain-of-thought actually buys you mechanically, structured outputs and schema-constrained decoding, treating prompts like versioned code with regression tests, and the decision framework for when a longer prompt is the wrong answer and RAG or fine-tuning is the right one.

    Pro Start →
  5. 5

    Advanced Prompting Techniques

    The technique landscape beyond few-shot and basic chain-of-thought: breaking a task into subtasks, sampling and voting across multiple reasoning chains, interleaving reasoning with tool calls (ReAct), searching over a tree of partial solutions, when a persona measurably helps versus is theater, letting a model or optimizer write the prompt for you (DSPy and friends), and what changes once the model itself does extended, budgeted reasoning at inference time. Each technique is presented the way an interviewer expects: what it costs, what it buys, and the concrete signal that tells you it's the wrong tool for the task in front of you.

    Pro Start →
  6. 6

    Fine-Tuning: SFT, LoRA/QLoRA, RLHF and DPO

    A deep, standalone treatment of fine-tuning for AI engineering interviews: the mechanics of SFT, LoRA/QLoRA, RLHF and DPO; realistic data requirements and where the data comes from; the sharp line between behaviour problems (fine-tune) and knowledge problems (RAG); training and serving cost order-of-magnitude; how to evaluate a fine-tune against catastrophic forgetting and a prompted baseline; and a worked decision scenario for a formatting-adherence bug in a support bot.

    Pro Start →
  7. 7

    Reasoning Models and Inference-Time Scaling

    A precise treatment of reasoning models for AI engineering interviews: what actually makes a 'thinking' model different from a standard model given a CoT prompt (visible vs. hidden traces, trained thinking-token budgets); the inference-time-scaling technique family building on the chain-of-thought baseline — best-of-N and self-consistency, iterative self-refinement, tree-of-thought search, and search against a trained verifier; a worked, real-numbers comparison of one large-model call against many small-model samples plus a verifier at equal compute; and the senior-engineer judgment call of when paying for a reasoning model or extra inference-time compute is worth it versus when it actively hurts.

    Pro Start →
  8. 8

    Training Reasoning Models: STaR, RLVR, and Reward Models

    The training-time half of the reasoning-model story for AI engineering interviews: SFT on model-generated reasoning traces (STaR, rejection-sampling fine-tuning); reinforcement learning with verifiable rewards (RLVR) and the DeepSeek-R1 recipe; GRPO vs. PPO at a glance — what GRPO drops and why; outcome vs. process reward models and auto-labeled process supervision without human step-level annotation; self-refinement as a training objective, distinct from inference-time self-refine; internalizing search (Meta-CoT, Stream of Search) so a model can search within one generation instead of needing external scaffolding; distilling reasoning traces into smaller models; named failure modes (overthinking, reward hacking, verifier coverage limits); and a worked reward-design example for a code-reasoning model using unit tests as the verifier.

    Pro Start →
  9. 9

    Context Engineering Fundamentals

    The discipline that succeeds prompt engineering once a system has retrieval, tool calls, and conversation history: treating the entire context window — not just the prompt string — as an assembled, budgeted, ordered artifact. Covers the anatomy of a real app's context window, token-budget allocation across fixed and variable regions, selection and ordering strategies, compaction (summarization, truncation, structured notes), isolation between trust boundaries, the four named context failure modes with repro sketches, and a fully worked, real-numbers example of assembling one turn of a support bot's context.

    Pro Start →
  10. 10

    Context Engineering for Agents

    The capstone application of context engineering to the hardest case a practitioner faces: a long-running, tool-using agent instead of a single prompt. Covers why agents break naive context management (monotonic growth, unpredictable tool-output size, compounding cost and error), how to write system prompts that hold up over 100+ turns, why tool names/descriptions/schemas are prompt text with a per-turn cost, compaction and structured handoff between context windows, sub-agents as context isolation, just-in-time retrieval versus pre-loading (the Claude Code model), where steering files like CLAUDE.md fit, and how to evaluate agent context strategies with a worked 50-turn token-budget trace.

    Pro Start →
  11. 11

    Multimodal LLMs and Vision

    Learn how multimodal LLMs actually process images: patch tokenization, vision-encoder pretraining, and how the model distinguishes between multiple objects in a scene through attention rather than bounding-box regression. Covers the practical gap between vision-language models and classical object detectors, common failure modes (counting, fine-grained discrimination, spatial relations), and prompting techniques (referring expressions, crops, set-of-mark) that make multi-object questions reliable in production.

    Pro Start →
  12. 12

    Text-to-Image Generation with Diffusion Models

    The deepest subject in this track's generative-image sequence: why diffusion, not autoregressive generation, is the default choice for text-to-image quality and a tunable steps-vs-quality dial; caption engineering at 500M-pair scale (BLIP-style re-captioning, CLIP-score filtering); the U-Net (downsampling/upsampling blocks with cross-attention) and DiT (patchify -> Transformer -> unpatchify) architectures; the forward/backward diffusion process and the predict-the-noise training objective at interview altitude; classifier-free guidance and DDIM step reduction as the two sampling techniques worth having ready; CLIP and CLIPScore introduced properly, once, for this subject and its two upcoming siblings to reuse; and the full production system — data, training, optimization, and inference pipelines, including the prompt-safety-to-super-resolution inference chain.

    Pro Start →
  13. 13

    Personalizing Image Generation: DreamBooth, LoRA and Textual Inversion

    The most product-shaped subject in this track's generative-image sequence: personalizing a pretrained diffusion model (`text-to-image-diffusion-models`) to one specific subject via three tuning methods on a single trade-off ladder — textual inversion (one new token embedding, cheap and weak), DreamBooth (full-model fine-tuning with a rare-token identifier and a class-specific prior preservation loss against overfitting and catastrophic forgetting, best fidelity), and LoRA (low-rank adapters, the same trick `fine-tuning-sft-lora-rlhf-dpo` teaches for LLMs, applied here to a diffusion U-Net) — argued and resolved for an AI-headshots product. Covers quality-gated data prep for a handful of user photos, the combined reconstruction-plus-prior-preservation training objective, hand-engineered sampling prompts, identity-fidelity evaluation (CLIP-I, DINO and dedicated face-recognition similarity, with the CLIP-vs-DINO reasoning spelled out), the asynchronous fine-tune-then-generate-then-verify production pipeline contrasted with subject 2's synchronous inference chain, and a deepfake/consent/PII safety sidebar.

    Pro Start →
  14. 14

    Case Study: Generative Fill — Inpainting, Outpainting and Object Removal

    Model interview answer for Generative Fill (remove / extend / generate-with-prompt): why the right design is a cascade — a LaMa-style single-pass GAN for object removal, mask-conditioned latent diffusion for prompt-driven fill, outpainting as inpainting with the mask outside the frame — argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject; crop-around-the-mask as the scaling trick that makes 8K interactive editing possible; seam-blending and a masked-region-specific evaluation suite (FID is not enough); and a system design with content-credential signing as a first-class pipeline stage.

    Pro Start →
  15. 15

    RAG Architecture End to End

    A deep, standalone treatment of RAG architecture: the offline ingestion path and online query path as one diagram, a named failure mode at every stage, why the 'naive RAG' first pass plateaus at mediocre quality, query rewriting and multi-query fan-out, the confidence gate and the 'I don't know' path, citations and grounding checks, advanced patterns (parent-child, HyDE, agentic retrieval) and when they earn their complexity, and a RAG maturity ladder for the 'how would you improve this system' follow-up.

    Pro Start →
  16. 16

    Chunking and Embedding Strategies

    Deep dive on the two decisions that most determine RAG retrieval quality: chunking strategy (fixed-size, structure-aware, semantic, parent-child), overlap, chunk size vs recall, embedding model selection, and embedding drift when you swap models — with worked recall@k examples.

    Pro Start →
  17. 17

    Vector Databases and Hybrid Search

    A systems-level treatment of the retrieval layer for RAG: pgvector vs dedicated vector databases (Pinecone, Weaviate, Chroma, Milvus) with a comparison table and a 'when Postgres is enough' decision rule; HNSW vs IVF ANN indexing and their recall/latency trade-offs; pre-filter vs post-filter vs filtered-ANN and why naive post-filtering silently starves results; hybrid dense+BM25 search with reciprocal rank fusion (RRF) worked by hand; and cross-encoder re-ranking with its O(k) cost model.

    Pro Start →
  18. 18

    Evaluating RAG Systems

    A deep, interview-ready treatment of RAG evaluation: why retrieval and generation must be measured separately, worked recall@k/MRR/NDCG@k arithmetic, generation metrics including faithfulness and answer relevance, how to build and version a gold evaluation set, LLM-as-judge rubric design and bias calibration, RAGAS-style automated pipelines, hallucination-detection techniques, and a worked diagnostic example that separates a retrieval problem from a generation problem.

    Pro Start →
  19. 19

    Tool and Function Calling

    Deep dive on tool/function calling: the request-execute-return loop that turns a text generator into something that can act, why schema and description quality is the real bottleneck on tool-selection accuracy, tool-choice modes (auto/forced/none), parallel tool calls and their latency payoff, error handling and argument validation, where the permission boundary actually lives, keeping results compact, and a side-by-side of Anthropic tool use and OpenAI function calling.

    Pro Start →
  20. 20

    Agent Design Patterns: From Workflows to Autonomous Agents

    The vocabulary layer above the agentic loop: the LLM-vs-workflow-vs-agent distinction and agency levels as a spectrum; the five named workflow patterns (prompt chaining, routing, parallelization's sectioning and voting sub-variants, orchestrator-worker, evaluator-optimizer) each with a mermaid diagram and a one-line heuristic; autonomous-agent techniques beyond plain ReAct (Reflexion's cross-attempt episodic memory, ReWOO's plan-then-execute-without-interleaving, LATS-style tree search over agent action sequences); and a worked cost/latency/reliability comparison of one product feature implemented as a single prompt, a prompt chain, and a full ReAct agent.

    Pro Start →
  21. 21

    Agent Architectures and the Agentic Loop

    A precise, vendor-neutral treatment of the agentic loop for AI engineering interviews: the observe-decide-act loop itself, the ReAct pattern and why explicit reasoning traces help tool selection, the plan-act-observe-reflect variant and when its extra LLM calls are worth it, single-agent vs multi-agent at a glance, termination and budget controls, failure modes with mitigations (loops, tool hallucination, runaway cost, context poisoning, analysis paralysis), a worked compounding-error-rate calculation, and a worked workflow-vs-agent decision.

    Pro Start →
  22. 22

    Agent Harness Engineering

    A vendor-neutral treatment of the harness — everything in an agent that is not the model. Covers the Agent = Model + Harness decomposition and why the harness, not the loop, is the security and reliability boundary; the five components (executor, sandbox, permission and approval tiering, context/memory plumbing, observability); a traced walkthrough of one tool call through every layer; a worked coding-agent harness design with real permission tiers and budgets; the discipline boundary between harness engineering, prompt engineering, and orchestration; harness-specific failure modes (permission gaps, validation/execution mismatch, trust-boundary carryover, approval fatigue, unobserved cost blowup, sandbox escape) with mitigations and arithmetic; and how harnesses are audited and evolved from trajectory data.

    Pro Start →
  23. 23

    Memory Systems for LLM Applications

    A working taxonomy for one of the most conflated terms in AI engineering: the conversation window (short-term), persistent facts across sessions (long-term), what-happened-last-time (episodic), and externalized scratchpads (working memory). Covers durability criteria for what to persist, eager-load-vs-retrieval-triggered design, the failure modes — staleness, injection, bloat, conflicting facts — and a concrete walkthrough of how CLAUDE.md and context compaction implement exactly this taxonomy in a real product.

    Pro Start →
  24. 24

    Multi-Agent Orchestration

    A vendor-neutral treatment of multi-agent LLM systems: the supervisor/worker pattern, handoffs versus shared state, pipeline-vs-barrier synchronization for parallel fan-out, a worked cost-multiplication example, and the honest heuristic for when a single stronger model beats a fleet of coordinating agents. Cross-links `mcp-and-subagents` as the concrete Claude Code implementation and `plan-and-loop-modes` for the pipeline and adversarial-verification mechanics.

    Pro Start →
  25. 25

    MCP and Tool Integration Protocols

    The concept and design-interview layer above MCP: why a standard client-server protocol replaced bespoke per-app tool integrations, the server/client/resource/tool abstractions at an architectural level, what MCP adds on top of plain function calling (and when bespoke calling is still the right answer), the trust boundary a third-party server introduces, and how to frame the 'build a server vs write a function' decision when an interviewer asks.

    Pro Start →
  26. 26

    Evaluating Agents and Multi-Agent Systems

    Agent- and multi-agent-specific evaluation for AI engineering interviews: outcome vs. trajectory evaluation and why trajectory matters even when the answer is right; concrete agent metrics (tool-call accuracy, step efficiency, task completion rate); pass@k vs. pass^k explained and measured with a worked numeric example, framed as evaluation metrics rather than a design-time cost calculation; LLM-as-judge applied to trajectories specifically, with the biases unique to judging a process rather than an answer; sandboxed benchmark environments (τ-bench/SWE-bench-style) as the rigorous alternative; cost-per-successful-task as the unifying production metric; multi-agent-specific coordination failures, credit assignment, and handoff loss; agent-specific regression testing and canarying; and a worked 200-task eval set where pass@1 and pass^3 disagree on a ship decision.

    Pro Start →
  27. 27

    Guardrails and Prompt-Injection Defense

    A deep dive on the topic every AI Engineer interview probes: guardrails as a pipeline wrapped around the model, not a sentence in the prompt. Covers input filtering (PII redaction, injection classifiers, topic filters), prompt injection as the defining LLM threat and why it has no clean code/data separation, defense-in-depth layers with worked examples, output filtering (schema, groundedness, PII, refusals), the hard limits of any classifier against an adaptive attacker, the latency/cost budget for a guardrail pipeline, and a full worked trace of an attack through every layer.

    Pro Start →
  28. 28

    Cost and Latency Engineering for LLM Apps

    The reusable toolkit behind every 'make it cheaper and faster' interview question: why input tokens dominate cost and output tokens dominate latency, how prompt-prefix caching and semantic caching work and where each breaks, model routing and cascades, streaming and parallel tool calls as latency levers, the TTFT-plus-decode model of p95 latency, and the cost-per-1,000-conversations math that ties every lever to a number leadership will ask for.

    Pro Start →
  29. 29

    LLM Observability and Evaluation

    Interview-ready coverage of running LLM systems in production: what a full request trace must capture, token/cost/latency dashboards as product metrics, the precise line between offline evals and online monitoring, LLM-as-judge as a general technique with its biases and calibration, human review workflows, CI regression suites that treat prompts and models as code, canarying prompt and model changes, and a worked dashboard-diagnosis example.

    Pro Start →
  30. 30

    Local LLM Deployment and Open-Weight Serving

    A hands-on-adjacent, numbers-first tour of the local/open-weight serving stack for AI engineering interviews: what Ollama, llama.cpp/GGUF, vLLM, and LM Studio each actually are and when you'd reach for each; OpenAI-compatible local endpoints as the integration pattern that makes a local model swappable behind existing client code; quantization (GGUF Q4/Q8, AWQ/GPTQ, bitsandbytes, and FP8 on Hopper-class hardware) and its quality-vs-memory tradeoff; the full VRAM math for weights plus KV cache, worked for a 32B model at Q4 on a 24GB consumer GPU; continuous batching and PagedAttention; speculative decoding and how it compares to quantization and distillation as three different levers for cheaper, faster inference; and the decision framework for when self-hosting actually beats an API call.

    Pro Start →
  31. 31

    Frameworks Landscape: LangChain, LlamaIndex, CrewAI, AutoGen, Semantic Kernel, Assistants API

    A trade-off-first tour of the LLM application framework landscape — LangChain, LlamaIndex, CrewAI, AutoGen, Semantic Kernel and hosted Assistants-API-style platforms — organised around what each abstraction gives you, what it hides, and when the honest answer is to roll your own thin orchestration instead.

    Pro Start →
  32. 32

    Case Study: Design a Customer-Support Assistant

    Model interview answer for designing an LLM customer-support assistant: requirements and non-goals, workflow-vs-agent decision, tenant-safe RAG over a help centre, tool calls for account lookups and refunds with confirmation, conversation state, escalation and the 'I don't know' path, layered guardrails, offline and online evaluation, cost and latency numbers with levers, and a rollout plan.

    Pro Start →
  33. 33

    Case Study: Design an Ask-the-Web Agent

    Model interview answer for designing a Perplexity-style 'ask the web' agent: query understanding and rewriting, routing by query type, a search API call, parallel fetch and readability extraction of source pages, dedupe and rerank of retrieved sources, a freshness/staleness gate, a cite-grounded streaming synthesis step, citation faithfulness evaluation, prompt-injection defense specific to untrusted web content, and a worked numeric cost/latency budget showing why the fetch step must be parallelized to hit a real end-to-end target.

    Pro Start →
  34. 34

    Case Study: Design a Deep Research Agent

    Model interview answer for designing a 'deep research' agent that produces a multi-step research report rather than a quick answer: a three-stage architecture (clarification and planning, parallel sub-agent execution in a sandbox, synthesis with redundancy removal and a dedicated citation agent), the model-tier allocation decision at each stage (cheap model vs. reasoning model), budget and termination controls for the overall multi-stage pipeline, evaluating a finished report on coverage, faithfulness and citation precision, and a worked example allocating cost across a 4-5 sub-question research query.

    Pro Start →
  35. 35

    Case Study: Design a Coding Agent

    Model interview answer for designing a terminal coding agent that reads, edits, runs and verifies code in a repository: requirements and threat model, the agentic loop, a minimal tool set with output caps, a permission model that separates read from write from execute, context management for a finite window (truncation, compaction, subagents, cached prefixes), layered memory, cost and step budgets, prompt-injection defence, evaluation on task suites, and observability.

    Pro Start →
  36. 36

    Case Study: Designing an RLHF / Preference-Tuning Platform

    Model interview answer for the platform-engineering capstone behind every instruction-tuned LLM: prompt sourcing and privacy filtering, comparison-UI and annotator-calibration design, SFT to reward-model to policy-training pipelines (and the DPO variant that collapses two stages into one), reward-hacking-aware evaluation, and a system design where rollout generation on inference-engine workers, not the learner, is the dominant cost. Grounded in InstructGPT, Bradley-Terry, DPO, RLAIF and open-source RLHF-infra patterns.

    Pro Start →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.