Match a job Paths Subjects Questions Quizzes Pricing
AI Engineering Advanced Pro

Local LLM Deployment and Open-Weight Serving

Ollama, llama.cpp, vLLM and LM Studio; quantization tradeoffs from Q4 to FP8; speculative decoding; the VRAM math that decides whether a model fits; and when running locally actually beats an API call

35 min read 8 views

A hands-on-adjacent, numbers-first tour of the local/open-weight serving stack for AI engineering interviews: what Ollama, llama.cpp/GGUF, vLLM, and LM Studio each actually are and when you'd reach for each; OpenAI-compatible local endpoints as the integration pattern that makes a local model swappable behind existing client code; quantization (GGUF Q4/Q8, AWQ/GPTQ, bitsandbytes, and FP8 on Hopper-class hardware) and its quality-vs-memory tradeoff; the full VRAM math for weights plus KV cache, worked for a 32B model at Q4 on a 24GB consumer GPU; continuous batching and PagedAttention; speculative decoding and how it compares to quantization and distillation as three different levers for cheaper, faster inference; and the decision framework for when self-hosting actually beats an API call.

Practice questions (5)

  • Ollama, llama.cpp, or vLLM for This Workload?

    Advanced · Free
    View →
  • When Q4 Quietly Breaks a Task

    Advanced
    View →
  • Sizing a 13B Model at Q8 on a 16GB GPU

    Advanced
    View →
  • Continuous Batching or PagedAttention — Which Fixes This Symptom?

    Advanced
    View →
  • Local Deployment or Just Call the API?

    Advanced
    View →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.