Local LLM Deployment and Open-Weight Serving
Ollama, llama.cpp, vLLM and LM Studio; quantization tradeoffs from Q4 to FP8; speculative decoding; the VRAM math that decides whether a model fits; and when running locally actually beats an API call
A hands-on-adjacent, numbers-first tour of the local/open-weight serving stack for AI engineering interviews: what Ollama, llama.cpp/GGUF, vLLM, and LM Studio each actually are and when you'd reach for each; OpenAI-compatible local endpoints as the integration pattern that makes a local model swappable behind existing client code; quantization (GGUF Q4/Q8, AWQ/GPTQ, bitsandbytes, and FP8 on Hopper-class hardware) and its quality-vs-memory tradeoff; the full VRAM math for weights plus KV cache, worked for a 32B model at Q4 on a 24GB consumer GPU; continuous batching and PagedAttention; speculative decoding and how it compares to quantization and distillation as three different levers for cheaper, faster inference; and the decision framework for when self-hosting actually beats an API call.
Practice questions (5)
-
View →
Ollama, llama.cpp, or vLLM for This Workload?
Advanced · Free -
View →
When Q4 Quietly Breaks a Task
Advanced -
View →
Sizing a 13B Model at Q8 on a 16GB GPU
Advanced -
View →
Continuous Batching or PagedAttention — Which Fixes This Symptom?
Advanced -
View →
Local Deployment or Just Call the API?
Advanced