Advanced
Pro
Local LLM Deployment and Open-Weight Serving Quiz
Test your understanding of the local-serving stack (Ollama, llama.cpp/GGUF, vLLM, LM Studio), OpenAI-compatible local endpoints, quantization tradeoffs, VRAM sizing math for weights plus KV cache, continuous batching and PagedAttention, and when self-hosting actually beats an API call.
10 questions
10 min
Pass: 70%