Match a job Paths Subjects Questions Quizzes Pricing

Local LLM Deployment and Open-Weight Serving

Ollama, llama.cpp, vLLM and LM Studio; quantization tradeoffs from Q4 to FP8; speculative decoding; the VRAM math that decides whether a model fits; and when running locally actually beats an API call

Overview Read

Local LLM Deployment and Open-Weight Serving

There is a specific moment in a lot of AI engineering careers where "just call the API" stops being the whole answer: a customer needs the model inside their own network, a team is burning real money on a high-volume steady workload, or someone just wants to know whether a 32B model actually fits on the GPU sitting under their desk before promising a stakeholder it will. how-llms-are-built-pretraining-to-chatbot makes the case for open-weight models as a category when closed APIs are structurally ruled out; this subject is what you actually do once you've landed there — the stack you serve those weights with, the arithmetic that tells you whether they fit, and the honest accounting of when self-hosting is worth the operational burden it adds.

Interviewers probe this territory because it separates candidates who have only ever called openai.chat.completions.create from candidates who understand that "the model" is inseparable from the hardware and serving stack running it — the same weights can be usable or unusable depending entirely on quantization choices and context-length settings that never come up when someone else is hosting the model for you. This subject assumes the KV-cache memory formula from transformers-for-ai-engineers and reuses it directly rather than re-deriving it — go there first if that formula isn't already solid.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.