Practice — Local LLM Deployment and Open-Weight Serving (5 questions)
Advanced
Open
Free
Ollama, llama.cpp, or vLLM for This Workload? Permalink →
A team is building an internal support-ticket triage feature that will serve roughly 200 concurrent employees during business hours, each sending occasional requests, on infrastructure the team fully controls. Someone proposes Ollama because "it's simple and we already use it on our laptops for testing."
- Is Ollama a reasonable choice for this production workload? Explain what specifically would or wouldn't scale.
- Which tool from this subject's stack is the better fit for this workload, and name the two specific techniques that make it a better fit.
- Would your answer change if this were instead a single internal tool used by one engineer at a time? Explain why or why not.
Share this question
Advanced
Open
Pro