Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Sizing a 13B Model at Q8 on a 16GB GPU

You need to size a 13B-parameter dense model at Q8 quantization for a 16GB GPU. Architecture: 40 layers, 40 query heads (head dim 128), and grouped-query attention with 8 KV heads. Target: a 16,000-token context window, batch size 1, fp16 KV cache.

  1. Compute the weight memory at Q8.
  2. Compute the KV cache memory at 16k tokens, batch size 1, using the formula from transformers-for-ai-engineers.
  3. Does this fit on a 16GB card, accounting for ~1-2GB of serving overhead? If it's close or doesn't fit, name the two cheapest levers to make it fit without changing the model.

Share this question

← Back to Local LLM Deployment and Open-Weight Serving practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.