Advanced
Open
Pro
Sizing a 13B Model at Q8 on a 16GB GPU
You need to size a 13B-parameter dense model at Q8 quantization for a 16GB GPU. Architecture: 40 layers, 40 query heads (head dim 128), and grouped-query attention with 8 KV heads. Target: a 16,000-token context window, batch size 1, fp16 KV cache.
- Compute the weight memory at Q8.
- Compute the KV cache memory at 16k tokens, batch size 1, using the
formula from
transformers-for-ai-engineers. - Does this fit on a 16GB card, accounting for ~1-2GB of serving overhead? If it's close or doesn't fit, name the two cheapest levers to make it fit without changing the model.
Share this question