The RAM Bill for 10 Million Embeddings
Your RAG corpus has 10 million chunks, each embedded as a 1536-dimension float32 vector.
Roughly how much memory do the raw vectors need — before any index overhead?
Show the bytes-per-vector arithmetic, and name two techniques that bring the footprint down.
C) ~61 GB — and that's before the index structure is built on top.
Per vector: 1,536 dimensions × 4 bytes (float32) = 6,144 bytes ≈ 6 KB.
10,000,000 × 6,144 bytes ≈ 61.4 GB of raw vectors. An HNSW graph index adds roughly another 50% on top for its link structure, pushing a naive in-memory deployment toward ~90 GB — beyond most single nodes you'd casually provision.
The rule of thumb to say out loud: vector RAM ≈ count × dimensions × 4 bytes, plus ~50% index overhead — 1M full-size vectors ≈ 6 GB.
Two ways to bring it down:
- Quantization — int8 cuts 4 bytes/dim to 1 (4× smaller, ~1% recall loss); binary quantization goes to 1 bit/dim (32× smaller) with a rescoring pass to recover accuracy.
- Smaller dimensions — Matryoshka-style embeddings truncated to 256–512 dims cut memory 3–6× with modest quality loss.
- Also worth naming: product quantization (PQ), and disk-based indexes (DiskANN-style) that keep only a compressed copy in RAM.
Share this question