Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Free

The RAM Bill for 10 Million Embeddings

Your RAG corpus has 10 million chunks, each embedded as a 1536-dimension float32 vector.

Roughly how much memory do the raw vectors need — before any index overhead?

Show the bytes-per-vector arithmetic, and name two techniques that bring the footprint down.

Solution

C) ~61 GB — and that's before the index structure is built on top.

Per vector: 1,536 dimensions × 4 bytes (float32) = 6,144 bytes ≈ 6 KB.

10,000,000 × 6,144 bytes ≈ 61.4 GB of raw vectors. An HNSW graph index adds roughly another 50% on top for its link structure, pushing a naive in-memory deployment toward ~90 GB — beyond most single nodes you'd casually provision.

The rule of thumb to say out loud: vector RAM ≈ count × dimensions × 4 bytes, plus ~50% index overhead — 1M full-size vectors ≈ 6 GB.

Two ways to bring it down:

  • Quantization — int8 cuts 4 bytes/dim to 1 (4× smaller, ~1% recall loss); binary quantization goes to 1 bit/dim (32× smaller) with a rescoring pass to recover accuracy.
  • Smaller dimensions — Matryoshka-style embeddings truncated to 256–512 dims cut memory 3–6× with modest quality loss.
  • Also worth naming: product quantization (PQ), and disk-based indexes (DiskANN-style) that keep only a compressed copy in RAM.

Share this question

← Back to RAG Architecture End to End practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.