Practice — RAG Architecture End to End (10 questions)
Diagnosing 'Plausible but Wrong' Answers Permalink →
Your team shipped a RAG assistant over a 20,000-article product knowledge base six weeks ago: fixed-size 512-token chunks, a single dense vector index, top-5 retrieval with no re-ranking, and no confidence gate. Support tickets show a pattern: the assistant answers fluently and confidently, but roughly 1 in 4 answers state a detail (a price, a limit, a supported feature) that is subtly wrong — not nonsense, just incorrect, and always delivered with total confidence.
- Walk through the pipeline stage by stage and name the most likely root cause(s) of this specific failure pattern (confident + subtly wrong, not garbled or off-topic).
- Propose the ordered set of changes you would make, and explain why that order.
- What would you add to make this class of failure visible in production instead of only visible in support tickets?
Share this question
What a Cross-Encoder Re-Ranker Actually Fixes Permalink →
Your RAG pipeline merges dense and BM25 results with reciprocal rank fusion (RRF) into a candidate pool of 30, then runs a cross-encoder re-ranking pass to pick the final top-5 for the prompt. A teammate asks: "RRF already produces a relevance-ordered list — why add a whole extra re-ranking pass on top of it?" What is the single most accurate answer?
A. A cross-encoder scores each (query, chunk) pair jointly in one forward pass, which is a more precise relevance signal than RRF's fusion of two independently-computed rankings. B. Cross-encoder re-ranking is computationally cheaper than RRF, so it's used to cut latency rather than to improve precision. C. Cross-encoder re-ranking retrieves from a larger candidate pool than the dense and BM25 passes did individually. D. Cross-encoder re-ranking removes the need for metadata/ACL filtering on the retrieved candidates.
Share this question
The One Trace Field That Catches Embedding-Model Mismatches
Unlock this question →The RAM Bill for 10 Million Embeddings Permalink →
Your RAG corpus has 10 million chunks, each embedded as a 1536-dimension float32 vector.
Roughly how much memory do the raw vectors need — before any index overhead?
Show the bytes-per-vector arithmetic, and name two techniques that bring the footprint down.
Share this question