Paths Subjects Questions Quizzes Pricing Search
Advanced Open Pro

The Silent Embedding-Model Mismatch

A team re-indexed their entire knowledge base overnight using a new, better embedding model, and updated the retrieval service to call that new model for query-time embeddings — but a canary of old application servers kept running for six hours during the rollout, still calling the old embedding model for queries against the new index.

  1. What happens to retrieval quality during those six hours, and why does it fail the way it does rather than erroring out?
  2. Why is this bug particularly dangerous compared to, say, a bad chunking decision?
  3. What would you change about the ingestion/deployment process so this class of bug becomes structurally hard to ship?

Share this question

← Back to RAG Architecture End to End practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.