Paths Subjects Questions Quizzes Pricing Search
Intermediate Open Pro

Pushback: 'Just Use the Big Context Window'

Your engineering lead proposes simplifying the architecture: "Our knowledge base is only 300,000 tokens total if we concatenate every article. The model we're using supports a 1-million-token context window. Let's drop the whole RAG pipeline — retrieval, chunking, re-ranking, the vector index — and just paste the entire knowledge base into every request. It's simpler to build and maintain, and the model can technically fit it."

Evaluate this proposal. Where is it right, where is it wrong, and what would you actually recommend? Be concrete — don't just say "RAG is better," explain the specific mechanisms.

Share this question

← Back to Tokenization and Context Windows practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.