Practice — Context Engineering Fundamentals (5 questions)
A Support Bot's Context Budget Blows Up on a Single Customer Permalink →
Your support bot has a 12,000-token practical per-turn budget, allocated roughly as: 1,200 system prompt + tool schemas (fixed), 4,000 retrieved policy content (capped, variable), 5,000 conversation history (capped, variable), 1,000 reserved for output, 800 margin. It has worked fine in testing.
In production, a customer with 300+ historical orders asks "what's the
status of all my recent orders?" The list_orders tool returns all 300
orders as raw JSON, which alone is around 18,000 tokens — blowing past
the entire budget before the model even sees the retrieved policy
content or gets to generate an answer. The request fails with a
context-length error.
- Diagnose which region of the context budget was actually unprotected, and explain mechanistically why this wasn't caught by the budget table above.
- Propose a fix for the
list_orderstool integration specifically — not a bigger budget. - Separately, is raising the practical budget from 12,000 to, say, 40,000 tokens a reasonable partial mitigation here? Justify with the mechanisms from context engineering, not just "more room is safer."
Share this question