The Quadratic Bill Hiding in a 50-Turn Conversation
Your chatbot resends the full conversation history on every turn. Each turn adds ~500 tokens (user message + assistant reply).
Across a 50-turn conversation, how many input tokens does the API actually process in total?
Show the arithmetic, and name the mechanism that makes this affordable in practice.
C) ~640K tokens — 25× the naive estimate.
The naive estimate is 50 turns × 500 tokens = 25,000 tokens. But turn k resends everything before it, so its input is ~500 × k tokens. The total is:
500 × (1 + 2 + … + 50) = 500 × (50 × 51 / 2) = 500 × 1,275 = 637,500 tokens
Input grows quadratically with conversation length — the last turn alone processes 25,000 input tokens, as much as the entire naive estimate for the whole conversation.
The rule of thumb to say out loud: an n-turn conversation processes ~n²/2 turn-sizes of input, not n — long chats are quadratic, not linear.
What makes it affordable:
- Prompt caching — the shared prefix (everything before the newest message) is a cache hit on every turn, typically billed at ~10% of the normal input price. The quadratic token count remains, but the quadratic cost mostly disappears.
- Beyond caching: summarize or truncate old turns, or cap history at a sliding window once the conversation stops needing early context.
Share this question