Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Free

The Quadratic Bill Hiding in a 50-Turn Conversation

Your chatbot resends the full conversation history on every turn. Each turn adds ~500 tokens (user message + assistant reply).

Across a 50-turn conversation, how many input tokens does the API actually process in total?

Show the arithmetic, and name the mechanism that makes this affordable in practice.

Solution

C) ~640K tokens — 25× the naive estimate.

The naive estimate is 50 turns × 500 tokens = 25,000 tokens. But turn k resends everything before it, so its input is ~500 × k tokens. The total is:

500 × (1 + 2 + … + 50) = 500 × (50 × 51 / 2) = 500 × 1,275 = 637,500 tokens

Input grows quadratically with conversation length — the last turn alone processes 25,000 input tokens, as much as the entire naive estimate for the whole conversation.

The rule of thumb to say out loud: an n-turn conversation processes ~n²/2 turn-sizes of input, not n — long chats are quadratic, not linear.

What makes it affordable:

  • Prompt caching — the shared prefix (everything before the newest message) is a cache hit on every turn, typically billed at ~10% of the normal input price. The quadratic token count remains, but the quadratic cost mostly disappears.
  • Beyond caching: summarize or truncate old turns, or cap history at a sliding window once the conversation stops needing early context.

Share this question

← Back to Cost and Latency Engineering for LLM Apps practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.