Practice — Cost and Latency Engineering for LLM Apps (6 questions)
Prompt-Prefix Caching in an Agent Loop Permalink →
You are building a code-review agent that loops for an average of 25 steps per review (reading files, running linters, drafting comments). Each step sends 18,000 input tokens, of which 6,000 tokens are a stable prefix — system prompt, tool schemas and review guidelines that are byte-identical on every step — and produces 150 output tokens. Assume input is $2 per 1M tokens, output is $10 per 1M tokens, and a cached prefix is billed at 10% of the standard input price.
- Compute the cost of one review without any caching.
- Assume the prefix is a cache hit on steps 2 through 25 (24 of the 25 steps). Compute the cost of one review with prefix caching.
- At 50,000 reviews/day, what is the daily dollar saving from turning caching on?
- Why must the 6,000-token prefix be placed at the very start of the prompt, and be byte-identical every step, for any of this to work?
Share this question
The Quadratic Bill Hiding in a 50-Turn Conversation Permalink →
Your chatbot resends the full conversation history on every turn. Each turn adds ~500 tokens (user message + assistant reply).
Across a 50-turn conversation, how many input tokens does the API actually process in total?
Show the arithmetic, and name the mechanism that makes this affordable in practice.
Share this question