Practice — Prompt Caching and Context Cost Optimization (5 questions)
The Cache Stopped Helping and Nobody Changed the Prompt Text Permalink →
Your team's customer-support assistant uses a system prompt with prompt
caching enabled: a ~2,000-token block of role instructions, policies,
and tool schemas, followed by the per-request retrieved context and the
user's message. Cost per request has been stable for months at roughly
$0.004/request (mostly cache hits on the 2,000-token block). Last week,
someone added a small feature: the system prompt now opens with
"Today's date is {current_date}. You are a support assistant..." so
the model can reason about date-relative questions ("is my order still
within the return window"). No other text changed. This week, average
cost per request has risen to roughly $0.011, and nobody connected it
to the date change because "it's the same prompt, just one more fact
in it."
- Explain exactly why this one addition destroyed the caching benefit, in terms of how prefix matching actually works — not just "the prompt changed."
- Propose a fix that keeps the date-awareness feature but restores the cache-hit rate. Be specific about where the date field should live.
- Roughly estimate the cost impact using the numbers given: what fraction of the original savings did this one change erase, and why does even "one more fact" have an outsized effect on prefix caching specifically, more than it would on a token-budget alone?
Share this question