Advanced
Open
Pro
Cost Routing and Evaluation Design
Your document assistant is live and usage has grown to 200,000 queries/day, a mix of simple lookups and deep synthesis/comparison questions. Finance flags that average cost per query is higher than projected. Separately, the eval team reports "recall@10 on our golden set is 94%, so retrieval quality is fine" — but user complaints about wrong answers haven't gone down.
- Using illustrative prices of $3/1M input tokens, $15/1M output tokens, and cached input at 10% of the input price, explain why a blended average cost figure is misleading here and what you'd report instead.
- Propose the routing strategy that reduces cost without touching quality, with rough numbers.
- Explain why "recall@10 is 94%" does not mean retrieval quality is fine for this system, and design the evaluation that would actually explain the complaints.
Share this question