Advanced
Open
Pro
Cost and Latency Levers Under Growth
Your ask-the-web agent handles 8,000,000 queries/day at the numbers worked through in this subject: search API at $0.005/query, a synthesis call at roughly $0.0186 (4,200 input / 400 output tokens, $3 and $15 per 1M tokens respectively), and negligible cost for rewrite/fetch/rerank. Latency to first token is roughly 2.65s, with the parallel fetch step consuming the largest single share of that budget (~1.5s of the ~2.65s).
- Compute the current daily spend split between the search API and the synthesis call.
- Finance wants a 40% cost reduction without a latency regression. Propose the levers in the order you'd pull them, and be explicit about which levers help cost without touching latency, and which would help both.
- A teammate proposes fetching only 4 candidate pages instead of 8 to cut fetch latency. Evaluate this proposal against the other levers you named — is it a good trade, and what does it risk?
Share this question