Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Cost and Latency Levers Under Growth

Your ask-the-web agent handles 8,000,000 queries/day at the numbers worked through in this subject: search API at $0.005/query, a synthesis call at roughly $0.0186 (4,200 input / 400 output tokens, $3 and $15 per 1M tokens respectively), and negligible cost for rewrite/fetch/rerank. Latency to first token is roughly 2.65s, with the parallel fetch step consuming the largest single share of that budget (~1.5s of the ~2.65s).

  1. Compute the current daily spend split between the search API and the synthesis call.
  2. Finance wants a 40% cost reduction without a latency regression. Propose the levers in the order you'd pull them, and be explicit about which levers help cost without touching latency, and which would help both.
  3. A teammate proposes fetching only 4 candidate pages instead of 8 to cut fetch latency. Evaluate this proposal against the other levers you named — is it a good trade, and what does it risk?

Share this question

← Back to Case Study: Design an Ask-the-Web Agent practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.