Paths Subjects Questions Quizzes Pricing Search
Advanced Open Pro

Modelling Cost and Latency for an LLM Feature

You are designing an assistant that will handle 500,000 turns per day. Each turn sends a 4,000-token prompt (of which 1,200 tokens are a fixed system prompt and few-shot examples) and produces a 250-token answer. Assume illustrative prices of $1.00 per 1M input tokens and $4.00 per 1M output tokens, time-to-first-token of 500 ms, and a decode speed of 50 tokens per second.

  1. Compute the daily cost and the end-to-end latency of one turn without streaming.
  2. The product owner wants the response to "feel fast" and the finance team wants at least 40% cost reduction. Propose concrete levers, with estimated effect, for each requirement.
  3. Which single lever changes latency the most, and which changes cost the most? Explain why they differ.

Share this question

← Back to LLM Application System Design practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.