Advanced
Open
Pro
Modelling Cost and Latency for an LLM Feature
You are designing an assistant that will handle 500,000 turns per day. Each turn sends a 4,000-token prompt (of which 1,200 tokens are a fixed system prompt and few-shot examples) and produces a 250-token answer. Assume illustrative prices of $1.00 per 1M input tokens and $4.00 per 1M output tokens, time-to-first-token of 500 ms, and a decode speed of 50 tokens per second.
- Compute the daily cost and the end-to-end latency of one turn without streaming.
- The product owner wants the response to "feel fast" and the finance team wants at least 40% cost reduction. Propose concrete levers, with estimated effect, for each requirement.
- Which single lever changes latency the most, and which changes cost the most? Explain why they differ.
Share this question