Advanced
Open
Pro
Cost and Latency Model for the Support Assistant
Your design handles 8,000,000 LLM turns per day. A large-model turn uses ~4,700 input tokens (of which ~1,400 are a fixed prefix: system prompt, few-shot examples and tool schemas) and ~250 output tokens. Assume illustrative prices of $0.50 per 1M input tokens and $1.50 per 1M output tokens, a small model at one tenth of those prices, time-to-first-token of 400 ms, decode of 60 tokens/s, and retrieval + re-ranking + input guardrails of 200 ms.
- Compute the daily cost and the first-token and full-answer latency of one turn.
- Finance asks for at least a 60% reduction. Propose levers in the order you would pull them and estimate the effect of each.
- The product owner separately wants first-token latency under 1 s at p95. Which of your cost levers help, hurt, or are neutral for that goal?
Share this question