Practice — Ranking & Recommendation System Architecture (5 questions)
Intermediate
Open
Free
Sizing the Funnel from a Latency Budget Permalink →
A video platform has a catalogue of 50 million videos. The ranking model costs 8 μs per (user, item) pair when batched, and the overall request budget (feature fetch + retrieval + ranking + re-ranking) is 120 ms p99.
- Show why scoring the full catalogue with the ranker is infeasible, with numbers.
- If feature fetch takes 15 ms and re-ranking takes 5 ms, how many candidates can the ranker afford to score within budget, and what does that number tell you about the candidate-generation stage's job?
- Candidate generation returns candidates from 4 sources (two-tower ANN, item-to-item, trending, subscriptions) with some overlap. Why is that better than relying on the two-tower source alone?
Share this question
Advanced
Open
Pro