Paths Subjects Questions Quizzes Pricing Search
Intermediate Open Pro

Compressing a Ranker for a Tighter Latency Budget

A cross-encoder re-ranker currently runs on GPU and takes 45 ms at p99 to score 100 candidate pairs, blowing the service's 100 ms total budget once feature fetch and retrieval are included. Offline accuracy (NDCG@10) is 0.71.

  1. Rank quantisation, distillation and pruning by how much latency improvement each is likely to deliver here, and explain why.
  2. Propose a concrete plan combining two of the three techniques, and state what you would re-check before shipping the compressed model.
  3. If the resulting model scores NDCG@10 = 0.685 offline, is that acceptable? What else would you want to know before deciding?

Share this question

← Back to Model Serving & Deployment practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.