Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Redesign a Click-Dominated Value Formula for Long-Term Value

A feed's current ranking score is score = 1.0*p(click) + 0.8*p(like) - 1.5*p(report). Leadership approves adding a long-term-value term driven by a 7-day session-return surrogate model, and accepts some near-term CTR cost.

  1. Two engineers propose w5 = 0.05 and w5 = 1.0 respectively for the new term. Using the item-A/item-B style comparison from this case study (one high-click/low-surrogate item, one lower-click/ high-surrogate item), explain concretely why one of these choices would ship with no real behavior change.
  2. What would you check, before launch, to confirm the chosen weight is "large enough to matter" without already needing the full long-horizon A/B result?
  3. What guardrail would you put in place to catch the value formula accidentally collapsing near-term engagement to zero?
Solution

1. Why the token weight changes nothing

Take item A: p(click) = 0.09, surrogate v = 0.20. Item B: p(click) = 0.15, surrogate v = 0.04. With w5 = 1.0: A scores 0.09 + 0.20 = 0.29, B scores 0.15 + 0.04 = 0.19 — A wins, ranking behavior actually changes. With w5 = 0.05: A scores 0.09 + 0.05(0.20) = 0.10, B scores 0.15 + 0.05(0.04) ≈ 0.152 — B still wins, exactly as it did before the redesign. A weight has to be large enough, relative to the existing terms' typical magnitudes, to actually flip rankings on real candidate pairs; otherwise "we added long-term value to the formula" is true on paper and false in production behavior.

2. Pre-launch checks without waiting for the full A/B

Run the new formula offline against a recent slate of real candidate pairs and measure how often the top-ranked item actually changes compared to the old formula — a near-zero reranking rate at the proposed weight is a direct signal the weight is still too small, independent of any online result. Also check the distribution of score contributions from each term across many requests (not just one hand-picked pair) to confirm the surrogate term is regularly comparable in magnitude to the click term, not swamped by it.

3. Guardrail against collapsing near-term engagement

Track near-term engagement metrics (session length, immediate CTR) as explicit guardrails during the rollout, not just the primary long-term metric — a formula that trades away all near-term engagement for long-term value is itself a failure mode (a feed nobody opens today can't accumulate the impressions the surrogate needs to keep learning). Set a guardrail threshold ("near-term engagement must not fall more than X% relative") and page/pause the rollout if it's breached, the same discipline used for any other guardrail metric in a staged release.

Share this question

← Back to Case Study: Long-Term Engagement Recommender practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.