Alignment + Accuracy: The Cascade Reward Representation for Preranking
Abstract
Prerankers in large-scale recommender systems select candidates for a downstream ranker under strict latency constraints. In practice, teams combine accuracy metrics with alignment losses to train and evaluate prerankers, but what these quantities should target—and how to combine them—remains ad hoc. We derive the Cascade Reward Representation: under mild assumptions on a fixed-retrieval, fixed-ranker pipeline, the expected change in user reward for preranker swaps with controlled overlap shift admits a calibrated first-order two-term representation \(\mathbb {E}[\mathbf {R}^E - \mathbf {R}^0] = \alpha \,\mathbb {E}[\Delta \hat{O}] + \beta \,\mathbb {E}[\Delta N] + \mathcal {R}\), where \(\Delta \hat{O}\) is a logged top-fraction overlap shift (alignment: agreement with the main ranker’s selections), ΔN is a threshold-conditioned precision shift (accuracy: engagement above a shared ranker threshold), and the remainder is bounded by \(O(\mathbb {E}[\Delta ^2]) + O_P(1/\sqrt {n})\). Both proxies are measurable on production-limited logged support without running the main ranker on the full retrieval pool. This representation motivates a calibrated offline metric and a matching two-branch training loss for the model family studied in our production system. We validate the representation in a large-scale industrial recommender system. A calibrated linear combination of the two metrics raises experiment winner prediction from 45–50% (accuracy-only) to 85% and Pearson r from ≤ 0.65 to 0.84 on held-out experiments. The matching training objective delivers +1.43% homefeed save rate over an accuracy-only baseline and +0.62% over a heuristic alignment+accuracy production model in two-week A/B tests.