Generalized Position-Based Model: Rethinking Position Weights in Ranking Off-Policy Evaluation
Abstract
Off-policy evaluation (OPE) estimates the performance of new recommendation policies using logged data, thus enabling fast, safe and inexpensive iteration prior to costly A/B tests. To evaluate ranking policies, existing OPE estimators all make structural assumptions about user behavior, leading to a spectrum of trade-offs between bias and variance. The recently proposed INTERPOL estimator navigates these trade-offs through a window system that defines how clicks at different positions are combined. However, this approach has two key limitations: the window configuration must be specified a priori, which can impede practical use, and all positions within a window are weighted uniformly, regardless of their relative utility, potentially limiting accuracy. To address these gaps, we introduce Generalized PBM (GPBM), an estimator that learns position-specific weights by minimizing an approximate upper bound on the estimation error. GPBM retains the unbiasedness guarantees of INTERPOL while significantly simplifying its hyperparameter selection. Our experiments demonstrate that GPBM provides more accurate and robust estimates across a wide range of position bias misestimation levels, logging policies, and data sizes.