Skip to content
Book Open access

DPEO: Dynamic Preference Evolution Optimization for Self-Evolving CTR Prediction

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 5044-5048 · 0 citations · 30 references
Computer Science

TL;DR

This paper proposes DPEO (Dynamic Preference Evolution Optimization), a co-evolutionary framework that transforms CTR modeling into a dynamic policy contention task, enabling the sub-learners to serve as alternating evolutionary benchmarks and 'self-evolve' toward the global optimum.

Abstract

Click-through rate (CTR) prediction is a pivotal component in large-scale industrial systems. Historically, CTR prediction paradigms have been confined to monolithic architectures governed by a single-policy optimization process. However, such isolated learning paths lack the intrinsic evolutionary mechanisms necessary for optimal convergence. Without policy diversity and internal competition, models tend to get trapped in local optima as performance reaches saturation, hindering further breakthroughs in modeling capacity. In this paper, we propose DPEO (Dynamic Preference Evolution Optimization), a co-evolutionary framework that transforms CTR modeling into a dynamic policy contention task. DPEO decouples the monolithic architecture into dual sub-learners to induce policy diversity, constructing an internal preference landscape without external rewards. A performance-driven Role Arbiter then dynamically designates the superior sub-learner as the Reference Policy and the other sub-learner as the Target Policy per batch, driving continuous model evolution. Through an asymmetric gradient flow, the target policy is optimized to surpass the reference policy in both probability and logit spaces. This process drives a co-evolution, enabling the sub-learners to serve as alternating evolutionary benchmarks and 'self-evolve' toward the global optimum. Extensive experiments on public benchmarks and a massive industrial dataset with over 10 billion samples demonstrate that DPEO significantly outperforms state-of-the-art models.

Read PDF

Similar papers

Preprint Aug 2026

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

Environment-Regularized Policy Optimization (ERPO) replaces the standard Policy-KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.

Xianlei Zhou, Xiangdi Meng, Yu He et al. · 0 citations
Preprint Aug 2026

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

This work proposes Single-rollout Autoregressive Policy Optimization (SAPO), a low-memory and compute-efficient framework in which the policy and value functions share a single autoregressive backbone, and introduces a trajectory-level generalized advantage estimator that combines lambda-returns with batch normalization.

D. Liang, Lang Feng, Bo An et al. · 1 citation
Preprint Aug 2026

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

This work proposes \methodname, a stability-guided active-set controller for controlled objective admission, a stability-guided active-set controller for controlled objective admission in reward-vector RLHF, which positions objective-entry timing as a concrete control variable in reward-vector RLHF.

Yong-Qi Tong, Z. Zhang, Ruirui Wang et al. · 0 citations
Preprint Sep 2026

CompEvo: Competition-Induced Evolution for Multi-Agent in News-Driven Time Series Forecasting

News-driven time series forecasting uses evolving textual events together with historical observations to predict future values, supporting applications such as market risk monitoring and resource scheduling. In multi-agent settings, two challenges still remain. The first is degeneration of thought, where agents converge to similar evidence-seeking behaviors. The second is insufficient theoretical grounding, where strategy updates are often heuristic and lack a principled formulation. To address the above challenges, we propose CompEvo, a competition-induced evolution framework for multi-agent news-driven time series forecasting. For theoretical grounding, we introduce an evolutionary game formulation to guarantee equilibrium existence and optimization convergence. Building on this formulation, we construct a trainable multi-agent evolution framework that integrates strategy execution, fitness-based differentiable selection, and competition-induced strategy evolution. CompEvo enables heterogeneous agents to explore diverse news evidence, converts forecasting feedback into differentiable influence weights, and evolves agent strategies under competitive pressure to preserve effective logic while maintaining diversity. Experiments on four real-world datasets show that CompEvo reduces RMSE by 27.3% and MAPE by 26.2% on average over strong baselines. Further analysis indicates that CompEvo successfully maintains diverse and specialized agent behaviors.

Yuxuan Zhang, Yang-Yang Feng, Yong Guan et al. · 0 citations
Preprint Aug 2026

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent and uses imitation only where group-relative rewards are uninformative, obtains the best result in all four scientific domains.

Yubo Zhang, Xin-Hong Ma, Zezhong Tan et al. · 0 citations
Open access 2026

Proportional Reward and Temporal Discounting for Monopoly-Free Heterogeneous Metaheuristic Portfolios

Real-world optimization landscapes are typically dynamic, high-dimensional, and uncertain, and a single meta-heuristic with fixed control parameters rarely sustains strong performance across such environments, as formalized by the No Free Lunch theorem. Existing adaptive frameworks attempt to address this through online operator or algorithm selection, but they suffer from two persistent limitations: coarse feed-back mechanisms that reward the frequency rather than the magnitude of improvements, and cumulative memory bias that allows early-performing algorithms to monopolize selection long after their advantage has faded. This work proposes a problem-agnostic adaptive framework that integrates the Relative Improvement Metric (RIM), a proportional reward quantifying the magnitude of each solver’s contribution, with a Sliding Window (SW) policy that discounts older rewards temporally so that algorithmic influence remains contingent on recent effectiveness. The framework orchestrates a heterogeneous portfolio of eight metaheuristics (GA, PSO, GWO, ACO, SSA, ABC, WOA, FA) through probabilistic selection driven by SW-RIM weights, with a small base probability that prevents any solver from being permanently excluded. Empirical evaluation across 23 standard benchmarks and the CEC2020 suite shows that the proposed mechanism reduces maximum solver participation from above 70% in the baseline configuration to under 40%, improves mean fitness on the majority of functions in both groups, and achieves statistically significant gains over both a baseline portfolio and a sliding-window-only variant (Wilcoxon p < 0.001; Vargha-Delaney A12 between 0.72 and 0.85, large effect across all comparisons). The two exceptions are functions with deceptive or ill-conditioned landscapes (a narrow-valley Rosenbrock-type function and a highly multimodal Schwefel-type function), where all three configurations perform comparably, indicating that SW-RIM’s benefit is contingent on the portfolio containing at least one solver structurally suited to the current landscape rather than on the selection strategy alone. The results support SW-RIM as a lightweight, general-purpose mechanism for sustaining diversity and impact-sensitive adaptation in complex continuous optimization, without the training cost of reinforcement-learning-based selectors.

Bilal Bataineh, Sofian Kassaymeh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.