Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· 0 citations· 31 references
TL;DR
LOGIC transforms the exploration paradigm: instead of exploring at the action level, it employs an integrated critic to guide a generator in discovering optimal, high-value goals (RTGs), enabling robust exploration beyond the dataset's frontier while avoiding the fragility of complex multi-loss objectives.
Abstract
Decision Transformers (DTs) have emerged as a powerful paradigm for sequential decision-making in offline learning, yet they face two intrinsic deficiencies: the challenge of specifying optimal Return-to-Go (RTG) targets and a fundamental lack of exploration beyond the offline dataset. While prior studies have attempted to address these issues separately, their solutions introduce new problems. In this paper, we introduce LOGIC (Learning Optimal Goals with an Integrated Critic), a holistic framework that resolves these dual challenges in a unified manner. LOGIC transforms the exploration paradigm: instead of exploring at the action level, it employs an integrated critic to guide a generator in discovering optimal, high-value goals (RTGs). This strategic shift enables robust exploration beyond the dataset's frontier while avoiding the fragility of complex multi-loss objectives. Simultaneously, a Backward Consistency Refinement (BCR) module, applied during both training and inference, ensures that all generated plans remain mathematically and semantically valid. Extensive experiments on diverse auto-bidding scenes demonstrate that LOGIC significantly outperforms state-of-the-art baselines, achieving superior cumulative returns and model robustness, thereby providing a principled and unified solution to goal-oriented offline learning.
A three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline is proposed, which provides superior scalability and stability, though often at the expense of representational depth.
Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al.· Discover Artificial Intellig...· 0 citations
Delta (Differential Testing for DRL Agents) is proposed, a novel and comprehensive framework that automatically identifies both safety-critical and optimality bugs in DRL agents and investigates the effectiveness of three offline RL algorithms in generating challenger agents.
Jun-Da He, Jie-Ke Shi, Zhou Yang et al.· Proceedings of the ACM on so...· 0 citations
Auto-bidding is central to computational advertising, where strategies must maximize advertisers'conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevail...
Ye-Wen Li, Peng Jiang, Yi-Tian Li et al.· 0 citations
Open-MOPD, a principled framework incorporating token-share balancing, gap-aware dynamic budget allocation, and student reward refresh, systematically restore cross-domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student.
Huan Gao, Haohan Chi, Yong Yan et al.· 12 citations· ⚡6
This work introduces Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups.
Jin Cui, Xin-Yue Long, Bo-Ran Zhao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.