GRB: A Generative Reinforcement Bidding Framework for Multi-Channel Online Advertising
Abstract
Auto-bidding has become a central component of modern advertising platforms. Recently, generative paradigms based on Decision Transformers (DT) have emerged as a promising alternative, modeling auto-bidding as sequence generation and using return-to-go (RTG) as a signal, thereby enabling long-horizon credit assignment and direct exploitation of long-context trajectory dependencies. However, existing generative bidding methods are confined to single-channel settings, and extending them to multi-channel settings is non-trivial due to the high-dimensional, coupled bid vector and cross-slot dependencies. Moreover, purely offline training on fixed logs is prone to impression bias and behavioral degeneration during rollout, whereas Reinforcement Learning is severely constrained in exploration by real-world revenue. To address these challenges, we introduce GRB (Generative Reinforcement Bidding), a large-scale generative bidding framework with an Offline-to-On-Policy (O2P) training paradigm. GRB learns a multi-channel DT policy using a score-based RTG that transforms long-horizon CPA objectives into a bounded signal suitable for sequence modeling through offline supervised learning. It then performs efficient RL-free on-policy refinement: the current stochastic policy explores against a generative evaluator, trajectories are selected via a baseline-gated criterion within an ad-wise replay buffer, and the policy is updated through a supervised Refinement Negative Log-Likelihood (RNLL) objective. To mitigate reward hacking, we further incorporate an adaptive constrained exploration module that enforces a KL trust region around the logging policy and a Lipschitz stability constraint, thereby providing provable safety and stability guarantees. Experiments on a large-scale Tencent Ads dataset spanning five channels, together with online A/B tests, demonstrate that GRB consistently outperforms strong online and offline baselines.