Skip to content
Preprint

AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning

Aug 2026 · 1 citation · 28 references
Computer Science

TL;DR

AReaL-DTE is presented, a snapshot-free Delta Transfer Engine that translates inference-visible weight sparsity into end-to-end system efficiency and achieves speedups of up to 19.9x over ByteCheckpoint and 3.2x over PULSE across clusters, and up to 7.6x and 7.4x within a cluster.

Abstract

Online agentic reinforcement learning implemented with micro-services separates policy training from rollout generation, improving scalability and modularity while potentially making frequent policy-weight synchronization a critical systems overhead. Shared storage naturally connects these services across clusters, but vanilla dense policy weight synchronization could incur model-scale construction, transfer, and application costs. Sparse synchronization reduces transferred data, yet checkpoint-oriented approaches can still retain a previous model and materialize complete intermediates to bridge heterogeneous training and inference layouts. We present AReaL-DTE, a snapshot-free Delta Transfer Engine that translates inference-visible weight sparsity into end-to-end system efficiency. Across our evaluated workloads, fewer than 2% of BF16 weight elements change between consecutive policy versions. AReaL-DTE reconstructs overwritten weights on demand by inverting AdamW updates, streams reconstructed and current parameters through converter-aligned BF16 change detection, and remaps changed elements directly into receiver-local coordinates. AReaL-DTE supports manifest-committed sparse transfer through shared storage across clusters and a deadlock-safe two-round protocol within a cluster, followed by direct application to inference shards. We evaluate AReaL-DTE on Qwen3-8B and Qwen3-30B-A3B across four online RL workloads. AReaL-DTE achieves speedups of up to 19.9x over ByteCheckpoint and 3.2x over PULSE across clusters, and up to 7.6x and 7.4x, respectively, within a cluster. In the same-cluster Qwen3-30B-A3B experiments, it reduces peak GPU memory by approximately 41% and peak CPU memory by at least 87%.

View source

Similar papers

#machine learning Preprint Aug 2026

HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees

HARTS is the first system to demonstrate arbitrary-rollout-tree prefix-sharing speedups on a real hybrid-attention model, and its numerical differences are comparable to baseline self-rerun variation, and its reward trend is similar to the baseline over the first 120 steps of SWE-bench training.

Bo-Yuan Meng, Pei-Hua Bao, Hong Liu et al. · 0 citations
#machine learning Preprint Jul 2026

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

Molt is a PyTorch-native training framework built to keep that cost small: a codebase compact and clean enough for a researcher to hold in their head, and for an AI coding assistant to read and reason about in its entirety, so the algorithm flow can be traced and changed end to end.

Jian Hu, Huiying Li, Hao Zhang et al. · 0 citations
Preprint Aug 2026

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward ha...

Yi-Ming Du, Yu-Xin Jiang, Tao Yuan et al. · 3 citations
Preprint Aug 2026

psRL: Efficient Training for Agentic AI via Training-Time Prefix Sharing

This paper proposes psRL (prefix sharing for RL), a new training system for agentic AI designed to exploit prefix redundancy among training samples, and introduces two novel prefix-sharing mechanisms that enable flexible, fine-grained workload distribution across GPU workers.

Mian-Jie Yu, Zi-Zhao Mo, Huanyu Qu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier by executing each task's own verifier.

Jun-Yao Yang, Yu-Cheng Shi, Zhong-Zhi Li et al. · 2 citations
Jul 2026

Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation

Adaptive FastOPD is proposed, a progress-aware strategy that expands the rollout horizon only when learning near the current boundary region has plateaued and the current horizon is sufficiently utilized, and remains robust across a range of hyperparameter settings.

Qian Tan, Huaifei Liang, Xuanyu Zhu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.