Skip to content

Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents

Sep 2026 · 1 citation · 36 references
Computer Science

TL;DR

SPACE induces two-level programmatic skills from successful trajectories, where subskill boundaries serve as direct chunk-boundary supervision from trajectory-induced programmatic skills, and distilled into a primitive-chunk policy via hybrid on-/off-policy optimization.

Abstract

Large language model (LLM) agents for long-horizon interactive tasks typically follow a ReAct-style protocol, issuing one primitive action per LLM round. While this enables frequent replanning, it is inefficient for long-horizon tasks where many rounds are spent on routine action sequences. A natural alternative is to let the agent emit variable-length action chunks. However, naively training such policies with standard reinforcement learning fails: the agent either collapses to single-action behavior or over-commits to excessively long sequences. Both failures share a common root cause: the inability to learn chunk boundaries. We propose SPACE, which addresses this challenge by distilling chunk-boundary supervision from trajectory-induced programmatic skills. We induce two-level programmatic skills from successful trajectories, where subskill boundaries serve as direct chunk-boundary supervision. This temporal structure is then distilled into a primitive-chunk policy via hybrid on-/off-policy optimization with chunk-aware credit assignment. Experiments on ALFWorld and ScienceWorld show that SPACE improves success rates by 7.0%-31.3% over the strongest baseline in each setting while reducing average LLM decision rounds by up to 78.9%.

View source

Similar papers

#machine learning Preprint Sep 2026

Action Chunking Proximal Policy Optimization with Feedback Correction

Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as act...

Sang J. Hahn, Jonghyun Choi · 0 citations
#artificial intelligence Preprint Sep 2026

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task success. Sequence-level re...

Youngjun Jun, Kyumin Choi, Young Min Kim et al. · 0 citations
#machine learning Preprint Sep 2026

FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world mo...

Shi-Du Ren, Qi Gu, Zheng-Hao Ni et al. · 0 citations
#artificial intelligence Preprint Sep 2026

APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering

Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either retrieve verbose task-specific traces that burden decision-making, or distill pr...

Jie Ding, Rui Sun, Xin-Yi Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks

World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over...

Guoheng Sun, Chen Chen, Jin Wang et al. · 0 citations
Preprint Aug 2026

Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses

Harness-RL is introduced, a structured reinforcement learning framework that combines Conflict-Aware Policy Optimization (CAPO) with interface-level black-box trajectory construction and supports both central-only and joint multi-agent training.

Xin-Ke Jiang, Zhi-Xin Zhang, Zhi-Bang Yang et al. · 1 citation

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.