Skip to content
Conference Open access

Beyond Scaling: A Survey of Data-Efficient Learning for LLM Agents

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · pp. 8050-8059 · 1 citation · 81 references

TL;DR

An agent-centric view of data-efficient learning is developed, an agentic learning loop is formulated, data efficiency is defined as capability improvement without proportional increases in supervision or real-environment trial-and-error, and methods are synthesized into experience augmentation, agent structural design, and learning paradigms.

Abstract

LLM-based agents perceive, reason, act, and adapt through interaction. While scaling remains important, agentic progress also depends on extracting more learning signal from limited experience, including demonstrations, feedback, reasoning traces, tool-use records, and interaction trajectories. This survey develops an agent-centric view of data-efficient learning. We formulate an agentic learning loop, define data efficiency as capability improvement without proportional increases in supervision or real-environment trial-and-error, and synthesize methods into experience augmentation, agent structural design, and learning paradigms. We also summarize representative application domains and open challenges in experience reuse and cost-aware evaluation. Overall, the agentic era requires smarter ways to acquire, structure, reuse, and learn from limited experience.

Read PDF

Similar papers

#machine learning Preprint Aug 2026

Learning Generalizable Behaviors for Terminal Agents

River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.

Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al. · 2 citations
Preprint Aug 2026

State2State: Environment-Derived Mid-Training for LLM Agents

State2State is proposed, an environment-derived mid-training method that converts explored environment states into training objectives, challenging agents to reach a specified target state by deriving tasks from environment exploration and verifying success through rule-based state matching.

Xuanyu Lei, Yi-Qi Zhu, Chen-Liang Li et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents

BQ-LoRA is proposed, a low-rank adaptation framework that organizes trajectory updates through a local behavior quotient manifold and contains two modules, i.e., behavior quotient balancing (BQB) and decision preserving compression (DPC).

Peng-Yang Zhou, Xiao-Bing Tu, Zheng-Xi Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orc...

Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé · 0 citations
#artificial intelligence Preprint Sep 2026

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee com...

Bo Mao, Hang He, Lin-Ting Wang et al. · 0 citations
Book Open access Aug 2026

Large Language Model (LLM) as an Excellent Reinforcement Learning Researcher in both Single-Agent and Multi-Agent Scenarios

A Self-Evolutional single-agent/multi-agent Reinforcement Learning (SE-RL) framework that utilizes a Large Language Model (LLM) to design various RL algorithm modules, such as agent model design, reward function, profiling, communication, and state imagination, by leveraging the LLM generating module output or code.

Vincent Fu, Xin-Xin Xu, Weichen Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.