Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods of...
Wan-Shun Su, Yang Shi, Fei Liu et al.· 2 citations
Long-horizon autonomous research tasks such as machine learning engineering require systems to make interdependent decisions under a limited budget. Existing LLM-based agents typically organize candidate-solution improvement through tree, graph, or chain structures, meaning that the search process determines how inform...
Shaokang Fu, Yulong Tao, Linbo Jin et al.· 1 citation
This work proposes PAO (Positive-Advantage-Only), a selective RL optimization method that selectively applies gradient updates only to retrieved items with positive advantages, effectively pulling query embed- dings toward high-reward regions while preserving global topo- logical stability.
Shao-Wei Wei, Chong Huang, Songtao Fang et al.· 0 citations
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumul...
Qi-Ming Shi, Yulong Tao, Linbo Jin et al.· arXiv.org· 2 citations
This work proposes BCSD (Bidirectional Context Self-Distillation), a framework that combines self-distillation with reinforcement learning to train LLM agents to use external skills more effectively, enabling agents to utilize external skills more effectively.
Tian Pan, Yuan Li, Hong-Da Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.