Generalization in multi-arm collaboration can be studied as composing familiar atomic skills in new ways across arms. However, existing evaluations offer limited insight into which training and architectural choices support this ability under different coordination requirements. We introduce \textbf{ACG-Bench}, a bench...
Zai-Bin Zhang, Bing-Hao Ran, Yu-Han Wu et al.· 0 citations
LatticeMind is presented, a conflict-aware structured memory that handles contradiction at write time, which maintains explicit item status, applies cheap symbolic conflict checks, and invokes LLM reconciliation only for unresolved semantic cases.
Heng Zhou, Lian Zhang, Yutao Fan et al.· 0 citations
Experiments across diverse models, benchmarks, and agent harnesses show that supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance, establishing verified data synthesis as an effective and scalable approach for skill-use training.
Zelin Tan, Yi-Qun Zhang, Hao Li et al.· 2 citations
MA-VLA decomposes cooperative behavior into mid-level atomic prompts and allocates them to individual arms, enabling explicit subgoal specification and compositional reuse across tasks and indicates that structured, per-arm atomic action assignment offers a practical route to scalable generalization in multi-arm embodi...
Zai-Bin Zhang, Jun-Lan Xiao, Zhong-Bo Zhang et al.· 4 citations
This work presents REAL, an agentic framework for open-world mobile manipulation, which establishes sim-to-real-consistent environment APIs without oracle perception and integrates a simulated user to enable human-in-the-loop interaction.
Boyu Mi, Mengchen Ma, Yifei Yao et al.· arXiv.org· 2 citations
SciOrch is presented, a framework that trains a lightweight 8B model to orchestrate frontier LLMs for scientific reasoning, and attains the best accuracy on both SGI and SFE with less than half the API cost of typical multi-agent methods.
Jingru Guo, Xiangyuan Xue, Lian Zhang et al.· arXiv.org· 0 citations
We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage normalization, to address two limitations of existing reward designs. Outcome reward models (ORM) evaluate only final-answer correctness, trea...
Zelin Tan, Zhouliang Yu, Bo-Cheng Lin et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.