Skip to content

Author

Long-Yuan Zhang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.

AgiBot Research Team, Renhang Liu, Wen-Zhi Zhao et al. · 0 citations
#small language model Open access Aug 2026

GSMultiAgent: A Multi-Agent Loop Framework atop Hermes Agent for Intelligent Design of Guidance Systems

The strong coupling among guidance laws, control loops, aerodynamics, and mission constraints poses growing challenges to the design of modern tactical missile guidance systems for autonomous flight. Conventional manual tuning and simulation-based trial-and-error result in long iteration cycles, limited reuse of design knowledge, and poor adaptability to changing scenarios. To address these limitations, we propose GSMultiAgent, a multi-agent collaborative cascade framework built atop Hermes Agent, which transforms natural-language mission requirements into optimized guidance system models through structured agent cooperation with feedback-driven iterative refinement. Three innovations are introduced: (1) a three-layer correction pipeline covering syntactic checking, deterministic mathematical verification, and semantic reasoning; (2) a bimodal experience repository supporting similarity-guided retrieval with access-count decay and best-quality retrieval for PPO warm-start initialization; and (3) a self-adjudicating optimizer that autonomously decides between PPO-based systematic parameter search and heuristic LLM-tuning guided by a reflection agent. Across four engagement scenarios, GSMultiAgent consistently attains high feasibility at a small fraction of the simulation budget required by conventional optimizers and single-agent baselines, and its design paths escalate autonomously from parameter tuning to structural law modification as task difficulty increases. Ablation studies confirm that the reflection agent, the optimization agent, and structured memory each contribute essential and complementary gains. These results establish multi-agent coordination with structured memory and self-adjudicating optimization as an effective paradigm for intelligent, reusable guidance system design.

Ji-Song Xiao, Chengwei Yang, Xiao Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.