A framework that combines decentralized asynchronous reasoning, lightweight information sharing, capability aware collaboration, and a unified action interface is proposed, enabling general purpose VLMs to generate robot specific actions executed by learning free experts without task or robot specific training.
Abstract
Multi-robot systems, particularly heterogeneous robot swarms, can improve the efficiency of complex task execution through parallel collaboration and complementary capabilities. However, conventional rule-based methods rely on predefined task models and specialized decision making programs, making it difficult to understand complex semantic instructions and coordinate heterogeneous robots. LLMs introduce strong language understanding and task reasoning capabilities, allowing multi-robot systems to interpret instructions, decompose tasks, and assign roles according to task semantics. VLMs further incorporate visual perception, enabling robots to reason about objects, regions, and spatial relationships in physical environments. Nevertheless, existing LLM/VLM based methods often depend on known maps, centralized and synchronized decision making, limiting their generalization to heterogeneous robots and unseen tasks. We therefore propose a framework that combines decentralized asynchronous reasoning, lightweight information sharing, capability aware collaboration, and a unified action interface, enabling general purpose VLMs to generate robot specific actions executed by learning free experts without task or robot specific training. Experiments across diverse scenarios and multiple VLMs show success rates above 70\%, with completion time reduced by up to 55.8\% relative to the geometric greedy baseline.
CoMuRoS enables runtime, event-driven replanning on physical robots and supports flexible multi-robot and human-robot collaboration across diverse scenarios.
Suraj S. Borate, Bhavish Rai B, Vipul Pardeshi et al.· Frontiers in Robotics and AI· 0 citations
Cloud-edge collaborative computing enables resource-constrained robots to leverage powerful cloudhosted models while retaining real-time on-device perception and control. Vision-Language Navigation in Continuous Environments (VLN-CE), which requires a robot to follow natural-language instructions through complex scenes, is a representative task that benefits from this paradigm: it relies on large Visual Language Models (VLMs) for multimodal reasoning yet demands responsive execution at the edge. However, existing VLM-based approaches remain constrained by limited context windows and insufficient planning capabilities for long-horizon tasks. We present EntityNav, an entity-centric stepwise planning framework for VLN-CE designed for cloudedge deployment. EntityNav comprises two integrated modules executed on the cloud: (1) Entity-Guided Stepwise Language Planning, which decomposes instructions into sequential, entity-centered sub-goals for explicit progress tracking, and (2) Entity-Aware Chain-of-Thought Reasoning, which generates a multi-stage structured reasoning chain whose hidden-state representations directly condition the action prediction head, regularized by a reasoning-action consistency loss. On the robot side, an edge-level module performs real-time visual capture and local trajectory refinement, with asynchronous communication overlapping cloud inference and physical motion to preserve responsiveness; an edge-side fallback mechanism further maintains safe navigation during transient cloud delays. Experiments on R2R-CE and RxR-CE benchmarks show that EntityNav achieves success rates of 62.7% and 60.3% respectively, demonstrating competitive performance against baselines. Real-world deployment on a quadruped robot further shows the framework's effectiveness under practical cloud-edge conditions.
Heng-Yi Yang, Yong Zhou, Shang Liu et al.· Fall Joint Computer Conferen...· 0 citations
Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.
Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks. In multi-agent settings, different agents inspect different image regions, video frames, or visual representations, so collaboration extends beyond distributed reasoning to distributed perception. This makes shared visual context a central problem in VLM agent collaboration. In this paper, we frame memory hierarchy, cross-agent sharing, and consistency mechanisms around the need to reconcile interpretations and update dependent reasoning. Effective collaboration requires agents to build on contributions from other agents, recover missing visual context, and reconcile differing interpretations as new evidence emerges. Shared visual memory preserves not only images or textual summaries but also the dependencies among observations, agent interpretations, and subsequent reasoning. Together, these design considerations shape how information flows and evolves across VLM agents. The proposed framework provides a foundation for building reliable and resource-efficient agent teams.
Hui-Xin Zhang, Shao-Jun Xia, Di Wang et al.· 0 citations
The World-Cognition Model is presented, a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime and introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks.
Autonomous property inspection requires more than robust robot navigation: a deployable system must connect heterogeneous sensing, reusable autonomy capabilities, multimodal scene understanding, human interaction, and enterprise response within a traceable operational loop. Existing quadruped inspection systems commonly integrate these functions through task-specific interfaces, making contextual coordination, knowledge reuse, and controlled adaptation difficult. This paper presents \textit{Harness Robotic OS} (HROS), a unified embodied-agent runtime, and Argos, its realization for residential-community inspection. HROS organizes the system into robot runtime, embodied autonomy skills, cognitive agent runtime, and interaction and operations planes. A shared context connects physical state with agent reasoning; streaming ASR/TTS supports voice-based mission interaction; hierarchical working, episodic, and semantic memory preserves operational knowledge; and a safety-gated self-evolution loop converts execution traces into versioned candidate updates without permitting unconstrained online modification. The Argos prototype integrates a Vbot quadruped, Fast-LIO2 localization and mapping, Hobot-Stereo depth perception, PCT-Planner global planning, EGO-Planner local motion generation, and OpenClaw-orchestrated Qwen3-VL inspection analysis. Experiments in a residential property environment achieved 100\% waypoint reachability, outdoor localization error below 10~cm, local obstacle-response latency below 200~ms, representative hazard-detection rates of 85--95\%, and 99\% success in alarm delivery and structured-report generation. These results validate the deployed navigation and inspection closed loop, while HROS provides an extensible software foundation for memory-augmented, voice-aware, and continuously improvable embodied inspection agents.
Yao-Yuan Yan, Zhi-You Heng, Haoxiang Jie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.