This work presents JITOMA (Just-In-Time On-demand Memory Activation), a closed-loop framework that unifies task reasoning, perception, and memory into a just-in-time growth process, and introduces JITOMA-Bench, a comprehensive suite for long-horizon multi-tasking and complex multi-step reasoning.
Abstract
While 3D Scene Graphs (3DSGs) provide crucial structured representations for embodied agents, conventional Ahead-of-Time, build-everything-then-filter pipelines conflict with the real-time, low-latency demands of edge platforms, inducing a perceptual saturation effect via severe observation redundancy. To resolve this, we present JITOMA (Just-In-Time On-demand Memory Activation), a closed-loop framework that unifies task reasoning, perception, and memory into a just-in-time growth process. Instead of exhaustively mapping the entire environment, JITOMA leverages a top-down task heatmap at the frontend to filter continuous observations, routing minimal streams to maintain a global foundation of low-cost, dormant anchors. Upon a cognitive query, the backend Large Language Model (LLM) parses the robotic intent to dynamically awaken task-relevant anchors, triggering resource-intensive operations -- such as dense node captioning and functional inference -- exclusively within the activated local subgraph. To evaluate these dynamic capabilities and study perceptual saturation trade-offs, we introduce JITOMA-Bench, a comprehensive suite for long-horizon multi-tasking and complex multi-step reasoning. Extensive experiments demonstrate that JITOMA substantially reduces active graph size and captioning latency, while maintaining stable processing time under long-horizon task switching.
Robots are now expected to execute increasingly complex long-horizon tasks in unstructured environments. Despite the strong potential of pretrained Vision-Language Models (VLMs) in task planning, their direct application to robotic manipulation is hindered by logical reasoning deviations and inadequate geometric scene...
Guang-Hui Ma, Jia-Hui Guo, Xin-Hua Tang et al.· Italian National Conference...· 0 citations
A modular mapping architecture is demonstrated that establishes 3D Semantic Scene Graphs (3DSSGs) as its foundational back-end, enabling the dense representation of extensive environments containing thousands of unique object instances and supporting open-vocabulary queries via CLIP features without requiring any addit...
Felix Igelbrink, Lennart Niecksch, Martin Günther et al.· Proceedings of the Thirty-Fi...· 0 citations
Prior-SG achieves state-of-the-art semantic region segmentation accuracy compared to recent baselines, robustly delineates distant functional boundaries in the absence of physical walls, and uniquely provides zero-shot ontological flexibility, enabling the robot to entirely restructure its spatial partitioning based on...
G. Tonetti, Laurent Kneip, Abel Gawel et al.· 0 citations
Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and executing continuous low-level actions. Existing methods often depend on LiDAR, panoramic cameras, or extra sensors; separate geometric-mapping and semantic-navigation visual...
Jian-He Zhao, Yan-Hua Qiu, Zhi-Yu Zhang et al.· 0 citations
HAM-VLN is presented, a decision-coupled, agent-authored memory that equips the robot with a persistent, depth-grounded world graph and reduces the context length by more than 65% compared to previous methods.
An Liu, Bingxi Liu, Hongyu Ding et al.· arXiv.org· 1 citation
Planning in complex environments requires task specifications grounded in representations that capture objects, relations, and affordances; scene graphs meet this need, but their size in large environments hinders efficient planning. While task-aware pruning and hierarchical abstractions have been explored, a general,...
Basak Sakçak, Francesco Verdoja· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.