Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inherent in the 3D physical world. This deficiency restricts their spatial awareness and adaptability for complex, high-precision manipulation. To bridge this crucial gap, we...
GRACE is pro-posed, an agent framework that adopts Executable Analytic Concepts (EAC) as a core knowledge annotation paradigm for object understanding and supports automated concept discovery, offering a scalable and modular solution for high-precision embodied intelligence.
Mingyang Sun, Jiu-De Wei, Qi-Chen He et al.· Proceedings of the Thirty-Fi...· 0 citations
This work proposes an analytic concept-centric memory framework for agentic embodied manipulation that improves task completion, retrieval accuracy, object re-identification, and cross-object skill generalization over unstructured and embedding-based memory baselines.