Aug 2026· 2026 IEEE International Conference on Mechatronics and Automation (ICMA)· pp. 1141-1146· 0 citations· 22 references
Abstract
This paper presents LVR-Draw, a fully local language–vision pipeline for robotic drawing that integrates structured scene generation, multimodal verification and correction, and deterministic execution within a unified Human–AI–Robot loop. Given a natural language prompt, a Large Language Model (LLM) generates a structured scene representation in a predefined format, which is rendered into an interpretable image. A Vision–Language Model (VLM) then performs visual inspection to detect inconsistencies in object placement and spatial relationships. These observations are processed by the LLM to produce structured editing operations, enabling iterative refinement of the scene. After validation, the refined scene is converted into executable robot instructions through a deterministic pipeline, supporting predictable and reproducible execution without using generative models for control. Experimental results suggest that the system can generate valid scene representations, support multimodal correction, and preserve drawing order during physical execution.
This letter introduces a language-guided object removal that combines neural field resampling with multiview-consistent progressive inpainting, a direct NeRF weight editing method utilizing knowledge distillation, and the first benchmark (NEO-Dataset) for quantitatively evaluating NeRF scene editing methods suitable fo...
M. Zieliński, David Hall, Dominik Belter et al.· IEEE Robotics and Automation...· 0 citations
Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can sc...
Yian Wang, Jun-Yi Cao, Xiao-Wen Qiu et al.· 0 citations
GenGX is a system that generates precise geometric diagrams from natural-language descriptions by combining large language model (LLM) interpretation with symbolic constraint solving by combining large language model (LLM) interpretation with symbolic constraint solving.
Kavi Wilson, P. Todd· SIGGRAPH Posters· 0 citations
Recent text-to-video systems can generate visually appealing clips from natural language prompts, yet narrative prompts often contain multiple implicit temporal stages that require the generator to infer scene decomposition, subject persistence, action ordering, and visual continuity from a single unstructured input. T...
Jing Chen· 2026 International Conferenc...· 0 citations
Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constraints, resulting in target ambiguity and physically implausible outcomes. To address this, we propose a plan--guide--edit framework that explic...
Sen Liang, Fengbin Guan, Youliang Zhang et al.· 5 citations
RoomWright is presented, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction, providing interactive environments for embodied AI and policy learning.
Zijian Xiao, Zi-Peng Ye, Jin-Kun Hao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.