Skip to content
Conference

LVR-Draw: A Language-Vision Pipeline for Robotic Drawing with Interactive Scene Verification and Correction

Aug 2026 · 2026 IEEE International Conference on Mechatronics and Automation (ICMA) · pp. 1141-1146 · 0 citations · 22 references

Abstract

This paper presents LVR-Draw, a fully local language–vision pipeline for robotic drawing that integrates structured scene generation, multimodal verification and correction, and deterministic execution within a unified Human–AI–Robot loop. Given a natural language prompt, a Large Language Model (LLM) generates a structured scene representation in a predefined format, which is rendered into an interpretable image. A Vision–Language Model (VLM) then performs visual inspection to detect inconsistencies in object placement and spatial relationships. These observations are processed by the LLM to produce structured editing operations, enabling iterative refinement of the scene. After validation, the refined scene is converted into executable robot instructions through a deterministic pipeline, supporting predictable and reproducible execution without using generative models for control. Experimental results suggest that the system can generate valid scene representations, support multimodal correction, and preserve drawing order during physical execution.

View source

Similar papers

Open access Jul 2026

NEO: NeRF It Once, Edit It Many Times for Continuous Object Manipulation

This letter introduces a language-guided object removal that combines neural field resampling with multiview-consistent progressive inpainting, a direct NeRF weight editing method utilizing knowledge distillation, and the first benchmark (NEO-Dataset) for quantitatively evaluating NeRF scene editing methods suitable fo...

M. Zieliński, David Hall, Dominik Belter et al. · 0 citations
Preprint Sep 2026

IM-ENGINE: Image Editing for Embodied Data Generation

Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can sc...

Yian Wang, Jun-Yi Cao, Xiao-Wen Qiu et al. · 0 citations
Book Open access Jul 2026

Natural-Language to Geometry Diagrams: A Constraint-Based Pipeline for Precise Visual Reasoning

GenGX is a system that generates precise geometric diagrams from natural-language descriptions by combining large language model (LLM) interpretation with symbolic constraint solving by combining large language model (LLM) interpretation with symbolic constraint solving.

Kavi Wilson, P. Todd · 0 citations
Conference Aug 2026

Scene-Level Planning for Temporally Coherent AI Video Generation Using Multimodal Representations

Recent text-to-video systems can generate visually appealing clips from natural language prompts, yet narrative prompts often contain multiple implicit temporal stages that require the generator to infer scene decomposition, subject persistence, action ordering, and visual continuity from a single unstructured input. T...

Jing Chen · 0 citations
Preprint Aug 2026

CoT-Edit: Let CoT Guide Instruction Video Editing

Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constraints, resulting in target ambiguity and physically implausible outcomes. To address this, we propose a plan--guide--edit framework that explic...

Sen Liang, Fengbin Guan, Youliang Zhang et al. · 5 citations
Preprint Aug 2026

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

RoomWright is presented, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction, providing interactive environments for embodied AI and policy learning.

Zijian Xiao, Zi-Peng Ye, Jin-Kun Hao et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.