Skip to content
Preprint

DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable

Aug 2026 · 2 citations · ⚡ 1 influential · 47 references
Computer Science

TL;DR

This work forms image-to-editable reconstruction, which recovers a structured, directly manipulable artifact from a raster image while preserving its visual and semantic content, and introduces DrawAI, comprising an agentic benchmark, DrawAI-Bench, and a reconstruction workflow, DrawAI-Flow.

Abstract

Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their raster outputs remain difficult to use directly because meaningful content and relationships are flattened into pixels, preventing users from inspecting, modifying, rearranging, or reusing individual components. We formulate image-to-editable reconstruction, which recovers a structured, directly manipulable artifact from a raster image while preserving its visual and semantic content. The central challenge is to jointly satisfy Fidelity and Editability, which often trade off in practice. To study this task, we introduce DrawAI, comprising an agentic benchmark, DrawAI-Bench, and a reconstruction workflow, DrawAI-Flow. DrawAI-Bench spans scientific figures, presentation slides, posters, and diagrams, combining real and AI-generated images to reflect practical visual-creation scenarios. It evaluates Fidelity and Editability through a hybrid protocol of 39 criteria: deterministic rule-based metrics measure properties with direct correspondences, while asset-specific vision-language rubrics capture semantic and perceptual qualities for which exact matching is misleading. Besides, we propose DrawAI-Flow, a two-stage agentic workflow in which a Parser Agent turns extracted elements evidence into an explicit reconstruction plan, and a Reconstruction Agent realizes the plan as executable graphics code through an iterative code-render-validate-revise loop. On DrawAI-Bench, we systematically evaluate thirteen models across five agent harnesses to study the effects of model capability, harness choice, and workflow design. The results show that reconstruction quality and costs vary substantially across model-harness configurations, while DrawAI-Flow consistently improves editable structure.

View source

Similar papers

Preprint Sep 2026

VectorHarness: Recovering Editable, Relation-Preserving Structure from Scientific Graphics

VectorHarness is presented, a multi-agent framework for raster-to-authoring reconstruction that recovers heterogeneous components using type-appropriate native representations that recovers heterogeneous components using type-appropriate native representations.

Jia-Hao Tang, Yi-Ren Song, Alex Jinpeng Wang · 0 citations
Preprint Aug 2026

ViSculpt: Visual-Centric Agentic Geometry Editing

This work presents a training-free multi-agent system that edits existing 3D meshes directly in Blender by emulating the iterative workflow of human artists, and views this work as an exploratory step toward visual-centric agentic geometry editing in professional graphics software.

Bo Pang, Jiaqi Pan, Xiao-Chen Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides

PPTBench establishes a measurable testbed for studying visual coding and advancing agents toward more reliable visual creation, and shows that agents can generally produce valid slide files, but struggle to produce high-quality reconstructions that faithfully recover the semantics and visual structure of the target.

Xiao-Qiu Wang, Yi-Zhe Chi, Wen-Yi Li et al. · 0 citations
Preprint Sep 2026

SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalable generation remains challenging. Recent agentic text-to-3D scene pipelines that rely on vision-language models (VLMs) can generate scenes of high fidelity but require costly iterative object placeme...

Xingjian Ran · 0 citations
#computer vision Preprint Sep 2026

Editable Visual Design

Editable Visual Design is proposed, a new paradigm driven by a Coding Agent that achieves both refined aesthetics and production-grade editability and faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers.

Jun-Yan Ye, Wei Liu, Dongzhi Jiang et al. · 0 citations
Preprint Sep 2026

Back2Struct: Making Structured Images Editable Again

Structured images, such as diagrams, charts, and flowcharts, are inherently symbolic and can be compactly represented in an editable format, yet in practice, they are often rendered as images, and therefore not graphically editable. This mismatch presents a significant challenge for researchers, engineers, and designer...

Peng-Yu Yan, Yi-Xin Wu, Yun-Jie Tian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.