Skip to content

Multimodal Thinking with Renderable Programs

Sep 2026 · 0 citations · 51 references
Computer Science

TL;DR

This work introduces SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks, and exploits the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process.

Abstract

Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domain, lacking tractability due to rasterized or latent representations of images. We introduce SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks. We exploit the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process. We provide a large curated dataset of SVG-based image editing dataset, as well as the paradigm to tune open-source VLMs. Experiments on a mathematical reasoning benchmark demonstrate that SVGLM achieves strong SVG generation power as well as think-with-image intelligence. Our results highlight SVG as a suitable medium for building more robust digital domain agents, bridging the gap between text-based thinking and pixel-based images.

View source

Similar papers

Preprint Sep 2026

Reasoning with Image Generation

Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.

Nishad Singhi, H. Rodriguez, Aditya Arora et al. · 0 citations
Preprint Sep 2026

GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models

Recent image generation systems increasingly combine multimodal understanding, reasoning, and synthesis, suggesting that they may do more than render plausible scenes. Yet existing evaluations emphasize aesthetics, prompt alignment, compositionality, or text-based answers, leaving unclear whether these systems can solv...

Chang-Peng Zhao, Yi-Ren Song, Jin-Peng Wang · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
Preprint Aug 2026

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent capabilities into image generation, but they either prescrib...

Jiahao Zhao, Xiao-Min Yu, ZhongXiang Sun et al. · 3 citations
Preprint Sep 2026

Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation

Recent advancements in unified generative models (UGMs) and world simulators have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level event alignment, leaving the capacity for high-level visual reasoning underexplored. True visual generative intellige...

Yu-Tong Liu, Nan Huang, Xu Cao et al. · 0 citations
#small language model Preprint Sep 2026

VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning

Across seven benchmarks and four base VLMs, VLM-in-Sandbox achieves the highest sample-weighted average accuracy among Vanilla VLM, Append-only Sandbox, and the proposed method, and is identified as a central abstraction for sandboxed VLM agents.

He-Xiong Yang, Mingrui Chen, Jie Cao et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.