Skip to content

SciForma: Structure-Faithful Generation of Scientific Diagrams

Jul 2026 · arXiv.org · Vol abs/2607.18091 · 0 citations · 66 references
Computer Science

TL;DR

Multi-Dimensional Conjunctive Preference Optimization (M-DPO) is developed, which enforces simultaneous correctness across all axes and adaptively routes gradients to the most deficient dimension in post-training, bringing open scientific diagram generation close to proprietary-level structural fidelity.

Abstract

Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can invalidate the entire figure, structural fidelity is inherently conjunctive: correctness on one axis cannot compensate for failure on another. Current open-source models fail to satisfy this criterion. Supervised fine-tuning (SFT) learns plausible layouts but cannot reliably ensure structural correctness, while scalar reward-based post-training obscures which structural dimension has failed. To address this, we introduce SciForma, a framework for the structure faithful generation of scientific methodology diagrams. Specifically, SciForma decomposes diagram quality into three structural axes: Component, Arrow, and Text, guided by a structural inventory. Built on this foundation, we curate SciFormaData-700K for structured training and SciFormaBench-2K for logic-verified evaluation. To close the gap left by SFT, we develop Multi-Dimensional Conjunctive Preference Optimization (M-DPO), which enforces simultaneous correctness across all axes and adaptively routes gradients to the most deficient dimension in post-training. The same structural inventory also enables iterative editing at inference time to correct residual errors. This combination allows SciForma-9B to exceed all open-source baselines and GPT-Image-1.5 on both SciFormaBench-2K and AIBench, bringing open scientific diagram generation close to proprietary-level structural fidelity. Our code and data will be available at: https://github.com/microsoft/SciForma.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

A framework for generating large-scale diagram-grounded instruction data by leveraging terminology derived from scientific curricula is introduced, and augmenting existing models such as LLaVA OneVision with SciGram establishes new state-of-the-art performance on diagram question answering.

Raúl Ortega, José Manuél Gómez-Pérez · 0 citations
#artificial intelligence Preprint Sep 2026

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear whether multimodal models can support realistic scientific-reading...

Shenxi Wu, Yu-Hong Liu, Hao-Song Zhang et al. · 0 citations
#computer vision Preprint Aug 2026

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.

Xue-Qing Wu, Ashwin Balasubramanian, Bingxuan Li et al. · 0 citations
Review Sep 2026

SciFigure2Code: An AI-Reconstructed Benchmark for Scientific Figure-to-Code

SciFigure2Code is introduced, an AI-reconstructed benchmark that instead evaluates presentation recovery: generating editable Python programs that preserve how a scientific panel is arranged and read, and provides an auditable testbed for agents that construct editable, visually faithful scientific figure presentations...

Wen-Tao Li, Yi-Bo Wu, Yi-Zhe Chen et al. · 0 citations

Skeleton-Guided Generation and Editing of Thematic

This work introduces thematic diagrams, which transform abstract structures into styl-ized diagrams that enable visual storytelling both structurally and semantically, and proposes a skeleton representation that abstracts diagrams into nodes and edges.

Jiayinhao Li, H. Prof.David, Laidlaw et al. · 0 citations
Preprint Aug 2026

Towards Physics-Faithful Generation of Scientific Diagrams

Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams, and curate and structurally annotate 4.3 million physics images, carry expert-level annotation, and adapt a unified multimodal backbone.

Ming-Hui Zhang, Jinxin Shi, Yi-Fan Chang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.