This work introduces thematic diagrams, which transform abstract structures into styl-ized diagrams that enable visual storytelling both structurally and semantically, and proposes a skeleton representation that abstracts diagrams into nodes and edges.
Structured images, such as diagrams, charts, and flowcharts, are inherently symbolic and can be compactly represented in an editable format, yet in practice, they are often rendered as images, and therefore not graphically editable. This mismatch presents a significant challenge for researchers, engineers, and designer...
Peng-Yu Yan, Yi-Xin Wu, Yun-Jie Tian et al.· 0 citations
A post-training paradigm that integrates multimodal alignment (MA) and structural perception (SP) is proposed, MA enhances element interpretation by grounding metadata-defined elements to their visual counterparts, mitigating semantic drift, and SP models layer-aware inter-element spatial relationships to improve hiera...
Yiyang Huang, Zhao-Wen Wang, Simon Jenni et al.· 0 citations
ReVision is presented, an AI-based tool that decomposes visual and textual references into editable conceptual interpretations and visual motifs, enables their recombination across conceptual and visual spaces, and renders each direction as divergent visual-form variations, supporting more divergent exploration during...
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to...
Cheng Yang, Chu-Fan Shi, Hui-Juan Wang et al.· 0 citations
Experiments on DocLayNet, FUNSD, and SROIE evaluate FRAGMENT alongside representative autoregressive, layout-only, and graph-based baselines, providing an empirical analysis of the characteristics and trade-offs of the proposed factorized graph generation framework.
Ayoub El Bouchtili, Guilhaume Leroy-Meline· 0 citations
VectorHarness is presented, a multi-agent framework for raster-to-authoring reconstruction that recovers heterogeneous components using type-appropriate native representations that recovers heterogeneous components using type-appropriate native representations.
Jia-Hao Tang, Yi-Ren Song, Alex Jinpeng Wang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.