Skip to content

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling

Jul 2026 · arXiv.org · Vol abs/2607.15619 · 0 citations · 48 references
Computer Science

TL;DR

StructGen is proposed, which employs a structured, dictionary-like format to encode multiple reference images, thereby enabling explicit and unambiguous specification of generation intentions, and consistently outperforms existing methods on both semantic alignment and detailed reference-generation consistency.

Abstract

Multi-reference image generation aims to synthesize images by integrating attributes from multiple reference images under textual instructions. As the number of references increases, the task necessitates complex semantic comprehension, such as correctly associating attributes with the intended subjects and planing out coherent spatial arrangement between subjects and their environments. Existing approaches, which rely solely on natural language instruction, often fail to capture these complex intentions precisely, leading to semantic misalignment and inconsistent generation. We identify two key factors behind these limitations: natural language instructions are often verbose and ambiguous, and high-quality multi-reference data is scarce. To address these issues, we propose StructGen, which employs a structured, dictionary-like format to encode multiple reference images, thereby enabling explicit and unambiguous specification of generation intentions. To support this design, we construct a structured dataset based on high-quality real images and develop a corresponding training framework, along with a dedicated benchmark for challenging multi-reference scenarios. Extensive experiments on both public benchmarks and our proposed benchmark demonstrate that StructGen consistently outperforms existing methods on both semantic alignment and detailed reference-generation consistency, especially under complex instructions with multiple references. The code is available at https://jianingpeng0382.github.io/StructGen/

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Cost-efficient Active Learning for Referring Image Segmentation and Grounding

This work formulates active learning for VG under the realistic setting where only raw images are available without accompanying text, and introduces Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates.

Junbeom Hong, Seonghoon Yu, Hyungsik Jung et al. · 0 citations
Preprint Aug 2026

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

Recognizing that diverse multi-reference tasks share a common set of atomic operations, this work formalizes four operators: Anchor, Disentangle, Disentangle, Apply, and Compose, and constructs TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8.

Haoran Wang, Chaofan Ma, Ran Yi et al. · 0 citations
Conference Aug 2026

Noun Presence Loss with Dynamic Weighting for Compositional Text-to-Image Generation

Compositional text-to-image generation requires faithful depiction of multiple objects with distinct visual attributes. While recent inference-time optimization methods have substantially improved attribute binding, entity neglect - where specified objects are absent from the generated image - remains an unresolved fai...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Preprint Aug 2026

CoCo-IR: Contextual Composed Image Retrieval

A new model based on a Large Multimodal Model (LMM) that functions as a context-aware reasoner for CoCo-IR is proposed, which interprets the entire interaction history to generate Transformable Image Embeddings (TIE) that evolve across turns.

Shengcao Cao, T. Dabral, Z. Ding et al. · 0 citations
Conference 2026

Zero-Shot Multi-Reference Personalization via MLLMs-Guided Layout Planning

This work proposes RIG (Regional Image-prompt Generation), a novel training-free framework for multi-reference personalized generation that significantly outperforms state-of-the-art adapter methods in terms of both personalization fidelity and text-layout alignment.

Junhao Feng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.