Skip to content
Preprint

Grounding Free-Form Instructions for Fashion Complementary Image Generation

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

StyleFlow instantiates the task with StyleFlow, a Rectified Flow Matching model that jointly conditions on the seed image and instruction within a single multimodal transformer, and consistently produces instruction-aligned and stylistically coherent garments while reducing architectural complexity and inference cost relative to auxiliary-module approaches.

Abstract

Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must interpret language in visual context. Existing CIG benchmarks rely on rigid template prompts (e.g.,"a photo of a skirt"), failing to reflect natural user queries and obscuring model behavior across levels of linguistic specificity. We introduce fashion complementary image generation with free-form instructions, a multimodal language-grounding setting where a model generates a compatible garment from a seed image and a natural-language instruction. To this end, we enrich three CIG benchmarks with low-, medium-, and high-specificity instructions generated by a vision-language model and validated by human annotators. We instantiate the task with StyleFlow, a Rectified Flow Matching model that jointly conditions on the seed image and instruction within a single multimodal transformer. Across image quality metrics, catalog-alignment analysis, ablations, and human evaluation, StyleFlow consistently produces instruction-aligned and stylistically coherent garments while reducing architectural complexity and inference cost relative to auxiliary-module approaches.

View source

Similar papers

Preprint Sep 2026

GenPuzzle: Benchmarking Visual Reasoning in Image Generation Models

Recent image generation systems increasingly combine multimodal understanding, reasoning, and synthesis, suggesting that they may do more than render plausible scenes. Yet existing evaluations emphasize aesthetics, prompt alignment, compositionality, or text-based answers, leaving unclear whether these systems can solve visual problems and faithfully express solutions in pixels. We introduce GenPuzzle, a benchmark for reasoning-centric image generation. GenPuzzle contains 2,005 problems across 12 tracks, spanning pattern completion, spatial construction, mazes, Sudoku, nonograms, tangrams, board games, matchstick puzzles, orthographic projection, and mathematical visual proof. Each task provides a visual puzzle and requires an image output that preserves the input state while executing a logically valid solution. GenPuzzle uses task-specific evaluation protocols: discrete grid outputs are transcribed and verified programmatically, while visually complex outputs are assessed with tiered, multidimensional, or binary multimodal large language model (MLLM) rubrics. We further select the automatic judge by measuring agreement with human reference scores. Across three frontier generators, the strongest model reaches only 40.57 Macro Overall, revealing frequent failures in logic, geometry, state preservation, and instruction execution. GenPuzzle provides a testbed for measuring progress from image rendering toward visual problem solving.

Chang-Peng Zhao, Yi-Ren Song, Jin-Peng Wang · 0 citations

RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes

Fashion Image Captioning (FIC) plays a vital role in enhancing user experience and product search in e-commerce platforms. Unlike natural scene image captioning, FIC requires fine-grained visual reasoning and knowledge of domain-specific terminology to capture subtle attributes such as neckline and closure types, graphic patterns, and dress silhouettes. Moreover, as fashion inventories evolve rapidly with new trends, styles, and frequently emerging vocabulary, developing training-free captioning solution becomes essential for scalability and real-world adaptability. Instruction-tuned vision-language models (VLMs) offer a promising solution to fashion image captioning dueto their strong zero-shot capabilities and natural language fluency. However, these general-purpose models often lack attribute-level coverage and precision, and tend to hallucinate or misidentify fine-grained fashion details, making them less suitable for high-fidelity applications like product cataloging or personalized recommendations. To address this, we propose RA-CoA (Retrieval-Augmented Chain-of-Attributes), a novel, training-free framework that disentangles fashion image captioning into two interpretable stages: (i) retrieval of relevant attribute sets from a product knowledge base, and (ii) attribute-level reasoning to generate the final caption. RA-CoA is a model-agnostic approach that works with frozen VLMs to improve fine-grained attribute precision in product captions without the need for fine-tuning. Extensive evaluations across diverse VLM model families under different prompting paradigms demonstrate that RA-CoA significantly improves caption quality, achieving an average gain of 26.3% METEOR score over zero-shot captioning. We make our code publicly available.

A. S. Penamakuri, Shreya Shukla, Anand Mishra · 0 citations
Jul 2026

TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models

This paper introduces the concept of instruction-dense visual jailbreaks, in which image-generation models produce detailed, readable, and actionable harmful instructions within images, and proposes TYPO, a black-box framework that exploits this safety gap by automatically generating adversarial TYPOgraphy prompts.

Meng Xie, Li Zeng, Hang Zhang et al. · 0 citations
Aug 2026

Pose-Star++: Semantic-Visual Understanding for Fine-Grained Fashion Image Editing.

Fashion image editing demands high-dimensional, fine-grained control to follow personalized, unpredictable natural-language instructions. Yet current methods are limited by a fundamental trade-off: fashion-specific approaches offer structural accuracy but lack semantic flexibility, while general text-driven editors are semantically flexible but structurally inaccurate. To bridge this gap, we propose Pose-Star++, a training-free, plug-and-play framework that introduces two core innovations: an LVLM-based Understanding Module that shifts from word- to sentence-level semantic-visual comprehension, eliminating cumbersome instruction pre-parsing and enabling robust understanding of complex natural language; a Bidirectional Calibration Module that co-optimizes semantic and structural constraints through forward pose-guided and backward attention-guided refinement, achieving precise, whole-body-reachable region calibration even under challenging in-the-wild poses. We further contribute the first real-world-oriented fashion-editing benchmark with diverse data, instructions, and tasks, exposing long-overlooked practical challenges. Extensive experiments demonstrate that Pose-Star++ significantly outperforms existing methods in semantic alignment, pose robustness, and in-the-wild generalization across complex scenarios, advancing toward practical, user-guided fashion creation.

Yuran Dong, Bo Du, Mang Ye · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.