Skip to content
Conference

Structure-Aware and Frequency-Guided Diffusion Framework for Multimodal Fashion Image Editing

2026 · Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada · pp. 4372-4379 · 0 citations

TL;DR

A Structure-Guided Textual Mask Network is designed to predict geometry-aware editing regions by leveraging refined textual structural cues and human-centric priors, where a structural prior reweighting mechanism is introduced to improve localization accuracy.

Abstract

Fashion image editing aims to modify target garment attributes under textual and reference guidance while preserving non-target contents. However, existing methods often suffer from inaccurate garment localization, insufficient preservation of high-frequency textures, and unnatural transitions near edited boundaries. To address these issues, we propose a structure-aware and frequency-guided framework for multimodal fashion image editing. Specifically, we design a Structure-Guided Textual Mask Network to predict geometry-aware editing regions by leveraging refined textual structural cues and human-centric priors, where a structural prior reweighting mechanism is introduced to improve localization accuracy. We further develop an adaptive frequency-domain texture enhancement module to inject high-frequency fabric details from a reference image during late denoising, and employ a boundary-band soft fusion strategy to ensure smooth visual transitions. In addition, we construct a new dataset, DFEdit, for fine-grained multimodal fashion image editing. Experimental results show that the proposed method achieves competitive performance in terms of editing fidelity, texture consistency, and visual quality, showing its effectiveness for intelligent fashion image editing applications.

View source

Similar papers

Open access 2026

Learning Dynamic Spectral Blending for Seamless Text-Guided Image Inpainting

Image inpainting is a fundamental task in computer vision and multimedia processing. With the rapid development of denoising diffusion models, text-guided image inpainting has gradually become a mainstream research direction, enabling flexible content creation and localized semantic editing. Although existing text-guid...

Xingguo Jiang, Chong-Guang Wang, Ming-Ju Chen et al. · 0 citations
Aug 2026

Precision-aware localization and bidirectional feature enhancement for local image editing

An instruction-guided image editing framework that integrates a Precision-Aware Localization Module and a Bidirectional Feature Enhancement Adapter that facilitates accurate and instruction-aware region localization without requiring any user-provided masks is proposed.

Yaya Lv, Li Liu, Daoguang Han et al. · 0 citations
Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Conference Jul 2026

Dual-path attention modulation for training-free text-guided image editing

A training-free Dual-path Attention Modulation (DAR) framework that decouples semantic edits while preserving source image structure is proposed and Adaptive Self-Attention (ASA) and Adaptive Cross-Attention (ACA) modules that dynamically regulate attention replacement are introduced.

Tong Cui, Jie Yang, Kai-Ru Li et al. · 0 citations
Jul 2026

Region-aware diverse image stylization: enhancing fidelity and diversity through object-background augmentation

The region-aware diverse stylization (RDS) method is proposed, which generates multiple distinct stylized images from a single-style image without additional training and significantly outperforms state-of-the-art approaches in both fidelity and diversity.

Yang Wen, Yu-Hang Zhuang, Wuzhen Shi et al. · 0 citations
Conference Aug 2026

HierDiff: hierarchical saliency-guided diffusion for template synthesis

Automated template generation plays a vital role in digital advertising, filmmaking, and virtual environments, yet challenges remain in spatial control particularly in maintaining overall coherence between salient foreground and background regions. We propose HierDiff, a novel saliency-guided diffusion framework that u...

Guangwu Liu, Li Liu, Jibin Liang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.