2026· Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada· pp. 4372-4379· 0 citations
TL;DR
A Structure-Guided Textual Mask Network is designed to predict geometry-aware editing regions by leveraging refined textual structural cues and human-centric priors, where a structural prior reweighting mechanism is introduced to improve localization accuracy.
Abstract
Fashion image editing aims to modify target garment attributes under textual and reference guidance while preserving non-target contents. However, existing methods often suffer from inaccurate garment localization, insufficient preservation of high-frequency textures, and unnatural transitions near edited boundaries. To address these issues, we propose a structure-aware and frequency-guided framework for multimodal fashion image editing. Specifically, we design a Structure-Guided Textual Mask Network to predict geometry-aware editing regions by leveraging refined textual structural cues and human-centric priors, where a structural prior reweighting mechanism is introduced to improve localization accuracy. We further develop an adaptive frequency-domain texture enhancement module to inject high-frequency fabric details from a reference image during late denoising, and employ a boundary-band soft fusion strategy to ensure smooth visual transitions. In addition, we construct a new dataset, DFEdit, for fine-grained multimodal fashion image editing. Experimental results show that the proposed method achieves competitive performance in terms of editing fidelity, texture consistency, and visual quality, showing its effectiveness for intelligent fashion image editing applications.
Image inpainting is a fundamental task in computer vision and multimedia processing. With the rapid development of denoising diffusion models, text-guided image inpainting has gradually become a mainstream research direction, enabling flexible content creation and localized semantic editing. Although existing text-guid...
An instruction-guided image editing framework that integrates a Precision-Aware Localization Module and a Bidirectional Feature Enhancement Adapter that facilitates accurate and instruction-aware region localization without requiring any user-provided masks is proposed.
Yaya Lv, Li Liu, Daoguang Han et al.· Multimedia Systems· 0 citations
Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
A training-free Dual-path Attention Modulation (DAR) framework that decouples semantic edits while preserving source image structure is proposed and Adaptive Self-Attention (ASA) and Adaptive Cross-Attention (ACA) modules that dynamically regulate attention replacement are introduced.
Tong Cui, Jie Yang, Kai-Ru Li et al.· International Conference on...· 0 citations
The region-aware diverse stylization (RDS) method is proposed, which generates multiple distinct stylized images from a single-style image without additional training and significantly outperforms state-of-the-art approaches in both fidelity and diversity.
Yang Wen, Yu-Hang Zhuang, Wuzhen Shi et al.· The Visual Computer· 0 citations
Automated template generation plays a vital role in digital advertising, filmmaking, and virtual environments, yet challenges remain in spatial control particularly in maintaining overall coherence between salient foreground and background regions. We propose HierDiff, a novel saliency-guided diffusion framework that u...
Guangwu Liu, Li Liu, Jibin Liang· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.