Aug 2026· Multimedia Systems· Vol 32· 0 citations· 53 references
TL;DR
An instruction-guided image editing framework that integrates a Precision-Aware Localization Module and a Bidirectional Feature Enhancement Adapter that facilitates accurate and instruction-aware region localization without requiring any user-provided masks is proposed.
Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
A Structure-Guided Textual Mask Network is designed to predict geometry-aware editing regions by leveraging refined textual structural cues and human-centric priors, where a structural prior reweighting mechanism is introduced to improve localization accuracy.
Xin Chen· Poster Volume 0007 The 2026...· 0 citations
A training-free Dual-path Attention Modulation (DAR) framework that decouples semantic edits while preserving source image structure is proposed and Adaptive Self-Attention (ASA) and Adaptive Cross-Attention (ACA) modules that dynamically regulate attention replacement are introduced.
Tong Cui, Jie Yang, Kai-Ru Li et al.· International Conference on...· 0 citations
Experiments on PIE-Bench and the EditRegion-Bench, with human-verified edit-region annotations for single- and multi-object addition and replacement, show that PC-Edit achieves the best editing quality and background preservation among methods without user-specified edit regions.
We present"overpainting", an image editing operation which offers both control over the location of the edit and awareness of the previous content in that location. The overpainted area is given by a trimap, where white-annotated pixels must be edited, gray-annotated pixels may be edited, and black-annotated pixels mus...
Sam Sartor, Iliyan Georgiev, Michael Fischer et al.· 1 citation
This work revisits one-step image editing from a spatially controlled perspective and proposes WhereEdit, a framework that reformulates one-step editing as localized adaptive editing that consistently outperforms existing one-step image editing methods, achieving superior editing quality while maintaining the efficienc...
Ming Hu, Ming-Yu Dou, Jian-Fu Yin et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.