Skip to content
Conference

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 574-579 · 0 citations · 21 references

Abstract

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-Image (T2I) sub-block within the MM-DiT joint self-attention naturally encodes highly discriminative spatial localization signals. In this paper, we propose MaskFlow, a training-free framework for spatially localized image editing in rectified flow models. By strategically extracting and aggregating these attention maps from edit-relevant tokens during the standard ODE forward passes, MaskFlow automatically constructs a soft spatial mask. This mask confines semantic edits to the target region while perfectly preserving the original background. MaskFlow operates as a lightweight, plug-and-play extension without retraining the baseline model. Experiments on the FlowEdit benchmark show that MaskFlow reduces LPIPS by 28.1% relative to the baseline while maintaining competitive semantic alignment.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.