Skip to content
Preprint

Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

This work takes a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes, and proposes EditMod, which compares source- and target-conditioned predictions under a shared autoregressive context.

Abstract

Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to-fine prediction. However, their potential for text-guided image editing remains largely underexplored. Existing training-free VAR editing approaches often formulate editing as target-conditioned regeneration guided or constrained by the source image, and may rely on inversion, test-time optimization, attention control, or user-provided masks. This generation-centric formulation does not fully exploit the multiscale source representations provided by VARs and may introduce additional computation or intervention. We instead take a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes. Based on this perspective, we propose \textbf{EditMod}, which compares source- and target-conditioned predictions under a shared autoregressive context, treats their difference as a scale-wise editing direction, and applies it as a residual update to source tokens at selected scales. Experiments show that EditMod achieves leading source-image fidelity while maintaining strong text alignment, and completes end-to-end editing of a 1K image in only 1.57 seconds on a single A100 GPU without per-image preparation.

View source

Similar papers

Jul 2026

WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing

This work revisits one-step image editing from a spatially controlled perspective and proposes WhereEdit, a framework that reformulates one-step editing as localized adaptive editing that consistently outperforms existing one-step image editing methods, achieving superior editing quality while maintaining the efficiency of one-step generation.

Ming Hu, Ming-Yu Dou, Jian-Fu Yin et al. · 1 citation
Preprint Sep 2026

Overpainting: Localized Context-aware Diffusion Image Editing

We present"overpainting", an image editing operation which offers both control over the location of the edit and awareness of the previous content in that location. The overpainted area is given by a trimap, where white-annotated pixels must be edited, gray-annotated pixels may be edited, and black-annotated pixels must not be edited. This enables both precise and loose control, depending on user intent. We implement overpainting by adapting a pretrained image editing diffusion model using a combination of joint attention and low-rank adaption across input images with attention-dropout to balance the information flow between noise, source and mask images. We present a novel, automated, training data generation pipeline that (1) generates a set of candidate image pairs leveraging existing language-based editing models, (2) carefully curates those pairs, and (3) extracts a trimap from each usable pair. We demonstrate the versatility of our overpainting model on a wide range of editing tasks.

Sam Sartor, Iliyan Georgiev, Michael Fischer et al. · 0 citations
Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-Image (T2I) sub-block within the MM-DiT joint self-attention naturally encodes highly discriminative spatial localization signals. In this paper, we propose MaskFlow, a training-free framework for spatially localized image editing in rectified flow models. By strategically extracting and aggregating these attention maps from edit-relevant tokens during the standard ODE forward passes, MaskFlow automatically constructs a soft spatial mask. This mask confines semantic edits to the target region while perfectly preserving the original background. MaskFlow operates as a lightweight, plug-and-play extension without retraining the baseline model. Experiments on the FlowEdit benchmark show that MaskFlow reduces LPIPS by 28.1% relative to the baseline while maintaining competitive semantic alignment.

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Preprint Aug 2026

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is the first gradient-based test-time alignment framework for next-scale autoregressive image generation, and introduces the mechanisms needed to make such optimization stable across scales, together with an extensible objective space that any differentiable constraint on cross-attention can plug into.

Hossein Shahabadi, Niki Sepasian, M. Baghshah · 0 citations
Jul 2026

OSVE: One Step Video Editing with One Step Diffusion Models

OSVE is presented, the first framework to successfully adapt one-step Text-to-Image (T2I) models for high-quality video editing, addressing the core challenges of inversion, editability, and temporal consistency.

Habin Lim, Gyeong-Moon Park · 0 citations
Jul 2026

InnoText: A Unified Model for Visual Text Generation and Editing

This work proposes InnoText, a unified DiT-based framework capable of performing both text generation and editing within a single model, and introduces a Font Size-Aware Modulation module to enhance representations across font scales, a Small-Character Aware Augmentation strategy to improve fine-grained fidelity, and a Task-Specific Region Weighted Loss for adaptive optimization.

Hao-Wei Liu, Runze He, Jian Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.