Skip to content
Preprint

PredErase: Training-Free Object-and-Effect Removal with Predictive Latent Guidance

Sep 2026 · 0 citations · 43 references
Computer Science

TL;DR

PredErase, a training-free inference procedure on frozen FLUX.2 and I-JEPA, improves the native FLUX.2 backbone and is supported is training-free object-and-effect editing of frozen Fill, not replacement of paired-data erasers.

Abstract

Removing an object is not the same as filling its mask. Cast shadows and contact shading usually lie outside the user-provided instance mask M_obj, so a frozen Fill model that edits only that mask leaves the object's photometric footprint on nearby surfaces. Supervised removers learn this joint erasure from paired clean plates. Training-free editors freeze pretrained weights, yet most still treat M_obj as the entire editable support and steer sampling with CLIP or DINO energies that do not predict the occluded scene. We present PredErase, a training-free inference procedure on frozen FLUX.2 and I-JEPA. The method separates where Fill may rewrite pixels from what structure should occupy the hole. A contact-band expansion M_flux of M_obj exposes local residuals on the supporting plane. I-JEPA, pretrained for masked token prediction, supplies a context-conditioned hole target in representation space; sparse projected gradients align decoded Fill completions with that target inside the instance, while coordinates outside the packed support stay locked. Under instance-only masks on RemovalBench, RORD-Val, and DEFACTO-Val, PredErase improves the native FLUX.2 backbone. Supervised removers remain stronger on several full-image appearance metrics; the supported claim is training-free object-and-effect editing of frozen Fill, not replacement of paired-data erasers.

View source

Similar papers

Preprint Sep 2026

UnfoldCRF: Structured Mask Refinement with Image-Conditioned Latent Regions

UnfoldCRF treats refinement as inference in a conditional random field over pixel labels and latent region variables, and zeroing individual messages shows where the corrections come from: the pairwise messages mostly fix boundaries, the region messages mostly fix non-boundary errors.

Chun-Ming He, Rihan Zhang, Lei Xu et al. · 0 citations
Preprint Sep 2026

Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement

This paper proposes a zero-shot framework for constrained latent inpainting with a frozen pretrained Stable Diffusion model, requiring no task-specific training or model fine-tuning and demonstrates effective object removal and context-consistent replacement content.

Arman Taghizadeh, U. Krumnack, Kai-Uwe Kühnberger · 0 citations
Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Preprint Sep 2026

Just Align $\bm{x}$: Aligning Predictions, Not Representations

Representation alignment has become an effective way to accelerate diffusion training, but its benefits do not transfer reliably to pixel-space clean-image prediction. In JiT, we find that auxiliary feature alignment can improve access to semantic features while reducing access to image variation needed for clean-image...

Yu-Yao Zhang, Yu-Wei Hu, Zi-Yang Mai et al. · 0 citations
Preprint Aug 2026

Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS

Feed-forward 3D Gaussian Splatting reconstructs an explicit Gaussian representation from multiple input images in one network execution, making 3D reconstruction increasingly accessible for casual captures. However, such captures frequently contain transient objects that appear in only a subset of the views. Such conte...

Kangmin Seo, Jae-Pil Heo · 0 citations
Preprint Sep 2026

TripleFlow: Training-Free Video Object Removal by Bridging Residual Editing and Native Generation

Video object removal presents a uniquely difficult editing challenge. Because a removal prompt specifies only what to erase rather than what to generate, the model must infer and reconstruct a highly specific occluded background entirely from the surrounding context. Existing training-free methods struggle with this be...

Song-He Wang, Li-Fu Wei, Shuo-Lin Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.