Skip to content

SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation

Jul 2026 · arXiv.org · Vol abs/2607.29367 · 0 citations · 15 references
Computer Science

TL;DR

Results suggest that VLM-assisted segment annotation is a practical route to data-efficient, spatially controllable satellite image editing.

Abstract

Satellite image editing requires spatially precise object-level control, but supervised editing datasets for overhead imagery are costly to build because object masks, semantic labels, and paired edits are rarely available at scale. We introduce SatEdit, a mask-conditioned satellite image editing framework that constructs training supervision from unlabeled imagery. SatEdit proposes object masks with a seg- mentation foundation model, assigns semantic la- bels to sampled segments with a Vision-Language Model, and applies lightweight human verification before generating paired addition and removal exam- ples through mask-guided inpainting. We fine-tune a high-resolution image editing backbone with LoRA on a SODA-A-derived dataset containing 1,014 im- ages and 852 verified object annotations across 91 classes. In controlled comparisons with open- source and proprietary image editing models, SatE- dit achieves the highest aggregate masked-region se- mantic alignment, with a CLIP score of 0.6322 and CLIP delta of 0.0726, while preserving the surround- ing scene qualitatively. These results suggest that VLM-assisted segment annotation is a practical route to data-efficient, spatially controllable satellite image editing.

View source

Similar papers

Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Preprint Sep 2026

Semantically Aligned Gradient-Driven Context-Preserving Image Editing

IAB edited achieves state-of-the-art instruction adherence performance on RealEdit and EMU Edit benchmarks based on embedding-based metrics and shows that gradient-aligned VLM distillation holds up under real-world-like surveillance and occlusion conditions.

Chiranjeev Chiranjeev, Muskan Dosi, M. Vatsa et al. · 0 citations
Preprint Aug 2026

RenderMatte: Exact-Alpha Rendering and Group-Relative Alignment for Image Matting

The RenderMatte dataset is constructed, a large-scale synthetic dataset combining 3D-rendered RGBA foregrounds with diverse multi-source assets that features exact strand-level alpha annotations and diverse background composites, demonstrating a scalable path toward high-fidelity matting in open-world scenes.

Ze-Cheng Ren, Ya-Fei Hu, Jianing Zhao et al. · 0 citations
#computer vision Preprint Aug 2026

MaskFlow: Precise, Consistent and Seamless Regional Image Editing

The proposed MaskFlow, a training framework for precise localization, consistent background preservation, and seamless boundary transitions, incorporates the mask into the probability path and flow-matching objective, coordinating generation within the editable region with source preservation outside it.

Rui Xu, Yang Yong, Shun-Zi Yang et al. · 0 citations
Preprint Sep 2026

Overpainting: Localized Context-aware Diffusion Image Editing

We present"overpainting", an image editing operation which offers both control over the location of the edit and awareness of the previous content in that location. The overpainted area is given by a trimap, where white-annotated pixels must be edited, gray-annotated pixels may be edited, and black-annotated pixels mus...

Sam Sartor, Iliyan Georgiev, Michael Fischer et al. · 1 citation
Preprint Sep 2026

Data-Efficient Crosswalk Segmentation from Overhead CCTV via Confidence- and Geometry-Guided Pseudo-Labeling

Pixel-level annotation of fixed traffic-camera imagery is expensive, while crosswalk models trained from street-level imagery face a substantial viewpoint and appearance shift when applied to elevated CCTV. We investigate a data-efficient target-domain pipeline using 241 manually annotated CCTV images and 5,926 unlabel...

Abdirashid A. Omar, Jonghyuk Park · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.