Skip to content
Preprint

MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs

Sep 2026 · 0 citations · 61 references
Computer Science

TL;DR

MinCU is introduced, a benchmark for grounded minimal-change understanding and Semantic-Guided Implicit Spatial Anchors (SG-ISA), a structured autoregressive method that decomposes prediction into a Think-Locate-Describe sequence, suggesting that an implicit intermediate spatial interface can be more effective than relying solely on model scale for grounded dual-image understanding.

Abstract

Localizing and describing fine-grained differences between near-identical images is a critical yet underexplored capability for multimodal large language models (MLLMs). Existing benchmarks largely assess semantic comparison or single-image grounding in isolation, without jointly requiring faithful description and physical localization. To bridge this gap, we introduce MinCU, a benchmark for grounded minimal-change understanding, where each sample consists of an image pair differing by a single atomic variation in object category, attribute, count, or spatial position, and models are evaluated on their ability to describe the change, localize the changed regions, and identify the changed entity. We further propose Semantic-Guided Implicit Spatial Anchors (SG-ISA), a structured autoregressive method that decomposes prediction into a Think-Locate-Describe sequence. SG-ISA first predicts a semantic cue for the changed concept, then uses discrete spatial anchors as an implicit localization scaffold, and finally generates the change description together with the grounding box. Experiments reveal that even the strongest closed-source MLLMs and recent R1-style reasoning models struggle on MinCU, with most failing to jointly produce accurate descriptions and grounding boxes. Compared to the previous chain-of-thought method, fine-tuning with SG-ISA yields substantial joint improvements in grounding accuracy and description quality while reducing reasoning-token overhead by approximately 26%. These results suggest that an implicit intermediate spatial interface can be more effective than relying solely on model scale for grounded dual-image understanding.

View source

Similar papers

Preprint Sep 2026

Spot-the-shift: Evaluating Grounded Image Difference Captioning of Long-term Changes

Long-term change understanding from images of the same place revisited over time is a challenging task with applications in map maintenance and urban infrastructure monitoring. Prior work addresses it either through pixel-level prediction or difference captioning, neither of which is sufficient to reliably measure how...

Benedetta Liberatori, Nermin Samet, Paolo Rota et al. · 0 citations
Sep 2026

Enhancing target identification and query discrimination for visual grounding

FSG-AID is presented, which integrates fine-grained semantic guidance with an attribute-aware iterative decoder and jointly exploits visual and language features to mine attribute semantics, initialize the target query, and iteratively refine the target representation.

Xiya Bu, Yu Liu, Jizhe Yu et al. · 0 citations
Preprint Oct 2026

TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows

Recent text-to-image models have made substantial progress in photorealism, aesthetics, and text-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or inconsistent spatial relationship...

Shuai Fu, Jing Gu, Jian Zhou et al. · 0 citations
Preprint Oct 2026

VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations

Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predic...

Xian-Da Du, Max W.F. Ku, Weiming Ren et al. · 0 citations
Preprint Aug 2026

ID-VTG: Image-Disambiguated Video Temporal Grounding

The Visually-Guided Disambiguation Aggregation Aggregation (VGD-Agg) framework is proposed, a framework based on a dual-branch fast-slow architecture that enhances discriminability via two learnable tokens and achieves state-of-the-art results on the proposed benchmarks.

Minghang Zheng, Jing Wei, Hong-Yi Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.