GroundBench is introduced, a diagnostic benchmark that separates explanations through six branch-and-merge conditions, each adding a controlled information bundle, and a counterfactual re-ask targeting a real alternate part visible in the same image.
Abstract
A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language models, with no model outperforming a constant baseline until the part was named. However, naming the part supplies information that a real system must infer, confounding visual grounding, mechanical reasoning, and category-to-action association. We introduce GroundBench, a diagnostic benchmark that separates these explanations through six branch-and-merge conditions, each adding a controlled information bundle, and a counterfactual re-ask targeting a real alternate part visible in the same image. Across three OpenAI models and 1,068 predictions, supplying the target region without its identity leaves action accuracy at or below the 0.53 majority baseline (0.26, 0.26, and 0.53), although the models largely reproduce the supplied region. Supplying identity without location instead yields 0.74, 0.68, and 0.68. Every above-baseline gain in this curated set occurs where the supplied part category itself determines the action. A no-vision control leaves GPT-5's scores unchanged or improved, providing evidence consistent with substantial category-to-action association. GPT-4o mini declines on one condition, so this interpretation is not universal. Adding joint type and motion axis does not improve accuracy across six model-stratum comparisons. On 74 counterfactual pairs from 32 objects, GPT-5 achieves 0.86 pair-weighted compliance with a 0.07 shortcut rate but fails all observed push-to-lift-vertical cases. GroundBench identifies which supplied information changes affordance behavior and tests whether apparently grounded performance can be reproduced through textual shortcuts.
Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fails. We separate two steps that affordance questions conflate: identifying which part of an object to act on, and knowing what action that part requires. Across 19 articulat...
Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicitly, requiring robots...
Aernaer Akelijiang, Jian-Nan Li, Zhi-Neng Chen et al.· 0 citations
Locator-Critic (LOCI) is proposed, a training-free framework that decouples visual search from evidence verification and improves accuracy for both open-weight models like Qwen3-VL and proprietary models like Gemini 2.5 Pro.
Walid Bousselham, Mathilde Caron, Arsha Nagrani et al.· 0 citations
The real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.
Siyu Xu, Yun-Ke Wang, Zi-Jian Wang et al.· 3 citations
PatchGate is proposed, a training-free framework that extracts prompt-free object evidence intrinsic to a frozen VLM before generation and uses it to narrow the gap between an intrinsic object set and final object mentions.
Jihyung Ko, Eunji Jung, Hyeongsub Kim et al.· 0 citations
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.