Skip to content
Preprint

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

Aug 2026 · 0 citations · 21 references
Computer Science

TL;DR

This work introduces ThinkAfford, which decouples high-recall affordance proposal generation from instruction-grounded reasoning and outperforming comparable 3D open-vocabulary and vision-language-model-based 2D-to-3D baselines.

Abstract

Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing methods either predict 3D masks directly or construct them by selecting and fusing intermediate 2D/3D regions. However, they remain vulnerable to two intertwined failure modes: the predicted or selected regions may miss the target interaction area or have unsuitable granularity, while language grounding may confuse visually similar alternatives under relational instructions. To this end, we introduce ThinkAfford, which decouples high-recall affordance proposal generation from instruction-grounded reasoning. Specifically, the Affordance Proposal Generation module first uses learnable affordance prompts and multi-level visual features to predict interaction-conditioned heatmaps, extracting a variable number of fine-grained proposals without parsed object or part names as segmentation prompts. Visual-Prompted Affordance Reasoning then reasons over labeled proposal overlays using the full instruction, returning identifiers in a structured"think-then-answer"response. Moreover, Group Relative Policy Optimization uses proposal-level rewards from lifted 3D overlap to align VPAR selection with final 3D grounding. On the SceneFun3D validation split, ThinkAfford achieves 10.69% AP50 and 25.46% AP25 under the official evaluator, outperforming comparable 3D open-vocabulary and vision-language-model-based 2D-to-3D baselines. Module-level diagnostics further show that APG attains 77.5% recall at 25% intersection-over-union, while GRPO-trained VPAR achieves 72.1% selection accuracy on APG-covered queries, compared with 63.4% under supervised fine-tuning.

View source

Similar papers

Preprint Aug 2026

AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning

Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assume pre-built object-centric 3D geometry and closed affordance ontologies, limiting deployment from raw RGB observations. We present AffordAny, an end-to-end framework tha...

Jun-Qiang Wu, Kai-Hua Tang, Xuan-Wen Chen et al. · 0 citations
Preprint Aug 2026

Zero-shot 2D Grounding with Novel Affordance Types

AffordAnything is proposed, a training-free method that leverages segmentation cues, motivated by the strong correlation between affordance regions and object subparts, and developed AffordAnything+, a trainable variant that learns to combine these cues.

Haomeng Zhang, Raymond A. Yeh · 0 citations
Jul 2026

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

This work introduces Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability, and proposes Mixture-of-Thought-Tokens, a new free-form multimodal grounding method that bridges the perception-reasoning gap.

Tianyi Gao, Han Fang, Tianyi Ding et al. · 0 citations
Preprint Aug 2026

EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation

EgoAfford is introduced, a benchmark designed to connect the semantic roles of participating objects with task-state-aligned visual observations and multi-step planning and EgoLens, a 3B multimodal large language model with role-specific mask decoders, as an in-domain reference model for this joint task.

Xinyuan Guan, Feifan Chen, Xinyu Zhan et al. · 0 citations
Preprint Aug 2026

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting

CausalSplat is a framework that integrates vision-language models with 3D scene graphs to disentangle explicit structural perception from implicit logical inference and achieves state of the art performance on reasoning benchmarks while showing strong generalizability on standard referring and open vocabulary 3D segmen...

Jiayu Ding, Meilu Song, Yun Chen et al. · 0 citations
Preprint Aug 2026

GuideGround: VLM-guided Semantic Understanding and Viewpoint-aware Reasoning for 3D Visual Grounding

This work proposes GuideGround, a VLM-guided framework that complements rather than replaces task-specific grounding models by leveraging VLMs for semantic enhancement and viewpoint-specific hypothesis verification and replaces auxiliary closed-set object classification with VLM-generated object semantic descriptions t...

Yiwen Wang, Yuyang Deng, Yihao Long et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.