Skip to content
Preprint

Leveraging Inter-object Affordances for Efficient Planning in Contact-rich Tasks

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

This work proposes a method that leverages a TAMP approach, defining object-centric abstractions of execution constraints, called Unified TAMP (U-TAMP), to execute robotic tasks involving interactions among objects with heterogeneous shapes, sizes, and materials.

Abstract

Traditional task-and-motion planning (TAMP) approaches primarily focus on defining sequences of actions along with the necessary geometric and kinematic constraints to execute long-horizon tasks. However, their applicability in real-world settings is limited, as they typically assume simplified object models that overlook key physical properties critical for the successful execution of contact-rich tasks. Moreover, they often use sub-symbolic reasoning during motion planning, which drastically increases planning time and decreases overall success rates. We propose a method that leverages a TAMP approach, defining object-centric abstractions of execution constraints, called Unified TAMP (U-TAMP), to execute robotic tasks involving interactions among objects with heterogeneous shapes, sizes, and materials. Using a Vision-Language Model (VLM), we generate abstractions of inter-object affordances for characterizing physical interaction constraints between objects in contact-rich tasks, such as grasp and support constraints. These constraints are used to enrich the U-TAMP planning domain to deal with objects with variable physical properties. We perform experiments in simulated kitchen table organization scenarios and compare our results with those of the original U-TAMP, as well as a state-of-the-art VLM-based planner that leverages common sense knowledge of objects'affordances for plan generation. Our approach achieves significantly higher planning success rates and improves planning times by one to two orders of magnitude compared to other methods.

View source

Similar papers

Preprint Sep 2026

RoboFolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-object manipulation, where policies must track changing states and execute relia...

Chen-Huan Liu, Yi Xu, Feng Wu et al. · 0 citations
Preprint Aug 2026

Evidence-Gated Task and Motion Planning with Vision-Language Models

Evidence Acquisition and Feasibility Gating (EAFG) is proposed, a framework that acquires visual evidence through VLM-generated exploratory subgoals and TAMP-based execution and applies a feasibility gate to decide whether to proceed with task planning, acquire further evidence, or halt.

Tsunehiko Tanaka, Matthew Stephenson, Alistair Macvicar et al. · 0 citations
Preprint Sep 2026

One Demonstration, Many Objects: Generalizing Manipulation via Local Contact Geometry

Dexterous manipulation with multi-fingered robot hands promises human-level dexterity, but collecting large-scale dexterous robot hand data remains difficult. Learning from human demonstrations has emerged as a scalable alternative to robot teleoperation, providing strong priors on object interaction and contact strate...

Satvik Sharma, Samrat Sahoo, Huang Huang et al. · 2 citations · ⚡1
Preprint Sep 2026

KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation

Humanoid loco-manipulation requires robots to interpret task instructions and scene semantics while executing coordinated whole-body motions. We propose a hierarchical framework that uses motion keyframes as an intermediate representation between Vision-Language Model (VLM) planning and Reinforcement Learning (RL) cont...

Si-Tong Chen, Fatemeh Zargarbashi, Jin Cheng et al. · 0 citations
Preprint Sep 2026

GraspTwin: Zero-Shot Task-Oriented Grasp Optimization via a Digital Twin

As robots transition from structured factory settings into homes, they are required to interact with an ever-increasing variety of objects. Many tasks require grasping, and often it is not sufficient to just pick up the target object. Consider a task like"pouring coffee"--- to facilitate the subsequent pouring, the rob...

Daniel J. Evans, Yin-Long Dai, Simon Stepputtis et al. · 0 citations
Preprint Sep 2026

Search, Ground, Plan: Functional Sufficiency for Task and Motion Planning under Incomplete Scene Knowledge

Foundation models (FMs) have expanded task and motion planning (TAMP) to manipulation problems specified through language and visual observations. However, incomplete scene knowledge leaves a critical gap between understanding what the task requires and knowing whether the physical scene can actually realize it. We int...

N. Vijayakumar, Nav Singhal, G. Varma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.