Skip to content
Conference

Evaluation of Vision-Language Models for Task-Oriented Robotic Grasping

Jul 2026 · Signal Processing and Communications Applications Conference · pp. 1-4 · 0 citations · 9 references

Abstract

This study comparatively examines the task-oriented grasping problem, which is of critical importance in robotic manipulation, through modern Vision-Language Models. Within the scope of the study, the performance rates of the GraspMolmo model, specifically trained for robotic tasks, and Gemini ER-1.5, a general-purpose multimodal AI, were analyzed. The evaluation process was conducted using the TaskGrasp-Image dataset, which encompasses a wide range of objects and tasks, through natural language commands and RGB-D images. The accuracy of the grasping coordinates generated by the models was systematically assessed across varying tolerance thresholds, revealing that both models achieved high accuracy rates. The robotics-specialized model demonstrated a notable advantage over the general-purpose model, particularly under strict tolerance conditions. Error analysis showed that the majority of failed predictions targeted functionally incorrect regions of the object rather than falling outside it entirely, indicating that semantic reasoning rather than geometric localization constitutes the primary challenge.

View source