This work added YOLO-based object and hand detection, stereo vision-based localization using the robot's built-in low-resolution fisheye cameras, and task-specific corrections for grasp execution to form a novel calibration-based grasping pipeline that does not require RGB-D cameras, motion capture, or external tracking systems.
Abstract
Robotic grasping requires accurate coordination between visual perception, object localization, inverse kinematics, and hand control. However, when movements planned in simulation are executed on a physical robot, the sim-to-real gap can cause small positioning errors that prevent successful grasping. In our previous work, we introduced a low-cost haptic calibration method that improved 2D reaching accuracy of the humanoid robot NICO. In this paper, we extend this approach from reaching to tabletop object grasping by adding YOLO-based object and hand detection, stereo vision-based localization using the robot's built-in low-resolution fisheye cameras, and task-specific corrections for grasp execution. Together, these components form a novel calibration-based grasping pipeline that does not require RGB-D cameras, motion capture, or external tracking systems. We also implemented a visual feedback model that aligns the robot hand with the detected object before grasping. Our results show that the fully nonlinear calibration model achieved the best performance inside the calibrated area, while the visual feedback model achieved the highest overall grasping success across the full tabletop workspace.
Developing a human-robot collaborative workplace is the solution to perform faster and more efficient tasks by merging human cognition, awareness, and consciousness with the robot’s power generation, capacity, and precision. In this paper, we address the problem of manipulating linear deformable objects such as cables, ropes, or textiles in a collaborative setup.The proposed method is based on a real-time model-based control algorithm used to position a given point belonging to the object, which is grasped by a human and a robot at its endpoints. The basis of this method lies in (i) the theory of catenaries for modeling the object’s deformation in real-time (ii) the formulation of an interaction matrix representing the robot controller gradient to reach the target position. The experimental results show that the proposed method is reactive to human motion during manipulation and able to reach the desired position accurately.
Racha Ghaddar, A. Koessler, Mourad Benoussaad et al.· 2026 IEEE/ASME International...· 0 citations
The problem of shaping soft objects is widespread in industrial, medical, and household settings. Hence, robotic Deformable Object Manipulation (DOM) is a field of research that has recently emerged to improve robotic systems’ ability to handle such objects. Indeed, human-robot collaboration is also relevant to applications featuring soft objects, since the decision-making and dexterity of a human operator are currently beyond reach.Our aim is to assess the feasibility of using fast finite element inverse simulation in collaborative shaping tasks. To this end, we propose a computationally efficient method for controlling the shape of an object grasped at both ends. In our experimental setup, a leader robot moves freely along unplanned trajectories, while a controlled robot maintains the desired shape despite these perturbations. We reach an update frequency of 20 Hz for the inverse simulation, with average steady-state shape errors below 10 mm in most cases. Those satisfying results allow us to envision our next milestone: the deployment of the inverse simulation in real human-robot shaping tasks.
A. Koessler, T. Raharijaona, H. Courtecuisse· 2026 IEEE/ASME International...· 0 citations
Dual-arm robots often encounter difficulties when handling easily deformable or structurally complex objects using traditional grasping-based manipulation. In addition, grasping and releasing operations introduce significant time overhead. To address these limitations, this paper proposes a vision-based predictive control framework for dual-arm nonprehensile transportation. The proposed method employs a hybrid end effector design that integrates an elastic tether with a tray, enabling flexible and stable transportation without direct grasping. A predictive control strategy is adopted to optimize dual-arm motion trajectories on the move under kinematic and safety constraints. To further enhance coordination accuracy, a direct visual servoing scheme is incorporated to dynamically regulate the arm velocities, minimizing relative motion between the end effectors and the object. This effectively suppresses oscillations induced by the elastic tether. Both simulation and experimental results demonstrate that the proposed approach ensures convergence to desired states and achieves continuous, stable, and safe object transportation, even in the presence of disturbances.
Chang Liu, Yuan Yang, Panfeng Huang et al.· 2026 IEEE International Conf...· 0 citations
Neural-network-based grasp detection has achieved remarkable success in robotic manipulation due to its efficiency and generalization ability. However, detected poses are often not optimized, leading to undesired object motion or collisions during physical execution. This paper proposes a motion-aware refinement framework that minimizes estimated object motion while enforcing collision avoidance. The seven-dimensional pose is decomposed into approach direction, engagement depth, planar projection, and gripper opening width, enabling efficient and interpretable optimization in lower-dimensional subspaces. To evaluate grasp stability beyond conventional success metrics, we introduce the observed success rate (OSR) together with quantitative motion measurements including translation, rotation, and tilt. Real-robot experiments show that, for high-profile objects, the full pipeline improves the measured success rate (MSR) from 93.33% to 100% and OSR from 83.33% to 97.78%. It also reduces the mean translation from <inline-formula> <tex-math notation="LaTeX">$6.099{\,}mm$ </tex-math></inline-formula> to <inline-formula> <tex-math notation="LaTeX">$2.684{\,}mm$ </tex-math></inline-formula>, rotation from 3.732° to 1.344°, and tilt from <inline-formula> <tex-math notation="LaTeX">$4.417{\,}mm$ </tex-math></inline-formula> to <inline-formula> <tex-math notation="LaTeX">$1.313{\,}mm$ </tex-math></inline-formula>, while requiring <inline-formula> <tex-math notation="LaTeX">$0.82\pm 0.42{\,}s$ </tex-math></inline-formula> on average. For low-profile objects that cannot be detected by the baseline point-cloud-based planner, the full pipeline achieves 100% MSR and OSR.
Tian Tan, Redwan Alqasemi, R. Dubey· IEEE Access· 0 citations
Teleoperating a robotic manipulator in industrial environments demands precision that camera-based interfaces alone struggle to deliver. The operator must align the end-effector with a target in clutter, under limited depth perception, and without colliding with the surrounding structures. This paper presents a shared-autonomy framework that assists the operator throughout this process. A single RGB-D camera captures the operator's arm motion and hand gestures without wearables, fiducials, or a calibration stage. The intended target is specified by a free-form text prompt, grounded by a vision-language model in the robot's gripper camera, and tracked across its onboard cameras by a promptable video-segmentation model, resulting in a grasp frame continuously separated from the obstacle map. Every commanded motion is executed by a GPU-accelerated model-predictive controller that enforces self- and environment-collision avoidance against an online volumetric reconstruction, while a potential field corrects the operator's reference toward the grounded target during the final approach. An autonomous mode can be gesture-triggered to complete the grasp on the same target without a separate perception pipeline. The framework is validated on a quadruped mobile manipulator. The interface achieves a positional RMSE of 59 mm relative to motion-capture ground truth, and the controller keeps the arm at least 18 cm from obstacles while the operator deliberately commands the arm into them by 6 cm. In an industrial valve manipulation and a pick-and-place task, the full framework succeeded in all trials, while ablating either the collision or the assistance module produced failures through complementary mechanisms, and autonomous execution succeeded in four of five trials per task.
Murilo Vinicius da Silva, Ricardo V. Godoy, Juliano Negri et al.· 0 citations
A deep examination of the vision-based object manipulation through collaborative robotics with respect to perception pipelines, object detection and recognition, pose estimation, grasp planning, and real time control integration is given.
Rahul Mehta· International Journal of Int...· 0 citations