Skip to content

Labimus: A Simulation and Benchmark for Humanoid Dexterous Manipulation in Chemical Laboratory

Jun 2026 · arXiv.org · Vol abs/2606.31037 · 1 citation · 48 references
Computer Science

Abstract

Laboratory automation has made remarkable progress through robotic platforms and AI-driven scientific reasoning. However, many laboratory operations (e.g., solid--solid transfer) remain inherently dynamic and require real-time adaptation to different materials and experimental conditions. Such precision-critical manipulations are difficult to standardize, motivating the use of humanoid robots with dexterous hands. Despite this opportunity, no existing benchmark evaluates humanoid manipulation in precision-critical laboratory environments. We present Labimus, to our knowledge, the first benchmark for humanoid dexterous manipulation in organic chemistry laboratories. Labimus reconstructs over 30 functionally faithful assets from real organic chemistry workstations through real-to-sim modeling, collectively covering the core operations of routine organic chemistry experiments. The benchmark integrates articulated laboratory instruments, particle-based powder physics, and closed-loop instrument readouts, enabling a complete manipulation-to-measurement pipeline. It further defines six atomic operations and a seven-step solid-weighing workflow derived from real laboratory standard operating procedures. We introduce a precision-aware evaluation protocol designed to jointly measure task completion, experimental precision, and long-horizon execution. We benchmark three representative policies under procedural layouts and environmental perturbations. Results reveal a precision gap: policies that successfully complete laboratory tasks can still fail to satisfy the quantitative tolerances required by experimental protocols. Our benchmark exposes a fundamental disconnect between task completion and experimental validity, providing a new testbed for developing reliable humanoid robots for scientific laboratories.

View source

Similar papers

Preprint Aug 2026

LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories

Autonomous laboratories hold great promise for accelerating scientific discovery. To achieve this vision, robots are supposed to dexterously manipulate diverse labware and instruments and execute long-horizon, state-dependent experimental procedures. Yet existing benchmarks do not jointly capture dexterous hand use, real-world laboratory interactions, and multi-stage experimental procedures, limiting systematic training and evaluation. To bridge this gap, we introduce LabDex, a large-scale real-world dataset and benchmark for dexterous manipulation in chemistry laboratories, organized around a hierarchical task taxonomy spanning atomic skills, compositional tasks, and long-horizon experiments. First, LabDex is cross-platform and, for the first time, unifies real-world and simulation platforms under a common framework, providing standardized task definitions, demonstrations, and evaluation protocols. Second, LabDex is large-scale and systematically organizes chemistry laboratory operations into three interconnected levels: Atomic Skills, which characterize fundamental dexterous manipulation capabilities; Compositional Skills; and Long-Horizon Laboratory Workflows. This hierarchical design not only supports the evaluation of end-task performance, but also enables the analysis of how fundamental dexterous skills compose and influence more complex laboratory operations. We conduct cross-level evaluations of representative robot learning methods in both real-world and simulation environments. The experimental results validate the effectiveness of the LabDex task design and demonstration data, and show that the benchmark supports the training and systematic evaluation of existing robotic policies across laboratory dexterous manipulation tasks at different levels, providing a foundation for further research and development of autonomous laboratory robots.

Zhipeng Tang, Sihan Chen, Sha Zhang et al. · 0 citations
Book Open access Jul 2026

Simulating a Dextrous Hand For Robotics With OpenUSD

Accurate, stable simulation is foundational to robot learning, digital twins, and physical AI. This hands-on course presents a workflow for configuring robot simulations using OpenUSD and PhysX in NVIDIA Isaac Sim, with an emphasis on physics tuning, stability, and asset authoring best practices. The course frames physics tuning as a prerequisite for meaningful training and evaluation, and situates OpenUSD’s composability within the broader physical AI ecosystem. Attendees work with a production Inspire underactuated robot hand asset, inspecting the scene—joints, masses, and collision shapes—to diagnose why the asset fails to simulate stably. Through guided activities, they configure colliders, select appropriate collision shapes, apply collision filters, and resolve overlapping-collider artifacts. They then configure joint drives and tune PD gains—stiffness and damping—using the Gain Tuner to achieve stable, responsive behavior. The session concludes with a demonstration of the tuned asset and an overview of downstream workflows in Isaac Lab. Outcomes include a repeatable workflow for preparing robot USDs for simulation, established best practices for colliders and joint drives, and hands-on experience with OpenUSD and PhysX in a production-grade toolchain.

Ji Yuan Feng, Alexandra Kissel · 0 citations
Open access Jul 2026

ChemGrasp: Affordance-Aware Dexterous Grasping for Laboratory Automation

The transition towards sustainable energy systems demands accelerated materials discovery. This requirement drives the development of fully automated chemical laboratories. While multi-fingered dexterous hands offer the kinematic flexibility required to manipulate complex laboratory glassware, deploying them in safety-critical chemical environments remains a formidable challenge. Existing data-driven grasping models prioritize geometric stability but largely overlook strict functional constraints, often producing grasps that occlude vessel openings and cause sample contamination. To address this critical bottleneck, we present ChemGrasp, an affordance-aware dexterous grasping framework tailored for laboratory automation. Our approach introduces a task-specific affordance module during the inference phase of a generative model, employing an energy-based optimization function to strictly penalize semantic violations. Furthermore, to evaluate execution feasibility in constrained simulated workspaces, ChemGrasp integrates a system-level motion planning pipeline featuring a phased hand execution strategy, enabling collision-free and kinematically reachable trajectories in simulation. Extensive physics-based simulations demonstrate that ChemGrasp significantly elevates the Safe Success Rate (SSR) by eliminating functional violations, reliably executing dynamic grasp sequences in both floating and full-pipeline tabletop settings. Ultimately, this framework demonstrates a simulation-validated step toward adapting robotic dexterity to laboratory safety protocols for autonomous clean energy research.

Xuanwei Liu, Tiewei Shang, Rui Wang et al. · 0 citations
Open access Aug 2026

Robotic safety in self-driving laboratories.

The emergence of autonomous laboratories is accelerating discovery in chemistry, drug discovery, materials science, and related fields by enabling high-throughput, data-driven experimentation. However, the integration of heterogeneous robotic systems, ranging from fixed manipulators to mobile platforms, introduces safety challenges that are not systematically addressed in newly established laboratories. In this context, this work aims to raise awareness of robotic safety among chemists and biologists leading laboratory automation projects who may have limited access to industrial robotics expertise. To support a preliminary evaluation of existing or newly developed automated laboratory systems, we explain and demonstrate the use of a simple, structured safety assessment methodology based on ISO standards and tailored to laboratory environments. The framework combines established robotics safety standards with laboratory-specific considerations, including chemical hazards, human-robot interaction, and dynamic workflows. To facilitate its adoption by scientists, the methodology is illustrated through a case study conducted at the Swiss CAT+ West Hub autonomous laboratory, focusing on a multi-instrument analytical platform integrating collaborative robotic arms and mobile robotic systems. The proposed framework follows a six-step iterative process encompassing system definition, hazard identification, risk estimation, risk reduction, and validation. Its applicability was evaluated through the case study, in which sixteen hazards were identified, with robot-human collisions and chemical exposure representing the most critical risks. Experimental force and pressure measurements further demonstrated that widely used collaborative robots may exceed accepted safety thresholds under realistic operating conditions, particularly as a consequence of end-effector design and task-dependent motion characteristics. Risk mitigation strategies based on dynamic safety zoning, sensor-based human detection, and operational mode control were implemented to ensure compliance with safety requirements. The results highlight the need for systematic, context-specific safety assessments in autonomous laboratories and demonstrate that collaborative robots are not inherently safe without rigorous validation. This work provides a practical framework for the safe deployment of robotic systems in autonomous and digital laboratory environments.

Edy Mariano, Maël Löwensberg, Théo Bloesch et al. · 0 citations
Open access Jul 2026

In vivo feasibility study of humanoid robots in surgery.

Recent advances in actuation, control and learning have rapidly pushed humanoid robots from a distant vision towards near-term real-world deployment1-18. Healthcare is a particularly pressing domain, in which staffing shortages and increasing care demand are widening the gap between clinical workload and available skilled labour19-21. Although current automation has largely focused on digital and logistical tasks22, much hospital work remains embodied, requiring mobility, manipulation and safe interaction in human-designed environments. Humanoid form factors offer unique potential, particularly for assisting with surgical tasks. Traditionally, robotic systems for surgery are purpose-built platforms such as Intuitive Surgical's da Vinci Surgical System23,24, and it remains unclear how close current humanoid systems are to meeting the precision, control and safety requirements of minimally invasive surgery. Here we present a systematic evaluation of contemporary humanoid technology for laparoscopic surgical tasks. We develop a humanoid-based laparoscopic teleoperation framework using general-purpose instruments and assess its abilities through benchtop characterization, dry-laboratory user studies spanning diverse surgical experience levels and in vivo porcine studies. Across these evaluations, we quantify technical feasibility, task performance and clinical readiness relative to established surgical platforms. Together, our study provides an evidence-based assessment of current humanoid abilities and limitations for surgical applications, highlighting both their promise and key technical challenges that must be addressed before clinical deployment.

Zekai Liang, N. Thareja, Peihan Zhang et al. · 1 citation
Preprint Jul 2026

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inadequate for laboratory instruments, which must be approached from their operating side while maintaining safe clearance from surrounding equipment. We introduce BioVLN, a simulation platform for developing and evaluating visual-language navigation agents in biomedical laboratories. BioVLN represents each instrument with three regions: its physical body, a surrounding clearance region, and an operation area in front of the usable side. This model is applied consistently to scene generation, target placement, navigation evaluation, and safety analysis, so success depends on reaching a position from which the instrument can be accessed. BioVLN supports procedural scene generation and manually designed environments, producing 47 scenes and 1667 episodes. Standardized navigation and reinforcement-learning interfaces enable trajectory collection and policy training. Experiments show that geometric exploration reaches 74.4--87.5% success, while sampling multiple valid positions in the operation area improves success to 83.3--92.5% and reduces unsafe proximity.

Zhe Liu, Quan Lu, Zhaohui Du et al. · 0 citations