Skip to content

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

Jul 2026 · arXiv.org · Vol abs/2607.13056 · 0 citations · 44 references
Computer Science

TL;DR

HRIBench is introduced, a diagnostic benchmark for intent-aware human-robot collaboration based on executable interaction scenarios that represents collaborative tasks as structured scenario scripts that explicitly model agent roles, temporal dependencies, coordination constraints, and human behavior distributions.

Abstract

Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled. However, real-world collaboration fundamentally requires coordination under shared agency, including intent understanding, temporal synchronization, protocol adherence, and safe interaction in dynamic environments. To address this gap, we introduce HRIBench, a diagnostic benchmark for intent-aware human-robot collaboration based on executable interaction scenarios. HRIBench represents collaborative tasks as structured scenario scripts that explicitly model agent roles, temporal dependencies, coordination constraints, and human behavior distributions. Building on this abstraction, HRIBench defines three representative interaction roles: Instructor, Collaborator, and Intruder, covering intent communication, joint coordination, and robustness under human intervention. The benchmark contains 13 role-conditioned tasks with over 650 evaluation episodes generated from diverse interaction trajectories and scene variations. Beyond binary task success, HRIBench introduces interpretable interaction-centric metrics spanning synchronization, responsiveness, protocol compliance, and safety. We evaluate adapted policies based on GR00T, pi0.5, and ACT under a unified protocol. Results show that current foundation robot policies struggle substantially in collaborative settings despite strong manipulation ability, revealing major limitations in temporal coordination and intent-aware behavior. Fine-tuning on HRIBench consistently improves collaborative performance. In a real-world adaptation study, simulation data generated by HRIBench improves GR00T N1.5's physical-task success rate from 0.10 to 0.43, demonstrating the benchmark's value for advancing interaction-centric robot learning.

View source

Similar papers

Jul 2026

WCM: World-Cognition Model for Generalizable Human-Robot Interaction

The World-Cognition Model is presented, a human-centered embodied agent built on the SLAK architecture (Sensing, Logic, Action, and Knowledge) and an asynchronous runtime and introduces a human-in-the-loop teaching mode that enables users to interactively teach the robot difficult or long-horizon tasks.

Yuzhen Chen, K. Zhou · 0 citations
Preprint Sep 2026

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual...

Ming-Yu Mei, Haojie Xu, Shi-Hao Jin et al. · 0 citations
Preprint Aug 2026

TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment, and shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation.

Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al. · 1 citation
Preprint Aug 2026

Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

This work proposes ReflexVLA, an efficient VLA model designed for reaction-critical manipulation without large-scale robot-data pretraining, which enhances temporal reasoning through latent future prediction and multi-frame temporal fusion within the vision backbone, while reducing deployment latency through batched vi...

Yuxuan Chen, Wanruo Zhang, Xiao Li · 0 citations
Aug 2026

Adaptive Interaction Network for Human Motion Prediction During Human–Robot Collaboration

Human motion prediction during human-robot collaboration is critical for achieving safe and efficient interactions in shared environments. Unlike previous studies that have primarily focused on humans and objects, we focus on the robot-aware human motion prediction task, which explicitly models the influence of robots...

Mengyuan Liu, Yangting Lin, Qiongjie Cui · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.