Skip to content

Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration

Sep 2026 · 0 citations · 54 references
Computer Science

TL;DR

IMPLIED is an implication modeling method that treats fixed-rule implications as an initial guide while learning to infer and revise accepted and rejected action labels over time, and predicts human implications more accurately than the fixed rule approach and LLM baselines, approaching the performance of a human-label oracle.

Abstract

In Human-Robot Interaction, the standard approach to learn a reward model that represents human preferences for robot behavior consists of three steps. First, the robot collects limited direct evidence from human feedback (e.g., positive or negative binary feedback). Then, the robot utilizes the direct evidence to derive accepted or rejected labels to feasible but unchosen actions using fixed implication rules. Finally, the robot updates the reward model with both the direct and derived evidence. Unfortunately, the fixed rule can hinder preference learning: in a user study with two collaborative simulation environments, human-provided implication labels often differed from the standard fixed rule, and using the human labels substantially improved reward learning with the Preference Learning from Implicit and Explicit Feedback (PIE) algorithm. Consequently, we propose IMPLIED, an implication modeling method that treats fixed-rule implications as an initial guide while learning to infer and revise accepted and rejected action labels over time. Across evaluations on recorded human-robot interaction trajectories and a physical robot pizza-making study, IMPLIED predicts human implications more accurately than the fixed rule approach and LLM baselines, approaching the performance of a human-label oracle. In turn, IMPLIED reduces preference-estimation error and leads to robot actions that are more often rational with respect to a combined reward (which includes the true preference reward and a task-specific reward) compared to baselines. By learning to reason about the implications of human feedback, this work enables more faithful and efficient robot behavior adaptation during human-robot collaboration.

View source

Similar papers

Conference Open access Sep 2026

Empirical Evidence and Analysis of a Critical Pitfall in Reward Learning from Human Feedback

This work provides the first comprehensive empirical investigation of humans' internal models play a mediating role in feedback behaviour through a randomized controlled trial and shows that this relationship is invariant across visual contexts and is robust to three common feedback types.

Taha Shaheen, S. West, Yu Zhang · 0 citations
Preprint Sep 2026

Simultaneous Forward and Inverse Human-in-the-Loop Optimization

Subjective user experience is important to human-robot interaction, but the outcomes users value, and how those preferences vary across individuals and contexts, are often unknown. While inverse learning approaches using human data can help identify user rewards, in many assistive settings the experimental costs of exe...

Kyeongwon Park, Steve H. Collins · 0 citations
Preprint Sep 2026

PreferenceFlow: Test-Time Guidance of Flow-Matching Robot Policies from Human Interventions

Flow-matching policies can represent complex robot behaviors but remain susceptible to local errors under distribution shift at deployment. Many reinforcement learning approaches to policy improvement require reward signals that are difficult to specify or obtain in real-world manipulation. We present PreferenceFlow, a...

Yiqi Tang, Di-Yuan Shi, Run-Ze Li et al. · 0 citations
Preprint Sep 2026

Pack It My Way: Triadic Human-Robot Collaboration for Personalized Autonomous Packing

Personalized autonomous packing requires robots to account for resident preferences that cannot be inferred from scene geometry alone. Expert teleoperators can interpret these preferences and translate them into feasible robot actions, but continuous expert involvement limits scalable deployment. In this paper, we inve...

Sandeep Chowdary Kotapati, Yan-Xin Gao, Tsung-Chi Lin · 0 citations
Open access Aug 2026

Toward Adaptive Interaction Strategies for Human Companion Robot via Deep Reinforcement Learning

In the field of Human-Robot Interaction (HRI), achieving flexibility in human-accompanying within real-world environments holds great potential for various applications but also poses significant challenges. Traditional methods typically restrict robots to fixed positions relative to humans, such as tracking from behin...

Cong-Thanh Vu, Yen-Chen Liu · 0 citations
#artificial intelligence Preprint Sep 2026

Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most common...

Bo-Hao Wang, Xiao-Yan Zhao, Yang Zhang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.