Skip to content
Preprint

Ludi${}_{\scriptscriptstyle 0.1}$: An Agentic System for Socially Intelligent Robots

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

Ludi is presented, an agentic system for socially intelligent robots that integrates interactive speech, multimodal reasoning, memory, navigation, and learned manipulation, and demonstrates a practical path toward fluid human-robot collaboration today while producing the multimodal interaction traces needed to develop a more deeply integrated foundation model for robots and people.

Abstract

Robot foundation models have substantially advanced perception and control, but natural human-robot collaboration requires more than executing isolated commands. A robot must recognize ambiguity, maintain context across turns, communicate its intentions, and revise ongoing behavior as the user's intent changes. We present $\scriptstyle\mathsf{Ludi}_{\scriptscriptstyle 0.1}$, an agentic system for socially intelligent robots that integrates interactive speech, multimodal reasoning, memory, navigation, and learned manipulation. Its decision-making core is a fine-tuned vision-language model trained on multi-turn interaction traces spanning ambiguous requests, clarifications, corrections, interruptions, mixed social and task dialogue, and multi-step tasks. A purpose-built harness manages the model-tool interaction loop, while specialized navigation and manipulation policies execute physical skills. Ludi${}_{\scriptscriptstyle 0.1}$ demonstrates a practical path toward fluid human-robot collaboration today while producing the multimodal interaction traces needed to develop a more deeply integrated foundation model for robots and people.

View source

Similar papers

Preprint Oct 2026

RobotUse: Allocating Computation, Context, and Decisions

Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUs...

Junhoo Lee, In-Chang Baek, Seungyeon Kim et al. · 0 citations

A Goal-Oriented Agentic Framework For Collaborative Branching Human-Robot Interactions

It is suggested that the benefit of the agentic framework lies primarily in interaction quality rather than conversational efficiency, and the agentic architecture displayed robustness by recovering from non-normative inputs while maintaining strict goal alignment.

Morten Roed Frederiksen · 2 citations
Preprint Aug 2026

$\tau_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

Xiaowei Cai, Yunuo Cai, Bing Chen et al. · 4 citations · ⚡1
Preprint Sep 2026

PhasePlan: Ordered Future-Phase Planning for Robot Brain Models

Robot brain models integrate vision, language, and robot state to generate actions for complex manipulation tasks. Most predict fixed-length action chunks that may span multiple task phases. This can obscure phase transitions and favor frequent action patterns, compromising action timing in dynamic environments. We pro...

Xiao-Yu Yang, Ya-Fen Zhang, Wen-Sheng Li et al. · 0 citations
Preprint Aug 2026

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee...

Dream Team, Rui Chen, Xiangxiang Chu et al. · 5 citations
Preprint Aug 2026

MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration

MistyPilot is presented, a multi-agent LLM framework that interprets high-level natural-language instructions and orchestrates the corresponding skills on the Misty social robot and attains high accuracy on routing, sensor-skill binding, task-state parsing, result reuse, and skill extension up to 100 skills, and lower...

Xiao Wang, Lu Dong, Ifeoma Nwogu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.