Skip to content

DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning

Sep 2026 · 1 citation · 62 references
Computer Science

Abstract

A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly that. Interactive imitation learning takes a step in that direction by letting a policy practice on its own and calling an expert when it goes wrong. Existing methods decide when to interrupt the learner. A further 2 decisions are left to whichever episode happened to trigger the interruption: which failure to correct, and where the demonstration should start. This paper is a first attempt at making both of them deliberately. DISEIL (Demonstration dIstillation for Sample-Efficient Imitation Learning) marks each failed episode at the step where the policy first becomes unreliable, represents that moment with a geometric descriptor, and groups the failures into recurring failure modes. A vision-language model and a language model read the selected mode and write a request for the next demonstration, and a store of task constraints checks that the request can be carried out before any expert time is spent. No model produces a robot action. Across 5 simulated tasks under state and image observations, changing only what the expert is asked for gives the highest mean held-out success rate in all 10 settings, with a tie in 1, and the margin is widest at the smallest budget we tested. The scope is narrow: a single round of practice at a time, in simulation, with experts that are mostly scripted. The longer-term aim is a learner that also tracks what its demonstration set already covers, and that asks a human teacher for the missing behavior in proportion to the effort each request costs them.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Encore: Few-Shot Agentic Discovery of Manipulation Strategies

Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We in...

Yi-Fan Kang, Zihan Wang, Zhi-Wen Fan et al. · 0 citations
Preprint Aug 2026

Rethinking Demonstration Unlearning in Imitation Learning for Robotics

Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. Retraining without them is the natural reference, but its cost grows with policy and dataset scale, motivating cheaper operators that edit a trained policy. Metrics inherited from machine unlearning, such as f...

Jia-Zhuo Li, Yu Zhang, Yi-Ming Fei et al. · 1 citation
Preprint Sep 2026

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

FIND is introduced, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace and reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem.

Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al. · 0 citations
#small language model Preprint Sep 2026

Memorize, Adapt, Ignore: Diagnosing Robot Learning Mechanisms under Training Data Variation

This work examines both model behavior and internal representations, using the empirical neural tangent kernel (NTK) as the primary diagnostic tool, and develops practical guidance for designing DR schemes, selecting models, and detecting shortcut learning.

Ke Zhang, Danica J. Sutherland, Chao Liu · 0 citations
Preprint Aug 2026

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token is introduced and its data-scaling and temporal-context behavior under the tested recipes are characterized.

Chunkai Yang, An-Dong Yang, Di Huang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.