Self-supervised multi-modal imitation learning under skewed trajectory demonstrations.
Abstract
Agent behavior consists of two modalities, state-action trajectories and paired videos. Nevertheless, the absence of viable reward functions in practical workflows often undermines reliable assessment and robust control, a challenge that imitation learning addresses by learning policies directly from expert demonstrations. However, imitation learning heavily relies on expert demonstrations, which are typically characterized by skewed trajectory distributions, leading to mode collapse and insufficient coverage of infrequent behaviors. To address the challenges, we propose a two-stage framework for multi-modal policy acquisition under trajectory distribution skew. Specifically, the first stage performs self-supervised pre-training for future state prediction to capture dynamic features more effectively, providing a strong foundation for subsequent policy learning. In the second stage, we encourage tight clustering of samples from similar modes in the latent space while separating those from different modes with a mode-aware contrastive objective, thus improving multi-modal disentanglement. We further introduce a frequency-aware dynamic scaling factor that reweights mode-conditioned learning signals according to mini-batch mode frequencies, thereby increasing the contribution of underrepresented modes. Beyond state-based metrics, we additionally report a complementary post-hoc assessment using Video-Mode Prototype Consistency Accuracy (VMPCA), which measures whether a generated rollout video is assigned to the expert-video prototype corresponding to its predefined behavior mode. The experimental results demonstrate that our method captured both common and rare expert behaviors from skewed trajectory distributions across diverse multi-modal scenarios. It achieves competitive or superior VMPCA performance, providing complementary evidence that generated rollout videos are frequently assigned to expert-mode prototypes corresponding to their conditioning modes.