Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning
A hierarchical machine teaching algorithm for reward learning that operates across multiple MDPs and achieves substantially lower regret and stronger generalization to held-out environments than uniform teaching baselines under identical feedback budgets, demonstrating the importance of multi-environment, multi-modal teaching for learning dynamics-robust reward functions.