Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
ST-Omni-R1 is proposed, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning, and results on three public spatial-audio benchmarks indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.