Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception
Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as timestamp tokens. However, this generative f...