Preprint
Jul 2026
Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models
The results suggest that competitive MLLM can emerge from alignment alone, reducing multimodal extension to a lightweight projector-training problem that generalizes across modalities and adapts rapidly to each new LLM release.
Xuanru Zhou, Yiwen Shao, Jiahong Li et al.
· 1 citation