MusicLayout, an explicit intermediate representation for controlling musical structure in text-to-music generation, is introduced, providing evidence that explicit layout planning can improve long-range structural organization and support layout-level control.
Shuyu Li, Ke-Jun Zhang, Jia-He Lei et al.· 0 citations
DuoTok is presented, a source-aware dual-track music tokenizer for vocal-accompaniment generation based on staged disentanglement, suggesting that tokenizer design is a core modeling problem for multi-track music generation, beyond compression alone.
DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a large language model (LLM) commonsense check...
Haolin He, Renhe Sun, Zheqi Dai et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.