Skip to content
Preprint

StepAudio 3 Music Technical Report

Sep 2026 · 0 citations · 28 references
Engineering Computer Science

TL;DR

This work introduces StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation, and achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems.

Abstract

We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-Hz stream from a 65536-entry single codebook, using semantically informed self-supervised and multi-task training to preserve musical structure and reconstruction-relevant information. A flow-matching diffusion Transformer (DiT) predicts continuous StepAudio VAE latents, which our VAE decoder converts into 48-kHz audio. This discrete-continuous design is guided by comparisons of single-codebook VQ, Semantic and Acoustic RVQ, and different DiT configurations. For explicit planning, a Mixture-of-Experts autoregressive model uses ABC notation to produce an intermediate arrangement plan (ABC-CoT) before predicting music tokens, making harmony, rhythm, and melodic structure part of the generation context. A progressive training curriculum and supervised fine-tuning support song and instrumental generation, accompaniment generation from dry vocals, and cover-song synthesis for up to 5 minutes and 30 seconds. With reinforcement learning via direct preference optimization (DPO), the final model achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems, with competitive SongBench results. On the preliminary Artificial Analysis Music Arena Vocals leaderboard, it obtains a Quality Elo of 1105, behind only Suno V5.5 and Mureka and ahead of Suno V5, MiniMax models, and other systems. Audio demonstrations are available at https://stepaudiollm.github.io/step-audio-3-music.

View source

Similar papers

Preprint Sep 2026

StepAudio 3 Gen Technical Report

This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.

Bin Lin, Bo Zhao, Bo-Yang Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Open-Qwen-Music: An Auditable Framework for LLM-Based Music Composition and Diffusion Rendering

We present Open-Qwen-Music, an open reconstruction of Qwen-Music and a fully specified research system for text-to-music generation that couples LLM-based semantic composition with diffusion-based acoustic rendering. The system comprises a 25 Hz single-codebook music tokenizer, a 3B-parameter autoregressive Music LLM,...

Yang-Bin Yu, Ming-Yu Yang · 0 citations
Preprint Aug 2026

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

MusicLayout, an explicit intermediate representation for controlling musical structure in text-to-music generation, is introduced, providing evidence that explicit layout planning can improve long-range structural organization and support layout-level control.

Shuyu Li, Ke-Jun Zhang, Jia-He Lei et al. · 0 citations
Preprint Sep 2026

Rethinking Music Tokenization: A Semantic Codec toward High-Fidelity LLM Music Generation

Discrete audio tokenization has become the critical interface between raw waveforms and autoregressive modeling in recent music generation. As a result, music tokenizers must simultaneously support high-fidelity reconstruction and produce discrete sequences that remain amenable to language modeling. Existing reconstruc...

Hua-Kang Chen, Guo-Bin Ma, Yue-Peng Jiang et al. · 0 citations
Preprint Sep 2026

DiffSynth-Music: Audio-Conditioned KV-Cache Adapters for Controllable Music Generation

Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Music (https://modelscope.cn/models/DiffSynth-Studio/DiffSynth-Music), a framework that adds composable audio conditioning to a music synthesis...

Zhong-Jie Duan, Sheng-Chuan Gao, Hong Zhang et al. · 0 citations
Preprint Aug 2026

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

ReLMCodec is a low-bitrate single-codebook speech codec built upon a preserve--control--refine principle that moves the empirical single-stream predictability--reconstruction frontier in the evaluations, with gains that carry over to downstream text-to-speech (TTS) synthesis in both intelligibility and speaker similari...

Zixiang Wan, Xusheng Yang, Zhengmeng Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.