Skip to content
Preprint

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

Aug 2026 · 0 citations · 34 references
Engineering

TL;DR

FullDiT is introduced, a conditional DiT that fuses eight frame-aligned RVQ streams with independently encoded captions and lyrics and uses non-causal self-attention over the complete acoustic latent sequence and outperforms five commercial systems on 15 of 18 automatic metrics.

Abstract

Hybrid music generators combine the long-range planning of an autoregressive language model with the fidelity of a diffusion- or flow-based acoustic renderer. Yet renderers are trained with clean, target-derived codec tokens but deployed with imperfect language-model predictions, creating codecinterface exposure bias. Rather than treating rendering as a simple reconstruction task,we formulate it as full-context generation from an imperfect discrete plan. We introduce FullDiT, a conditional DiT that fuses eight frame-aligned RVQ streams with independently encoded captions and lyrics and uses non-causal self-attention over the complete acoustic latent sequence. During training, Error-Matched Distractor Conditioning (EMDC) matches per-codebook replacement rates to teacher-forced top-1 error rates and samples near-miss tokens from cosine-KNN neighborhoods without changing the acoustic target. At inference, four-way classifier-free guidance (4-CFG) independently scales codec, lyric, and caption guidance increments. Matched ablations show that EMDC improves ViSQOL by 0.77 under synthetic corruption and is clearly preferred in non-tied comparisons with fixed languagemodel tokens. Further ablations show gains from full-song context and renderer-side text conditioning. The complete system outperforms five commercial systems on 15 of 18 automatic metrics and ranks among the top three on the Artificial Analysis Music with Vocals Leaderboard. The demo page is available at https://selinacloudl.github.io/fulldit-demo/.

View source

Similar papers

Preprint Aug 2026

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

ReLMCodec is a low-bitrate single-codebook speech codec built upon a preserve--control--refine principle that moves the empirical single-stream predictability--reconstruction frontier in the evaluations, with gains that carry over to downstream text-to-speech (TTS) synthesis in both intelligibility and speaker similarity.

Zixiang Wan, Xusheng Yang, Zhengmeng Wang et al. · 0 citations
Preprint Aug 2026

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

MiDashengLM-Gen is an end-to-end framework that couples a pre-trained Large Language Model (LLM) with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation and drastically improves speech intelligibility over existing unified models.

Xingwei Sun, Heinrich Dinkel, Gang Li et al. · 0 citations
#machine learning Preprint Sep 2026

TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context

Text-to-speech systems often face a trade-off between natural prosody and efficient inference: higher perceptual quality typically comes at increased computational cost and latency. We present TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU. Speech is encoded by the hierarchical DualCodec representation at 12.5 Hz, which separates a semantic stream from successive acoustic refinements. Our design assumes that prosodic structure is largely established when the semantic stream is generated, and allocates capacity accordingly: a Qwen3-1.7B-derived transformer predicts that stream and thereby the utterance duration, while three progressively smaller Qwen3-0.6B-derived transformers each add one acoustic refinement. Text is tokenized per character rather than by subword. Paired text and audio markers at shared positions support long-form generation with bounded context, and overlapping DualCodec reconstructions are mapped into the VibeVoice acoustic latent space and decoded causally, enabling streaming despite DualCodec's noncausal decoder. The model accepts up to one minute of reference audio for voice conditioning and is designed primarily for English and German, with additional multilingual support. The four predictors total 2.9B parameters; on a single RTX 5090 the streaming path reaches approximately 200 ms to first audio. In separate non-streaming measurements, the end-to-end real-time factor (RTF) is 0.08 for one input and the aggregate RTF is 0.02 across eight concurrent inputs. On our LLM-as-a-judge audiobook-reading benchmark, TontaubeV1 matches ElevenLabs Flash v2.5 and outperforms Fish Audio S2 Pro, the April 2026 Gradium API, and Cartesia Sonic 3 on prosody. The model weights are released on Hugging Face under the Tontaube Community Model License 1.0.

Fritz Cremer, Jonathan Cremer · 0 citations
Preprint Aug 2026

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

CuteTTS is presented, a compact continuous-autoregressive TTS system that reconciles high-fidelity generation with the latency demands of real-time interaction and introduces guidance-step distillation, which absorbs classifier-free guidance and multiple solver steps into a single interval-conditioned student.

Yuqian Zhang, Yao Shi, Kexin Huang et al. · 0 citations
Preprint Aug 2026

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

SIEDD, a discrete diffusion framework for text-guided speech inpainting and editing over hierarchical codec tokens, is introduced and results demonstrate that explicitly modeling the codec hierarchy substantially improves context-preserving speech reconstruction and editing.

Iftach Shoham, Tali Dror, Oren Gal et al. · 0 citations
Preprint Aug 2026

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis, is addressed, making it substantially faster than gradient-based TTA while requiring no additional training.

A. Shukla, R. Thakur, Aryan Das et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.