Skip to content
Review Open access

Challenges and opportunities of generative artificial intelligence models in audio/acoustic domain: a comprehensive survey

Aug 2026 · Artificial Intelligence Review · 0 citations

TL;DR

This survey serves as a foundational reference for advancing generative AI techniques across the audio domain, and is the first work to jointly cover all three audio domains and all generative model families, while also providing dedicated evaluation-metric analysis and cross-domain comparative assessments.

Abstract

Generative artificial intelligence (AI) has transformed image and text processing, but its adoption in the audio/acoustic domain remains underexplored due to inherent challenges in modeling long-term temporal dependencies in one-dimensional signals and achieving human-perceptible coherence. This comprehensive survey addresses these gaps by systematically reviewing state-of-the-art generative AI models, including generative adversarial networks (GANs), diffusion/flow-matching models, variational autoencoders (VAEs), recurrent neural networks (RNNs), transformers, and Neural Codec Language Models (Codec LMs), organized around three primary application domains: (1) speech synthesis , encompassing text-to-speech conversion, neural vocoding, voice conversion, and zero-shot voice cloning; (2) music generation , covering both symbolic and acoustic composition, multi-track generation, and style transfer; and (3) general audio synthesis, sound effects, and source separation , including text-to-audio generation, audio restoration and enhancement, data augmentation, and conditional source separation. We provide a detailed taxonomy of architectures, functionalities, comparative strengths, and limitations, supported by common evaluation metrics. Structured comparisons with existing surveys demonstrate that this is the first work to jointly cover all three audio domains and all generative model families, while also providing dedicated evaluation-metric analysis and cross-domain comparative assessments. We further highlight emerging opportunities in AI applications such as healthcare monitoring (e.g., symptom analysis and mental health assessment) and biometric authentication, demonstrating the potential of synthetic audio to address real-life challenges. Through research gap identification, model efficacy comparison, and future direction outlining, this survey serves as a foundational reference for advancing generative AI techniques across the audio domain.

Read PDF

Similar papers

Preprint Aug 2026

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

A unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features, and a fine-tuning strategy with auxiliary losses on reconstructed speech to improve both intrusive and non-intrusive metrics is presented.

Yihui Fu, Zhengyang Li, Tim Fingscheidt · 0 citations
Preprint Aug 2026

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

MiDashengLM-Gen is an end-to-end framework that couples a pre-trained Large Language Model (LLM) with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation and drastically improves speech intelligibility over existing unified models.

Xingwei Sun, Heinrich Dinkel, Gang Li et al. · 0 citations
Preprint Aug 2026

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

FullDiT is introduced, a conditional DiT that fuses eight frame-aligned RVQ streams with independently encoded captions and lyrics and uses non-causal self-attention over the complete acoustic latent sequence and outperforms five commercial systems on 15 of 18 automatic metrics.

Yun-Jia Li, Meng-Li Wu, Jun-Yu Dai et al. · 0 citations
Review Open access Aug 2026

Deep learning for speech enhancement: architectures, paradigms, and emerging trends

The need for further research in the development of lightweight models for mobile devices, multi-distortion suppression methods, and the integration of neural network noise suppression with generative models to achieve a new level of speech signal restoration quality is demonstrated.

D. Ivanko · 0 citations
#artificial intelligence Preprint Sep 2026

DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis

The proposed DiTAR+, a dual-optimization framework for continuous-latent Autoregressive Diffusion Transformer models, introduces Dilated Context Sampling and Hierarchical Acoustic Masking to prevent shallow layers from attending to acoustic pre-context, explicitly decoupling semantic alignment from acoustic detail reco...

Zi-Yu Zhang, Tian-Lun Zuo, Han-Zhao Li et al. · 0 citations

Models are Zero-Shot Text

Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.

Unknown authors · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.