Skip to content

StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios

Sep 2026 · 0 citations · 54 references
Computer Science

TL;DR

StrixAE, an agent based on a multimodal large language model (MLLM), outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.

Abstract

Audio enhancement in real-world scenarios involves complex distortion couplings and requires personalized enhancement. Existing solutions struggle to address both simultaneously. To improve robustness and enable autonomous operation in such scenarios, we propose StrixAE, an agent based on a multimodal large language model (MLLM). StrixAE leverages the MLLM as a controller to coordinate multiple audio enhancement and personalization models. To further enhance system robustness, reduce artifacts, and improve generalization across diverse real-world scenarios, StrixAE is trained through a two-stage process: first, CoT supervised fine-tuning on AcoustBench to ground basic reasoning and tool invocation; second, Audio Perception Reinforcement Learning (APRL), a reward design specifically tailored for audio restoration pipelines that jointly optimizes format validity, structural coherence, and perceptual quality. Unlike generic RL fine-tuning, APRL introduces structured rewards that enforce executable pipelines and logical section ordering, enabling the agent to produce reliable, interpretable enhancement plans without hallucinated tools. Based on real-world test datasets, our proposed method outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.

View source

Similar papers

Preprint Sep 2026

OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning

Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The chall...

Xing-Ming Shui, Da-Peng Chen, Bo-Wei Liu et al. · 0 citations
Preprint Aug 2026

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching

Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing, and VoxAudio, a causal autoregressive flow matching model that addresses this problem from three complementary aspects...

Wenxiang Guo, Changhao Pan, Ziyue Jiang et al. · 0 citations
Preprint Aug 2026

DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation

DAVE is presented, a decoupled audio-visual enhancement framework for real-world speech separation that applies scene routing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics.

Wei Zhou, Wan-Yi Ning, Yi Guo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence

In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation:...

Shilong Zou, Shi-Lin Zhang, Yingji Zhang et al. · 0 citations
Preprint Sep 2026

ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a unified diffusion-ba...

Jia-Ran Cai, Xing-Pei Ma, Shen-Neng Huang · 0 citations
Preprint Sep 2026

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

FIND is introduced, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace and reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem.

Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.