Skip to content
Preprint

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis, is addressed, making it substantially faster than gradient-based TTA while requiring no additional training.

Abstract

Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference. We address these failures with PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis: severe acoustic noise induces a low-rank affine shift in the multimodal latent space, with more than 90% of distortion energy confined to the leading 60 principal components. PRISM estimates and reverses this distortion from an unlabeled target batch using frozen text prototypes as geometric anchors via three closed-form geometric corrections compiled into a single static projection matrix by Affine Bias Regression. At inference, adaptation reduces to one matrix-vector multiplication in 0.0009 ms, making it substantially faster than gradient-based TTA while requiring no additional training. On UrbanSound8K, PRISM improves over the zero-shot baseline by 12.94 percentage points and surpasses an oracle-assisted TTA baseline by 9.41 percentage points, despite never observing its privileged augmented noise prompts. We further identify the Polyphonic Trap, a principled failure mode of subspace deflation for broadband classes, and resolve it via Confidence-Aware Regression (CAR), recovering up to 8.16 percentage points for the worst-affected class.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis

Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, they still suffer from limited decoding stability when synthesizing long utterances or complex linguistic structures. This instability primarily stems from a restricted historical receptive field and an acoustic inertia dependency within the diffusion decoder, which causes the model to ignore semantic conditions. To address these challenges, we propose DiTAR+, a dual-optimization framework. First, we introduce Dilated Context Sampling to expand the macro-level historical receptive field without violating physical temporal continuity, thereby preventing cumulative error propagation. Second, we propose Hierarchical Acoustic Masking to prevent shallow layers from attending to acoustic pre-context, explicitly decoupling semantic alignment from acoustic detail reconstruction. Extensive experiments show that our framework effectively mitigates pronunciation errors and semantic hallucinations, enhances generation robustness on challenging sentences, and maintains exceptionally high speaker similarity throughout the entirety of long-form utterances. On the linguistically challenging ZH-Hard set, DiTAR+ reduces the word error rate from 12.478% to 9.893%, and on extended utterances of 25 to 35 seconds it improves speaker similarity from 0.741 to 0.759 while simultaneously lowering the word error rate from 2.778% to 2.173%, outperforming both discrete-token and pure flow-matching baselines.

Zi-Yu Zhang, Tian-Lun Zuo, Han-Zhao Li et al. · 0 citations
Jul 2026

Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping

Results show that the original full-resolution LAM is the strongest, that separately trained lightweight models are the most competitive learned approaches, and that representation alignment between the upsampler and LAM matters more than model complexity.

Philipp Schmidt, Huw Cheston, Juan Azcarreta et al. · 0 citations
Preprint Aug 2026

SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification

Fully few-shot class-incremental audio classification (FFCAC) requires recognizing new sound classes from only a handful of labeled examples per session, without forgetting previously learned classes and without any large base dataset. Existing methods typically freeze a pre-trained audio--language encoder and classify with point prototypes, but they suffer from significant performance degradation throughout the sessions due to generic feature representations. We propose SPECTRA, a framework built on a frozen encoder which adds three components. (i) a lightweight trainable adapter that calibrates the generic embeddings to the task; (ii) subspace feature replay, an exemplar-free anti-forgetting scheme that replays old classes by sampling from the low-rank subspace of their stored features; and (iii) a transductive optimal-transport refinement of prototypes at test time. Our central finding is that the subspace structure of the replay diminishes forgetting and outperforms naive Gaussian replay of equal variance. On three FFCAC benchmarks (NSynth-100, FSC-89, LS-100), SPECTRA improves average accuracy and reduces forgetting over current state-of-the-art methods, and our ablations statistically validate each component.

Giries Abu Ayoub, Loay Mualem, Simon Korman · 0 citations
Preprint Aug 2026

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

FullDiT is introduced, a conditional DiT that fuses eight frame-aligned RVQ streams with independently encoded captions and lyrics and uses non-causal self-attention over the complete acoustic latent sequence and outperforms five commercial systems on 15 of 18 automatic metrics.

Yun-Jia Li, Meng-Li Wu, Jun-Yu Dai et al. · 0 citations
Open access Sep 2026

Blind Device-Response Interpolation for Intermediate-Domain Unsupervised Adaptation in Cross-Device Acoustic Scene Classification

Recording-device mismatch is one of the dominant failure modes of acoustic scene classification: a model trained on one microphone degrades sharply on an unseen one. Aligning source and target in a single adversarial step is unstable when the device gap is large, and recent intermediate-domain formulations mitigate this by inserting bridge domains, but they build those bridges from hand-specified parametric device operators that are unavailable at deployment time. We propose BRIDA, an intermediate-domain unsupervised adaptation framework whose bridges are derived from the device response itself. A blind estimator recovers the relative log-mel magnitude response between the labelled source devices and the unlabelled target device from long-term spectral statistics alone, with an optional class-balanced refinement for the case where the two unlabelled pools differ in scene composition. Because a magnitude filter is additive in the log-mel domain, fractional powers of the estimate yield a continuum of label-preserving bridge domains at essentially zero cost; we traverse that continuum with a curriculum and regularise it with an ordinal path-position adversary, a residual maximum-mean-discrepancy anchor and confidence-gated target prototypes. Experiments use 1648 real recordings and 66 measured microphone impulse responses under a leave-one-device-out protocol with disjoint source, target and evaluation content. Averaged over 3 unseen devices and 3 seeds, BRIDA attains 53.00 ± 4.29% target accuracy and 52.99 ± 4.58% macro-F1, improving on source-only training by 7.83 points and on a discrete two-bridge intermediate-domain baseline by 0.61 points, with no inference-time overhead.

Claire Whitmore, Lucas Bennett · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.