Aligned Continuous Integrate-and-Fire, a highly efficient framework for zero-shot speech processing that dynamically compresses continuous acoustic frames into the exact discrete token length of the target text utilizing explicit Dynamic Time Warping alignments is introduced.
Abstract
While Large Language Models excel in natural language processing, efficiently extending their capabilities to spoken input remains a significant challenge. Existing methods for building SpeechLLMs often rely on computationally expensive full-model fine-tuning, or employ parameter-efficient projectors that suffer from inefficient token sequence lengths and costly full-model supervision. In this paper, we introduce Aligned Continuous Integrate-and-Fire, a highly efficient framework for zero-shot speech processing. Our method dynamically compresses continuous acoustic frames into the exact discrete token length of the target text utilizing explicit Dynamic Time Warping alignments. This allows our initial training stage to establish a robust acoustic-to-semantic bridge using lightweight distance metrics, entirely bypassing the computationally expensive LLM forward pass. For subsequent fine-tuning, we propose a memory-efficient knowledge distillation objective that targets a single LLM layer, performing competitively with full-model cross-entropy training at a fraction of the computational cost. Through extensive evaluations on Automatic Speech Recognition and Speech Translation, we demonstrate that our method achieves superior performance compared to prior parameter-efficient baselines.
Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We...
FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction, is introduced, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction that substantially outperforms all baselines and remains robust under real ASR transcript...
Mithilesh Vaidya, Stephen W. Bailey, Sumukh Badam et al.· 0 citations
This work proposes Phoenix TTS, a unified framework that tightly couples representation learning with generative acoustic modeling and is optimized to reconstruct self-supervised features to maintain semantic richness, while simultaneously receiving direct supervision from a Flow Matching training loss.
Pei-Jie Chen, Zhuanling Zha, Zhipeng Nie et al.· 0 citations
This paper introduces PersianVox, a fully automated pipeline designed to generate high-quality speech corpora from unlabeled web data, using a novel prosody-aware segmentation strategy that utilizes acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling.
Saeedreza Zouashkiani, Soheil Khalesi, Saman Soleimani Roudi et al.· 0 citations
A frozen Audio Encoder trained on diverse speech understanding tasks as a semantic teacher to regularize the audio feature space and improves text-speech alignment and stabilizes autoregressive generation while keeping the overall system simple.
Fei-Yu Shen, Kun Xie, Yi-Chen Wu et al.· 5 citations· ⚡1
Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can...
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.