Skip to content

Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

Aligned Continuous Integrate-and-Fire, a highly efficient framework for zero-shot speech processing that dynamically compresses continuous acoustic frames into the exact discrete token length of the target text utilizing explicit Dynamic Time Warping alignments is introduced.

Abstract

While Large Language Models excel in natural language processing, efficiently extending their capabilities to spoken input remains a significant challenge. Existing methods for building SpeechLLMs often rely on computationally expensive full-model fine-tuning, or employ parameter-efficient projectors that suffer from inefficient token sequence lengths and costly full-model supervision. In this paper, we introduce Aligned Continuous Integrate-and-Fire, a highly efficient framework for zero-shot speech processing. Our method dynamically compresses continuous acoustic frames into the exact discrete token length of the target text utilizing explicit Dynamic Time Warping alignments. This allows our initial training stage to establish a robust acoustic-to-semantic bridge using lightweight distance metrics, entirely bypassing the computationally expensive LLM forward pass. For subsequent fine-tuning, we propose a memory-efficient knowledge distillation objective that targets a single LLM layer, performing competitively with full-model cross-entropy training at a fraction of the computational cost. Through extensive evaluations on Automatic Speech Recognition and Speech Translation, we demonstrate that our method achieves superior performance compared to prior parameter-efficient baselines.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

VETO: Video Efficient Token Optimization for Vision Language Models

Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We...

Gueter Josmy Faure, Hao-Ping Wang, Min-Hung Chen et al. · 0 citations
#machine learning Preprint Sep 2026

FuseAlign: Forced Alignment in the Wild

FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction, is introduced, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction that substantially outperforms all baselines and remains robust under real ASR transcript...

Mithilesh Vaidya, Stephen W. Bailey, Sumukh Badam et al. · 0 citations
Preprint Aug 2026

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

This work proposes Phoenix TTS, a unified framework that tightly couples representation learning with generative acoustic modeling and is optimized to reconstruct self-supervised features to maintain semantic richness, while simultaneously receiving direct supervision from a Flow Matching training loss.

Pei-Jie Chen, Zhuanling Zha, Zhipeng Nie et al. · 0 citations
Preprint Sep 2026

PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data

This paper introduces PersianVox, a fully automated pipeline designed to generate high-quality speech corpora from unlabeled web data, using a novel prosody-aware segmentation strategy that utilizes acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling.

Saeedreza Zouashkiani, Soheil Khalesi, Saman Soleimani Roudi et al. · 0 citations
#natural language process... Preprint Oct 2026

Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization

Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can...

Jeeyoung Yun, Seohwan Yun, Sungwoong Kim · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.