Jul 2026· International Conference on Signal Processing and Communications· pp. 1-5· 0 citations· 20 references
Abstract
End-to-end on-device Automatic Speech Recognition (ASR) systems have demonstrated remarkable accuracy and efficiency in recent years. However, challenges persist in correctly transcribing infrequent named entities (e.g., geographical locations, business entities, person names, etc.) and handling diverse user accents, which remain underrepresented in training datasets. While information retrieval augmentation or Retrieval Augmentation Generation (RAG) for correction of named entities has shown promise in knowledge-grounded NLP tasks when paired with large language models (LLMs), its integration into real-time on-device systems is non-trivial due to computational constraints. We introduce a novel lightweight method combining phonetic-aware retrieval, vector-based semantic search and generative correction. The system leverages a lightweight phonetic index for rapid candidate entity retrieval and a dense vectorembedding module to refine predictions as well as model the ASR error output distribution in generative space. Additionally, we introduce a novel approach to model ASR errors in natural language. Experiments on test sets emphasizing place names, monuments, airports, and landscapes yielded an increase in correct Named Entity (NE) recognition accuracy by 9.6% compared to baseline. These gains underscore the efficacy of hybrid retrieval-generation paradigms in resource-constrained environments.
This paper introduces PersianVox, a fully automated pipeline designed to generate high-quality speech corpora from unlabeled web data, using a novel prosody-aware segmentation strategy that utilizes acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling.
Saeedreza Zouashkiani, Soheil Khalesi, Saman Soleimani Roudi et al.· 0 citations
A phoneme-guided text-to-speech augmentation pipeline for ASR that connects multilingual speech generation with candidate-text selection and reference-speech quality control, and proposes phoneme-frequency-guided selection (PFGS), which uses phoneme frequencies from real ASR training transcripts to prioritize candidate...
Zhen Wang, Tian-Rui Wu, Rong-Qi Han et al.· 0 citations
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility r...
Chuan-Meng Bian, Da-Ren Chen, Pei-Xin Chen et al.· 2 citations
Aligned Continuous Integrate-and-Fire, a highly efficient framework for zero-shot speech processing that dynamically compresses continuous acoustic frames into the exact discrete token length of the target text utilizing explicit Dynamic Time Warping alignments is introduced.
Abderrahmane Issam, Yusuf Can Semerci, Jan Scholtes et al.· 0 citations
This work curates a human-annotated topic-utterance judgments dataset and examines the superior performance of lightweight LLM matchers over embedding and regex models when equipped with natural language descriptions.
Saman Rahbar, Xi-Liang Zhu, Irvin Cardoza et al.· 0 citations
ParaASR is introduced, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step and shows that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal...