Skip to content

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

Jul 2026 · arXiv.org · Vol abs/2607.23606 · 0 citations · 44 references
Computer Science Engineering

Abstract

Recent Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription rely on Grapheme-to-Phoneme (G2P) labels, but the phoneme labels are not necessarily phonetically faithful. To investigate this issue, we evaluate zero-shot phonetic classification on Chinese aspiration and Japanese moraic nasals. A PFM trained on G2P-labeled data excluding these two languages yields poor accuracy on both tasks, showing that multilingual coverage with discrete IPA tokens is not sufficient for unseen settings. To overcome this limitation, we propose a classification method based on continuous Articulatory Feature (AF) vectors extracted from each frame. This AF-based approach outperforms discrete token-based methods, particularly for rare phones. We further show that it is crucial to adopt the optimal temporal aggregation of AF vectors for the target distinction: single-frame classification is best for aspiration, while segmental classification substantially improves nasal classification.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.