This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.
Abstract
A major limitation in the speech sciences is access to naturalistic data in experimental settings and the difficulty of translating laboratory designs to real-world contexts. Researchers studying speech production or perception often rely on in-lab recordings, which constrain the sociolinguistic contexts examined and limit ecological validity. The Toolkit for Acoustic–Phonetic Analysis (TAPA) is an open-source pipeline that automates the acquisition, transcription, speaker diarization, forced alignment, and per-segment acoustic analysis of naturalistic, single/multi-speaker audio. The current release supports vowel formant extraction, stop voice onset time, and fricative spectral moments. We demonstrate TAPA on the 2016 U.S. presidential debate, extracting nearly 33,000 segments from a 90-min recording, and validate each measurement type against hand-coded annotation. Vowel formant agreement with expert measurements was high (F1 r = 0.89, F2 r = 0.87). Stop VOT showed reliable aggregate means but poor per-token agreement (r = − 0.04) because of a training–deployment mismatch in the neural VOT classifier. Fricative spectral standard deviation agreed strongly with hand-coded values overall (r = 0.80), and center of gravity agreed strongly for sibilants (/s/ r = 0.87, /ʃ/ r = 0.94), while non-sibilant moments diverged systematically. These findings suggest that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts’ supervision.
Japanese vowel duration is a key phonetic feature for distinguishing lexical meaning and stylistic register, but its realization patterns across different speech contexts remain insufficiently quantified. This study constructs a bimodal corpus involving 20 native Japanese speakers, including 50 minutes of spontaneous c...
Leading commercial and open-source Text-to-Speech (TTS) models fail to emulate the regional phonetic diversity of Brazilian Portuguese (pt-BR). By aggregating disparate dialects into a single training distribution, they generate a synthetic"diluted"accent: a phonetic profile attempting to represent all regional distrib...
Pedro H. L. Leite, Pedro Benevenuto Valadares, L. Biscainho· 0 citations
This work proposes an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception.
This paper introduces PersianVox, a fully automated pipeline designed to generate high-quality speech corpora from unlabeled web data, using a novel prosody-aware segmentation strategy that utilizes acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling.
Saeedreza Zouashkiani, Soheil Khalesi, Saman Soleimani Roudi et al.· 0 citations
This study investigates the role of prosodic-acoustic features in distinguishing two inland varieties of Brazilian Portuguese, Paraíba (PB) and São Paulo (SP), and examines how these features influence automatic speech recognition (ASR) performance. Grounded in sociophonetic theory and dialect-aware speech modeling, th...
Leônidas Silva, L. Tenani, João Marcelo Monte· Cadernos de Linguística· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.