Skip to content
Open access

Toolkit for acoustic–phonetic analysis of naturalistic speech data

Sep 2026 · Behavior Research Methods · Vol 58 · 0 citations · 61 references
Medicine

TL;DR

This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.

Abstract

A major limitation in the speech sciences is access to naturalistic data in experimental settings and the difficulty of translating laboratory designs to real-world contexts. Researchers studying speech production or perception often rely on in-lab recordings, which constrain the sociolinguistic contexts examined and limit ecological validity. The Toolkit for Acoustic–Phonetic Analysis (TAPA) is an open-source pipeline that automates the acquisition, transcription, speaker diarization, forced alignment, and per-segment acoustic analysis of naturalistic, single/multi-speaker audio. The current release supports vowel formant extraction, stop voice onset time, and fricative spectral moments. We demonstrate TAPA on the 2016 U.S. presidential debate, extracting nearly 33,000 segments from a 90-min recording, and validate each measurement type against hand-coded annotation. Vowel formant agreement with expert measurements was high (F1 r = 0.89, F2 r = 0.87). Stop VOT showed reliable aggregate means but poor per-token agreement (r =  − 0.04) because of a training–deployment mismatch in the neural VOT classifier. Fricative spectral standard deviation agreed strongly with hand-coded values overall (r = 0.80), and center of gravity agreed strongly for sibilants (/s/ r = 0.87, /ʃ/ r = 0.94), while non-sibilant moments diverged systematically. These findings suggest that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts’ supervision.

Read PDF

Similar papers

Open access Aug 2026

A Study on the Differences in Vowel Duration in Japanese in Natural and Written Speech: An Empirical Analysis Based on a Corpus

Japanese vowel duration is a key phonetic feature for distinguishing lexical meaning and stylistic register, but its realization patterns across different speech contexts remain insufficiently quantified. This study constructs a bimodal corpus involving 20 native Japanese speakers, including 50 minutes of spontaneous c...

Peng Liu · 0 citations
Preprint Sep 2026

Synthetic speech detection in Brazilian Portuguese through accent-related features

Leading commercial and open-source Text-to-Speech (TTS) models fail to emulate the regional phonetic diversity of Brazilian Portuguese (pt-BR). By aggregating disparate dialects into a single training distribution, they generate a synthetic"diluted"accent: a phonetic profile attempting to represent all regional distrib...

Pedro H. L. Leite, Pedro Benevenuto Valadares, L. Biscainho · 0 citations
Preprint Sep 2026

Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations

This work proposes an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception.

João Lima, Lucas H. Ueda, P. Costa · 0 citations
Preprint Sep 2026

PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data

This paper introduces PersianVox, a fully automated pipeline designed to generate high-quality speech corpora from unlabeled web data, using a novel prosody-aware segmentation strategy that utilizes acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling.

Saeedreza Zouashkiani, Soheil Khalesi, Saman Soleimani Roudi et al. · 0 citations
Open access Sep 2026

The Influence of Prosodic-Acoustic Features in Automatic Speech Recognition of Two Inland Brazilian Dialects: A Pilot Study

This study investigates the role of prosodic-acoustic features in distinguishing two inland varieties of Brazilian Portuguese, Paraíba (PB) and São Paulo (SP), and examines how these features influence automatic speech recognition (ASR) performance. Grounded in sociophonetic theory and dialect-aware speech modeling, th...

Leônidas Silva, L. Tenani, João Marcelo Monte · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.