Backchannel-Aware Transcription for Multiparty Conversation: ASR Omissions, Acoustic Recovery, and Cross-Corpus Transfer
Abstract
Collective-state research increasingly treats automatic transcripts as a substitute for raw audio. We show this is not free under one widely used pipeline: WhisperX omits approximately 48% of hand-labeled backchannels (“mm-hmm”, “yeah”) in close-talk multi-party recordings, and the omission replicates across held-out recording groups. Ablations (disabling voice-activity detection, prompt-biasing the decoder) are consistent with a decoder/language-model-stage failure in the tested pipeline rather than a voice-activity-detection failure. Because backchannels are a vocal listener-response cue of the kind an established literature already treats as a direct proxy for dominance, cohesion, rapport, and engagement, this is a measurement-validity concern for collective-states research using this class of transcription system, not only a transcription-quality complaint. We present an acoustic recovery pipeline (a candidate-generation front end plus a fine-tuned self-supervised speech classifier) and report its ceiling: event-level precision/recall of approximately 0.60/0.26 on held-out groups, bounded by a candidate-generation ceiling and base-rate effects under the tested recording conditions. We test cross-corpus transfer using the AMI Meeting Corpus: the detector transfers zero-shot (AUC 0.86), but pretraining on AMI’s 12, 573 human-labeled backchannels and fine-tuning on our data does not improve recovery over training on our data alone under the tested transfer protocol. Finally, backchannel rate is associated with self-reported voice inclusion (Spearman ρ up to 0.31, participant-clustered 95% CI [0.10,0.50], n = 111, Holm-adjusted p =.006 across six tested associations); a weaker association with self-reported engagement does not survive that correction. We report these as exploratory, single-corpus findings and discuss what they do and do not establish for collective-state measurement from multi-party audio.