Skip to content
Review

Developing a naturalistic approach for an experimental question: Perception of acoustically ambiguous phonemes within semantically disambiguating discourse

Aug 2026 · Journal of the Acoustical Society of America · Vol 159, pp. A321-A321 · 0 citations

Abstract

While acoustic details such as voice onset time can be informative as to the identity of a phoneme (e.g., GOAT versus COAT), these details are not always reliable. Fortunately, disambiguating cues can often be found elsewhere, such as the semantic context. Indeed, acoustically ambiguous phonemes can be perceptually biased toward a semantically congruent percept (e.g., “The girl milked the ?OAT,” resulting in the percept GOAT). However, these findings come from well-controlled studies where participants repeatedly attend to specific contrasts produced in isolated sentences by a hyper-articulate speaker, which may limit generalizability to real-world listening tasks. Therefore, the current project aims to reconsider this research question in a more naturalistic listening scenario. A two-speaker referential communication task was used to elicit English discourse about visual scenes. Subsequently, statements containing critical minimal-pair referents were replaced with experimental trials containing acoustic × semantic manipulations. Participants in the current study will be told to listen to the conversation and show they’re paying attention by clicking referents in the on-screen image. Supposedly, their primary task will be to complete intermittent survey items about the speakers’ social communication. In reality, we will use eyetracking to infer their phonological and lexical processing during the embedded trials.

View source

Similar papers

Aug 2026

EXPRESS: Measuring Individual Differences in Word-Meaning Disambiguation.

Theories of the mental lexicon must explain how people use context to interpret ambiguous words (e.g., "internal organ" vs. "musical organ") and why people vary on this core aspect of comprehension. Theoretical development has been hampered by the lack of reliable tests of disambiguation skill. We introduce a task in w...

L. M. Blott, A. Gowenlock, A. J. Parker et al. · 0 citations
Preprint Aug 2026

When Vocal Tone and Literal Meaning Diverge: An Acoustic-Semantic Incongruity Study for Large Audio-Language Models

Affective cues across modalities may be incongruous (e.g., sarcasm or mocking praise), potentially leading to misinterpretation when relying on a single modality. Large Audio-Language Models (LALMs) have recently gained popularity and been applied to multimodal emotion recognition, but their ability to disentangle acou...

Yu-Wen Chen, William Ho, Maxim Topaz et al. · 0 citations

When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection

Sarcasm detection is addressed through sarcasm detection, evaluating Qwen2.5-Omni and Qwen3-Omni on Mandarin Chinese and English under five modality conditions that decompose the contributions of lexical content, vocal semantics, and prosodic structure to reveal a shared stereotype of expressive prosody.

Yong-Jian Chen, Pengfei Wei, Yiqun Sun et al. · 1 citation
Open access Sep 2026

Neural tracking of surprisal and semantic distance in naturalistic movie viewing

Understanding speech requires listeners to integrate incoming input with prior linguistic and thematic knowledge to access meaning, a task greatly aided by prediction. Surprisal and related phenomena (e.g., next word prediction) tend to be associated with broad activation of language regions during listening. A major c...

Ryan M. O'Leary, Hailey C. Smith, Emily B. Myers et al. · 0 citations
Open access Sep 2026

Verbal recoding underlies visual ambiguity resolution, even in a nonlinguistic context

In language, word order can be used to help disambiguate the grammatical category (noun, verb, etc.) to which an unfamiliar or ambiguous word belongs. We have recently demonstrated that the order of stimuli impacts how ambiguous stimuli are subsequently categorized, even in a nonlinguistic context. However, it remains...

Angelle Antoun, Jacqueline Zhang, Fei Xu et al. · 0 citations
Preprint Sep 2026

Louder, Longer, Livelier: Acoustic Shortcuts and Underspecified Rationales in Speech LLM Judges

LLM-as-a-judge is widely used for evaluating text, but extending this paradigm to speech requires models to interpret acoustic as well as linguistic evidence. This introduces a modality-specific risk: a speech judge may treat a perceptually salient cue as evidence of quality even when that cue is irrelevant to the targ...

Ming-Yue Huo, Shivam Mehta, Bhavin Jawade et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.