Jan 2026· Trends in Hearing· Vol 30· 0 citations· 73 references
Medicine
Abstract
Neural representations of task-relevant sounds are emphasized when they are attended to, compared with when they are ignored. Classic markers of this modulation, such as amplitude change of event-related potentials (ERPs) and attentional modulation indices (AMIs) derived from envelope-tracking analyses, provide robust quantitative measures of top-down attentional strength. However, it remains unclear how well modern auditory attention decoding (AAD) algorithms applied to short electroencephalography (EEG) segments reflect these established neural signatures. Here, we used a two-stream, colocated listening paradigm with fixed and highly regular temporal structure, enabling precise isolation of ERPs and reliable computation of AMI. Participants attended to one of two simultaneous speech streams and detected occasional pitch deviants, while a 64-channel EEG was recorded. We compared three decoding pipelines—a forward linear model-based decoder, a backward linear model-based decoder, and a convolutional neural network (CNN) decoder—in their ability to classify the attended stream from single 4-s trials, a window short enough to reveal performance differences while still supporting above-chance decoding. Importantly, we examined how decoding outcomes relate to classical attentional modulation, including ERP peak amplitudes and AMI. All models achieved significant AAD performance, with the CNN decoder yielding the highest accuracy. Decoding success of all models aligned with known attentional modulation of ERPs, while the forward model decoder exhibited stronger alignment to the N1 peak-related AMI. These findings demonstrate how fixed temporal structure and colocation provide a testbed linking attention decoding to underlying neural mechanisms.
Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech.
Ilia Semenkov, Daria Kleeva, I. Dakhtin et al.· 0 citations
Electroencephalography (EEG) is a vital tool to measure and record brain activity in neuroscience and clinical applications, yet its potential is constrained by signal heterogeneity, low signal-to-noise ratios, and limited labeled datasets. In this paper, we propose FoME (Foundation Model for EEG), a novel approach usi...
Enze Shi, Kui Zhao, Qilong Yuan et al.· Computerized Medical Imaging...· 0 citations
The human auditory system represents sounds at multiple levels, from low-level acoustic features to abstract category- and object-level information. Although selective attention enables listening in complex natural soundscapes, it remains unclear which representational levels are modulated by attention and how this d...
Onnipekka V. Varis, I. Muukkonen, P. Wikman· Communications Biology· 0 citations
Models of auditory attention propose that multiple sound features are integrated into a perceptual object and that selective attention enhances the fidelity of these object-based representations. Similar attentional mechanisms appear to operate on representations held in auditory working memory, yet evidence has largel...
David Peita, Sung-Joo Lim· Psychonomic Bulletin & Revie...· 0 citations
Auditory attention detection (AAD) identifies which of several competing talkers a listener is attending to, a key step toward neuro-steered hearing devices for real-world listening environments with multiple speakers. Most AAD work to date has examined non-tonal languages, leaving tonal languages underexplored even th...
Shalong Samretngan· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.