Auditory timing precision is determined by spectro-temporal continuity across sounds, not individual sound acoustics
Abstract
Predicting when an event will occur is key to perceiving and evaluating dynamic sensory inputs, such as speech. The utility of temporal predictions hinges on perceptual timing precision (PTP), the sensitivity to deviations between actual and predicted sound onset times. Past work suggests that PTP is primarily determined by acoustic properties of target sounds, regardless of the acoustics of preceding sound cues. In contrast, in four experiments, we discover that PTP is best when context and target sound spectra and temporal envelopes match, with only minor effects of individual sound acoustics. This ability to utilize contextual continuity improved with musical aptitude. Theoretical modeling showed that spectrally tuned oscillatory entrainment to sound onsets and offsets is necessary to account for our results. In sum, the precision of temporal predictions is defined not by the acoustics of individual sounds but by the continuity of spectral and temporal content within a sound stream.