Electroencephalography and behavioral experiments reveal how statistical learning improves sound source segregation.
Humans leverage statistical regularities in acoustic scenes to perceptually segregate target sound sources from other competing sounds. Acoustic regularities present in natural sounds like speech, including temporal coherence and harmonicity, support bottom-up grouping. Other higher-level regularities like linguistic structure must be learned to form a mental “schema” (stored knowledge) of statistical patterns in the environment. Schema-based segregation has been proposed to rely on predictive modeling of the acoustic scene; yet, prior studies have not established a direct link between prediction and segregation. To address this gap, we combined electroencephalography (EEG) and behavioral experiments. Following established statistical learning paradigms, we exposed listeners to sequences of speech syllables with specific syllable-transition probabilities. Preliminary data suggest that post-exposure, detection of a target syllable in competition improves when learned knowledge of between-syllable transition probabilities in the attended stream predicts the target. This prediction benefit is smaller for speech perception in quiet compared to in competition, suggesting that statistical prediction aids source segregation. Preliminary EEG data suggest that the parietal P300 event-related potential accompanying target-syllable detection occurs earlier when the target is predictable compared to when it is not. These experiments provide initial insights into prediction benefits in auditory scene analysis.
No takes yet. Share an insight, caveat, or question.
Viswanathan et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: