Experimental analysis shows how distortion affects auditory scene analysis, suggesting potential benefits in comprehension.
Separating sounds into distinct causal components (i.e., auditory scene analysis, as exemplified by the “cocktail party problem”) is fundamental to hearing. We investigated how human comprehension of a single voice in a mixture is affected by various distortions (e.g., reverberation, filters, power compression, etc.), which are common in real-world scenes, in telecommunications technology, and induced by hearing aids. We demonstrate that distorting the mixture of voices reduces comprehension (as expected), but that distorting the target voice only (by comparable amounts) can aid comprehension. This shows distortion can sometimes aid listeners, presumably by making the voices more separable. In another experiment, we assess the human ability to recognize distortion itself, by asking listeners to match equivalent degrees of distortion across recordings of different voices. We show humans can robustly recognize both speaker identity and degree of distortion when both vary unpredictably. As distorted sounds are structured by both the source and the particular structure of the distortion, to successfully recognize the source and/or the transmission channel the auditory system must separately infer the two causal factors. Thus we propose that hearing a single distorted sound may itself be an auditory scene analysis problem, and can be studied as such.
No takes yet. Share an insight, caveat, or question.
Stolley et al. (2025) studied this question.