Experimental analysis finds phonological factors impact identification accuracy in speech, suggesting variability due to speaker and listener characteristics.
Speech perception is influenced by communicator characteristics and linguistic factors. Using a multi-speaker sound identification task, this study explored the relative contributions of speaker and listener variability. 384 triphones were extracted from 21 native English speakers reading a list of phonetically rich words. The words and the resulting triphones were balanced in terms of sound frequency (five groups), context prominence (three tiers), and word stress (levels: monosyllabic, unstressed/stressed polysyllabic). The mean length of the words was 6.0 sounds (range: 3–14). A total of 22 listeners were instructed to type the consonants they identified in audio clips drawn randomly from all the speakers using a Latin square design yielding a total of 16 408 responses. The mixed effects logistic regression model showed identification accuracy was lowered by lower frequency sounds (z = 4.0), less prominent context (z = 5.7), lack of stress (z = 3.7), and number of sounds in the word from which the triphone was extracted (z = 3.3). Speaker intercept standard deviation was 87% greater than that of listeners. The results corroborate a model of sound identification that considers frequency and context prominence. The results suggest that our experimental design is sufficiently powered to further explore the components of the speech chain.
No takes yet. Share an insight, caveat, or question.
Aalto et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: