This analysis evaluates algorithm accuracy in identifying creaky voice in vocal samples, suggesting algorithm reliability issues.
Recent research suggests that creaky voice varies in its acoustics, with some types demonstrating irregular cycles, period doubling, or low fundamental frequency, among other characteristics (Keating et al., 2015; Keating et al., 2023). Much of the research on creaky voice uses audio-visual criteria for determining instances of creaky voice (Dallaston and Docherty, 2020), making the identification of creaky voice largely subjective. There has been an increase in the use of COVAREP (Degottex et al., 2014), an algorithm for detecting creaky voice, particularly in the speech of voice disorder populations (Marks et al., 2023; Roy et al., 2024). However, validation of the algorithm beyond its initial development remains to be tested. In this study, we will examine more than 1500 sound files produced by vocally healthy speakers that have been hand-coded for the presence of creaky voice and compare these to the COVAREP output. In particular, we will examine overall accuracy (same labels for human coders and COVAREP), as well as the audio-visual characteristics of false alarms (COVAREP indicates creaky voice when there is none) and misses (COVAREP indicates no creaky voice when it is present).
No takes yet. Share an insight, caveat, or question.
Bellavance et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: