Longitudinal data reveal how prosody in child-directed speech aligns with both universal and language-specific structures.
Infant-directed speech (IDS) is highly rhythmic, and in European languages is dominated by patterns of amplitude modulation (AM) peaking at ~2Hz (reflecting prosody) and ~5Hz (reflecting individual syllables). The rhythm structure of spoken Japanese is thought to differ from European stress-timed and syllable-timed languages, depending on moraic units (~10Hz) comprising any onset phoneme and vowel phonemes within a syllable, PA-N-DA. As infant brain must be prepared to acquire any human language, initial speech encoding is likely to utilize language-universal physical acoustic structures in speech. These physical structures are however probabilistic, and may thereby simultaneously accommodate language-specific structures like morae. Here a language-blind computational model of linguistic rhythm based on the amplitude envelope (AE) is used to compute the physical acoustic stimulus characteristics for Japanese. Using ~18,000 samples of natural IDS and child-directed speech (CDS) recorded longitudinally over the ages 0 - 5 years, the data show that the temporal modulation patterns that characterise the AE of Japanese are similar to those found for stress-timed and syllable-timed European languages. However, the AM band corresponding to the syllabic level in CDS/IDS in European languages (2-12Hz) was elongated in Japanese (2.5-17Hz), possibly accommodating the faster modulation peaks reflecting morae. Further, the phase synchronization ratios between the two slowest AM bands were as likely to be 1:3 as 1:2, differing from European languages where 1:2 ratios (delivering the perceptual experience of a temporally regular beat) are dominant. Accordingly, the amplitude-driven physical acoustic structures important for cortical speech tracking flexibly accommodate both universality and specificity.
No takes yet. Share an insight, caveat, or question.
Daikoku et al. (2025) studied this question.