This discourse analysis uncovers phonological patterns in various languages, suggesting a standardized annotation approach.
The paper presents a corpus-driven methodology for discourse analysis of connected speech phenomena in typologically diverse languages (Russian, English, Chinese, Evenki). The study focuses on developing a multilingual speech corpus with a unified annotation system that can be used for comparative analysis of non-canonical phonological patterns across discourse types. The material comprises speech databases including English (news, academic, and regional varieties), Chinese (spontaneous speech, commercial and social advertisement), Evenki (INEL and Amur region corpora), and Russian (educational discourse). Within corpus-driven approach, the following methods and tools were used: automatic alignment tools (Montreal Forced Aligner, BAS WebServices), manual expert validation, file format conversion (XML, EXMARaLDA), and Python scripting (for data processing). As a result, a standardized corpus annotation system has been developed to compare natural phonetic modifications across languages. The research demonstrated the effectiveness of automated processing tools at the same time emphasizing the necessity of manual expert correction. Query templates for the EXAKT corpus manager have been designed to investigate modification frequency and contextual patterns. Future research directions include corpus expansion, development of machine learning algorithms for automatic modification detection.
No takes yet. Share an insight, caveat, or question.
Veronika G. Karavaeva (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: