Assessing phoneme distribution for speech modeling

Jesus A. Parra, Carlos Calvache,Matias Zanartu

18th International Symposium on Medical Information Processing and Analysis（2023）

引用 0|浏览1

暂无评分

摘要

Phonetically balanced texts are used to study different voice and speech characteristics. In the context of clinical work and research, these texts provide a standard for quantifying perceptual, acoustic, or aerodynamic assessments. Recent modeling efforts are being devoted to describing long-term speech behaviors based on a collection of sustained phonemes. However, comprehensive descriptions of phoneme distributions representative of connected speech are not readily available. Thus, the present study introduces a method to estimate phoneme distributions using text data mining, as an alternative to existing power law methods. The procedure used for the decomposition of texts into phonemes, the estimation of the phonetic distributions and the comparisons between different texts, conversational speech, and standard reading passages are discussed. The results are presented using histograms and R-squared determination coefficients for the case of the English language, although the approach can be easily applied for other languages. A discussion of the proposed method, results, and limitations is presented.

查看译文

关键词

Phonetically balanced, voice, speech, histograms

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要