3D Convolutional Neural Network for Speech Emotion Recognition With Its Realization on Intel CPU and NVIDIA GPU.

Mohammad Reza Falahzadeh,Edris Zaman Farsa,Ali Harimi,Arash Ahmadi,Ajith Abraham

IEEE Access（2022）

引用 3|浏览7

暂无评分

摘要

Due to the high level of precision and remarkable capabilities to solve the intricate problems in industry and academia, convolutional neural networks (CNNs) are presented. Speech emotion recognition is an interesting application for CNNs in the field of audio processing. In this paper, a speech emotion recognition system based on a 3D CNN is suggested to analyze and classify the emotions. In the proposed method, the three-dimensional reconstructed phase spaces of the speech signals were calculated. Then, emotion-related patterns formed in these spaces were converted into 3D tensors. Accordingly, a 3D CNN for speech emotion recognition applied to two datasets, EMO-DB and eNTERFACE05, using a speaker-independent technique achieved 90.40% and 82.20% accuracy, respectively. By employing gender recognition, the accuracy rates on EMO-DB increased to 94.42% and on eNTERFACE05 rose to 88.47%. Realization of the introduced 3D CNN on both Intel CPU and NVIDIA GPU is also explored. The results of the implemented 3D CNN without and with regard to gender recognition show that GPU-based running is faster for the EMO-DB and eNTERFACE05 datasets than CPU-based executions (using Python).

查看译文

关键词

Three-dimensional displays,Speech recognition,Mutual information,Emotion recognition,Tensors,Image reconstruction,Feature extraction,3D convolutional neural networks (3D CNNs),speech emotion recognition,reconstructed phase space,3D tensor

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要