Speech Driven Talking Face Generation From a Single Image and an Emotion Condition.

Sefik Emre Eskimez,You Zhang,Zhiyao Duan

IEEE Transactions on Multimedia（2022）

引用 33|浏览144

暂无评分

摘要

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an end-to-end talking face generation system that takes a speech utterance, a single face image, and a categorical emotion label as input to render a talking face video in sync with the speech and expressing the condition emotion. Objective evaluation on image quality, audiovisual synchronization, and visual emotion expression shows that the proposed system outperforms a state-of-the-art baseline system. Subjective evaluation of visual emotion expression and video realness also demonstrates the superiority of the proposed system. Furthermore, we conduct a pilot study on human emotion recognition of generated videos with mismatched emotions between the audio and visual modalities, and results show that humans reply on the visual modality more significantly than the audio modality on this task.

查看译文

关键词

Faces,Visualization,Face recognition,Emotion recognition,Synchronization,Speech processing,Lips,Audiovisual,emotion,multimodal,talking face generation

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要