Synthesized Speech Attribution Using The Patchout Spectrogram Attribution Transformer.

Kratika Bhagtani,Emily R. Bartusiak,Amit Kumar Singh Yadav,Paolo Bestagini,Edward J. Delp

IH&MMSec（2023）

引用 0|浏览13

暂无评分

摘要

The malicious use of synthetic speech has increased with the recent availability of speech generation tools. It is important to determine whether a speech signal is authentic (spoken by a human) or is synthesized and to determine the generation method used to create it. Identifying the synthesis method is known as synthetic speech attribution. In this paper, we propose the use of a transformer deep learning method that analyzes mel-spectrograms for synthetic speech attribution. Our method known as Patchout Spectrogram Attribution Transformer (PSAT) can distinguish new, unseen speech generation methods from those seen during training. PSAT demonstrates high performance in attributing synthetic speech signals. Evaluation on the DARPA SemaFor Audio Attribution Dataset and the ASVSpoof2019 Dataset shows that our method achieves more than 95% accuracy in synthetic speech attribution and performs better than existing deep learning approaches.

查看译文

关键词

deep learning, audio forensics, synthetic speech, transformers, melspectrograms

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要