Presentation Information
[A-11-05]Construction and Preliminary Analysis of an Audio Dataset for Special Fraud Prevention
◎Kazuha Miyazono1, Takahiro Baba1 (1. Kurume Institute of Technology)
Keywords:
Special Fraud,Speech Dataset,Automatic Speech Recognition,Whisper,Text Classification
Special fraud has become a serious social issue in Japan, increasing the need for automatic detection technologies based on telephone conversations. However, publicly available Japanese speech datasets for special fraud research remain limited. In this study, we collected special fraud audio recordings published by prefectural police departments and constructed a Japanese special fraud speech dataset. Audio data were transcribed using Whisper, and fraud-type labels were assigned to each sample. We then conducted frequency analysis and TF-IDF-based keyword extraction, revealing distinctive vocabulary characteristics for different fraud categories. In addition, a binary classification experiment targeting refund fraud and impersonation fraud achieved an accuracy of 0.917 using TF-IDF features and a Support Vector Machine (SVM). Future work includes expanding the dataset and developing dialogue generation methods for fraud-prevention training systems.
