Home /Research /Robat-e-Beheshti: A Persian Wake Word Detection Dataset for Robotic Purposes
OTHER

Robat-e-Beheshti: A Persian Wake Word Detection Dataset for Robotic Purposes

Parisa Ahmadzadeh Raji, Yasser Shekofteh

Year
2022
Citations
2

Abstract

In this paper a dataset is explained which was collected for a wake word detection project which works on classifying audio data. The data are audio files in Persian language. The number of 5738 audio files with different formats and a maximum length of 3 seconds were collected. Afterwards the format of audio files and their sample rates were changed to a specific one. There are positive and negative examples for this dataset. Different audio recorders were used for recording the audios and most of the data were gathered from 187 individuals and the other files were collected from the open source ShEMO dataset: a large-scale validated database for Persian speech emotion detection. In this paper the process of collecting data and how they were organized in different files, are explained and the data analysis is also presented. This dataset will be free for academic usage with official request to the authors.

Keywords

Computer sciencePersianWord (group theory)Sample (material)Speech recognitionNatural language processingArtificial intelligenceDatabase

Related papers

Browse all OTHER papers