Deep Learning-Based Sound Classification Algorithms for Enhanced Service Robots Audio Capabilities
Lorena Muscar, Toma Telembici, Corneliu Rusu
- Year
- 2024
- Citations
- 3
Abstract
This paper presents the way in which different audio signals are classified using at least 24 Mel Frequency Cepstral Coefficients (MFCCs) features extracted. Our audio dataset for service robots’ application in Romanian language which consists of 148 classes with 30 audio signals/class was used for analysis. We examined the impact of various feature extraction numbers on classification accuracy. Four neural networks classifier, Convolutional Neural Networks (CNN), Artificial Neural Network (ANN), Long-Short Term Memory – Recurrent Neural Network (LSTM-RNN) and Support Vector Machine (SVM) are utilized to evaluate the effectiveness of each feature set in recognizing each audio signal. The features that perform the best in the field of audio signal recognition were identified by comparing the correct classification rates. Data was split in 90% for the training part and 10% for the testing part. In order to validate the results, we used 10-fold cross validation repeated 10 times for ANN and SVM classifiers and 10-fold cross validation for LSTM-RNN and CNN. We showed that by using CNN and Mel spectrograms as input we reach an overall correct classification rate of 99.77% in the testing phase. This study makes contributions to the field of audio recognition from audio signals, facilitating advancements in service robots audio capabilities.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002