Papers
4
Total Citations
135
H-Index
3
About
Masato Akagi is a leading researcher in speech emotion recognition and auditory perception, whose work bridges the gap between human hearing and machine understanding. His primary research areas include dimensional emotion recognition from speech, auditory front-end processing, and binaural signal processing for human-robot interaction. Akagi’s most impactful contribution is the development of a speech emotion recognition system using 3D convolutions and attention-based sliding recurrent networks, which has garnered 83 citations. This work models how the human auditory system tracks emotional dynamics through intensity and fundamental frequency, enabling robots to interpret speaker intentions more naturally. He further advanced the field with multi-resolution modulation-filtered cochleagram features for LSTM-based emotion recognition (44 citations), improving the tracking of temporal emotional shifts. Akagi also explored modulation spectral features and recurrent neural networks for dimensional emotion recognition, providing frame-level feature sequences essential for dynamic interaction. Beyond emotion, his work on direction-of-arrival (DOA) estimation using equalization-cancellation theory contributes to binaural speech enhancement for auditory humanoid robots. With a career spanning foundational auditory theory and cutting-edge deep learning, Akagi’s research is pivotal for creating emotionally aware, perceptually intelligent machines.
Research Focus
Key Achievements
Top Papers
- 1
- 2
- 3
- 4A DOA estimation algorithm based on equalization-cancellation theory3 citations · 2010