Papers

4

Total Citations

135

H-Index

3

About

Masato Akagi is a leading researcher in speech emotion recognition and auditory perception, whose work bridges the gap between human hearing and machine understanding. His primary research areas include dimensional emotion recognition from speech, auditory front-end processing, and binaural signal processing for human-robot interaction. Akagi’s most impactful contribution is the development of a speech emotion recognition system using 3D convolutions and attention-based sliding recurrent networks, which has garnered 83 citations. This work models how the human auditory system tracks emotional dynamics through intensity and fundamental frequency, enabling robots to interpret speaker intentions more naturally. He further advanced the field with multi-resolution modulation-filtered cochleagram features for LSTM-based emotion recognition (44 citations), improving the tracking of temporal emotional shifts. Akagi also explored modulation spectral features and recurrent neural networks for dimensional emotion recognition, providing frame-level feature sequences essential for dynamic interaction. Beyond emotion, his work on direction-of-arrival (DOA) estimation using equalization-cancellation theory contributes to binaural speech enhancement for auditory humanoid robots. With a career spanning foundational auditory theory and cutting-edge deep learning, Akagi’s research is pivotal for creating emotionally aware, perceptually intelligent machines.

Research Focus

Key Achievements

3
H-Index
4
Papers
135
Total Citations
34
Avg Citations/Paper
🏆 Most Cited Paper
Speech Emotion Recognition Using 3D Convolutions and Attention-Based Sliding Recurrent Networks With Auditory Front-Ends
83 citations · 2020
📈 Most Prolific Year: 2020 (1 Papers)
🤝 Key Collaborators: 7
🏛 Institutions: Japan Advanced Institute of Science and Technology

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago