Home /Research /Enhanced Robot Speech Recognition Based on Microphone Array Source Separation and Missing Feature Theory
OTHER

Enhanced Robot Speech Recognition Based on Microphone Array Source Separation and Missing Feature Theory

S. Yamamoto, Jean-Marc Valin, Kazuhiro Nakadai, Jean Rouat, François Michaud, Tetsuya Ogata, Hiroshi G. Okuno

Year
2006
Citations
69

Abstract

A humanoid robot under real-world environments usually hears mixtures of sounds, and thus three capabilities are essential for robot audition; sound source localization, separation, and recognition of separated sounds. While the first two are frequently addressed, the last one has not been studied so much. We present a system that gives a humanoid robot the ability to localize, separate and recognize simultaneous sound sources. A microphone array is used along with a real-time dedicated implementation of Geometric Source Separation (GSS) and a multi-channel post-filter that gives us a further reduction of interferences from other sources. An automatic speech recognizer (ASR) based on the Missing Feature Theory (MFT) recognizes separated sounds in real-time by generating missing feature masks automatically from the post-filtering step. The main advantage of this approach for humanoid robots resides in the fact that the ASR with a clean acoustic model can adapt the distortion of separated sound by consulting the post-filter feature masks. Recognition rates are presented for three simultaneous speakers located at 2m from the robot. Use of both the post-filter and the missing feature mask results in an average reduction in error rate of 42% (relative).

Keywords

Computer scienceMicrophone arrayFeature (linguistics)MicrophoneFilter (signal processing)Humanoid robotRobotSpeech recognitionArtificial intelligenceDistortion (music)

Related papers

Browse all OTHER papers