首页 /研究 /Computational Auditory Scene Analysis and its Application to Robot Audition
HRI

Computational Auditory Scene Analysis and its Application to Robot Audition

Hiroshi G. Okuno, Kazuhiro Nakadai

发表年份
2008
引用次数
18

摘要

Robot capability of hearing sounds, in particular, a mixture of sounds, by its own microphones, that is, robot audition, is important in improving human robot interaction. This paper presents the robot audition open-source software, called "HARK" (HRI-JP Audition for Robots with Kyoto University), which consists of primitive functions in computational auditory scene analysis; sound source localization, separation, and recognition of separated sounds. Since separated sounds suffer from spectral distortion due to separation, the HARK generates a time-spectral map of reliability, called "missing feature mask", for features of separated sounds. Then separated sounds are recognized by the missing-feature theory (MFT) based ASR with missing feature masks. The HARK is implemented on the middleware called "FlowDesigner" to share intermediate audio data, which enables near real-time processing.

关键词

Computer scienceRobotFeature (linguistics)Artificial intelligenceComputer visionSpeech recognitionDistortion (music)Source separation

相关论文

查看 HRI 分类全部论文