Home /Research /Robot-directed speech detection using Multimodal Semantic Confidence based on speech, image, and motion
MANIPULATION

Robot-directed speech detection using Multimodal Semantic Confidence based on speech, image, and motion

Xiang Zuo, Naoto Iwahashi, Ryo Taguchi, Shigeki Matsuda, Komei Sugiura, Kotaro Funakoshi, Mikio Nakano, Natsuki Oka

Year
2010
Citations
10

Abstract

In this paper, we propose a novel method to detect robot-directed (RD) speech that adopts the Multimodal Semantic Confidence (MSC) measure. The MSC measure is used to decide whether the speech can be interpreted as a feasible action under the current physical situation in an object manipulation task. This measure is calculated by integrating speech, image, and motion confidence measures with weightings that are optimized by logistic regression. Experimental results show that, compared with a baseline method that uses speech confidence only, MSC achieved an absolute increase of 5% for clean speech and 12% for noisy speech in terms of average maximum F-measure.

Keywords

Computer scienceMeasure (data warehouse)Speech recognitionTask (project management)Object (grammar)Motion (physics)Artificial intelligenceBaseline (sea)RobotSpeech processing

Related papers

Browse all MANIPULATION papers