Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid Robot
Tatsuya Kawahara, Koji Inoue, Divesh Lala, Katsuya Takanashi
- Year
- 2018
- Citations
- 2
Abstract
This paper addresses audio-visual signal processing for conversation analysis, which involves multi-modal behavior detection and mental-state recognition. We have investigated prediction of turn-taking by the audience in a poster session from their multi-modal behaviors, and found out that the eye-gaze provides an important cue compared with head nodding and verbal backchannels. This finding has been applied to audio-visual speaker diarization by combining eye-gaze information. We are now investigating engagement recognition in human-robot interaction based on the same scheme. Robust and realtime detection of laughing, backchannels and nodding is realized based on LSTM-CTC. We introduce a latent “character” model to cope with the subjectivity and variations of engagement annotations. Experimental evaluations demonstrate that (1) the latent character model is effective, (2) automatic behavior detection is robust and does not degrade the engagement recognition accuracy, and (3) the eye-gaze is the most important feature among others.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002