Affective Voice Recognition of Older Adults1
Alexander Hong, Yuma Tsuboi, Goldie Nejat, B. Benhabib
- Year
- 2016
- Citations
- 4
Abstract
Older adults (>75 years old) may suffer from social isolation, social inactivity, or loneliness due to physical and cognitive disabilities as well as lifestyle adjustments resulting from old age [1,2]. Socially assistive robots can be used as an effective technology for the elderly to provide social interaction and cognitive assistance with activities of daily living. For example, they can support older adults with self-maintenance tasks (e.g., eating, grooming, and dressing), recreational activities (e.g., playing music and games), etc.In order to promote natural and social human–robot interaction (HRI), and provide the elderly with suitable assistance, robots would need to be equipped with emotional intelligence. For example, they would need to have the ability to consider and respond to the emotions, moods, or affect of the person with whom they are interacting [2].Older adults, including those with dementia, communicate their affective states using facial expressions, body language, and vocal intonation [3]. Our research focuses on the implementation and testing of emotion-based bidirectional interactions, to provide social and cognitive stimulation to older adults, via the intelligent socially assistive robot Brian 2.1 (Fig. 1).Our previous work with Brian 2.1 has focused on the detection of facial expressions [4] and body language [5] of the user. In this paper, we focus on the recognition and identification of affective vocal intonation of older adults as an input to determine Brian's corresponding assistive behaviors. For example, we present the development of an architecture to automatically recognize and classify affective states of older adults from vocal intonation.It has been shown that classifying affective states through voice is challenging, particularly for person-independent recognition and, furthermore, that recognition rates for older adults are lower compared to younger age groups [6]. The aging process directly affects the quality of the voice, as well as its production as a result of various physiological and anatomical changes on the vocal system [7]. For example, a valence detector was investigated in Ref. [8] using elderly voices. However, overall, with respect to automated recognition and classification of affect encompassing states of both arousal and valence during HRI scenarios, current research has not targeted the elderly population [9].Herein, we investigate the recognition and classification of the following combination of positive, neutral, and negative affective states: happy, sadness, anger, and neutral. Happiness is important to detect as for older adults it can indicate well-being, health, and longevity [10]. Sadness and anger are important to detect as they can be the signs of depression as a result of aging, for example, they are often observed in people suffering from dementia [11]. Neutral, which represents an experience of little or no noticeable feelings, is also useful to detect as a baseline for comparing other affective states.Our proposed automated vocal affect detection and classification architecture consists of three main modules: voice recognition, affect feature extraction (AFE), and affect classification (AC, Fig. 2).The VR module is responsible for capturing the audio signal of the elderly speaker and processing it into a file to be used by the AFE module in order to extract voice features from the signal (in our case, a 16-bit 11,025 Hz.wav file). This process was automated for real-time analysis by the robot. Each audio clip is 2–3 s in duration.The AFE module determines the vocal features used to classify the affective states of the elderly. In our work, we utilized the QA5 SDK Version 5.5 software by Nemesysco to identify these features. The.wav files are analyzed based on signal features such as thorns (which are local extrema in amplitude found in the second voice sample in three consecutive voice samples in a clip) and plateaus (local flatness in the voice in the c
Keywords
Related papers
The Organization of Behavior
D. O. Hebb
2005
The spread of true and false news online
Soroush Vosoughi, Deb Roy, Sinan Aral
2018
Fractional Brownian Motions, Fractional Noises and Applications
Benoît B. Mandelbrot, John W. Van Ness
1968
Review of deep learning: concepts, CNN architectures, challenges, applications, future directions
Laith Alzubaidi, Jinglan Zhang, Amjad J. Humaidi +7 more
2021