首页 /研究 /Modelling emotional valence and arousal of non-linguistic utterances for sound design support
OTHER

Modelling emotional valence and arousal of non-linguistic utterances for sound design support

Ahmed Khota, Eric W. Cooper, Yu Yan, Mate Kovacs

发表年份
2022
引用次数
2
访问权限
开放获取

摘要

Non-Linguistic Utterances (NLUs), produced for popular media, computers, robots, and public spaces, can quickly and wordlessly convey emotional characteristics of a message. They have been studied in terms of their ability to convey affect in robot communication. The objective of this research is to develop a model that correctly infers the emotional Valence and Arousal of an NLU. On a Likert scale, 17 subjects evaluated the relative Valence and Arousal of 560 sounds collected from popular movies, TV shows, and video games, including NLUs and other character utterances. Three audio feature sets were used to extract features including spectral energy, spectral spread, zero-crossing rate (ZCR), Mel Frequency Cepstral Coefficients (MFCCs), and audio chroma, as well as pitch, jitter, formant, shimmer, loudness, and Harmonics-to-Noise Ratio, among others. After feature reduction by Factor Analysis, the best-performing models inferred average Valence with a Mean Absolute Error (MAE) of 0.107 and Arousal with MAE of 0.097 on audio samples removed from the training stages. These results suggest the model infers Valence and Arousal of most NLUs to less than the difference between successive rating points on the 7-point Likert scale (0.14). This inference system is applicable to the development of novel NLUs to augment robot-human communication or to the design of sounds for other systems, machines, and settings.

关键词

Speech recognitionComputer scienceValence (chemistry)FormantLoudnessArousalKeyword spottingMel-frequency cepstrumArtificial intelligenceFeature extraction

相关论文

查看 OTHER 分类全部论文