Speech gesture generation from the trimodal context of text, audio, and speaker identity

Youngwoo Yoon,Minsu Jang,Jaehong Kim,Joo-Haeng Lee,Jaeyeon Lee,Bok Cha,Geehyuk Lee

doi:10.1145/3414685.3417838

Abstract

For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human-agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it is difficult to generate human-like gestures due to the lack of understanding of how people gesture. Data-driven approaches attempt to learn gesticulation skills from human demonstrations, but the ambiguous and individual nature of gestures hinders learning. In this paper, we present an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures. By incorporating a multimodal context and an adversarial training scheme, the proposed model outputs gestures that are human-like and that match with speech content and rhythm. We also introduce a new quantitative evaluation metric for gesture generation models. Experiments with the introduced metric and subjective human evaluation showed that the proposed gesture generation model is better than existing end-to-end generation models. We further confirm that our model is able to work with synthesized audio in a scenario where contexts are constrained, and show that different gesture styles can be generated for the same speech by specifying different speaker identities in the style embedding space that is learned from videos of various speakers. All the code and data is available at https://github.com/ai4r/Gesture-Generation-from-Trimodal-Context.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Speech gesture generation from the trimodal context of text, audio, and speaker identity

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Graphics

Lead the way for us

Journal: ACM Transactions on Graphics	Publication Date: Nov 27, 2020
Citations: 172

Similar Papers

Embodied Gesture Processing: Motor-Based Integration of Perception and Action in Social Artificial Agents
Amir Sadeghipour ... Stefan Kopp
Cognitive Computation | VOL. 3
Amir Sadeghipour, et. al.Amir Sadeghipour ... Stefan Kopp
11 Nov 2010
Cognitive Computation | VOL. 3

Synchronous Colored Petri Net Based Modeling and Video Analysis of Conversational Head-Gestures for Training Social Robots
Aditi Singh ... Cheng Chang Lu
-
Aditi Singh, et. al.Aditi Singh ... Cheng Chang Lu
04 Nov 2021
04 Nov 2021

Proceedings of the 8th International Conference on Human-Agent Interaction
...
-
, et. al. ...
10 Nov 2020
10 Nov 2020

Pattern analysis of EEG responses to speech and voice: Influence of feature grouping
Lars Hausfeld ... Milene Bonte
NeuroImage | VOL. 59
Lars Hausfeld, et. al.Lars Hausfeld ... Milene Bonte
30 Nov 2011
NeuroImage | VOL. 59

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Speech gesture generation from the trimodal context of text, audio, and speaker identity

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Graphics