MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation

Gwantae Kim,Insung Ham,Hanseok Ko,Seonghyeok Noh

doi:10.1109/icassp49357.2023.10095344

Abstract

When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and generalized encodings, we propose a novel framework with a multimodal pre-trained encoder for co-speech gesture generation. In the proposed method, the multi-head-attention-based encoder is trained with self-supervised learning to contain the information on each modality. Moreover, we collect full-body gestures that consist of 3D joint rotations to improve visualization and apply gestures to the extensible body model. Through the series of experiments and human evaluation, the proposed method renders realistic co-speech gestures not only when all input modalities are given but also when the input modalities are missing or noisy. The project page is available here <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.
Jing Liu ... Jinhui Tang
IEEE transactions on pattern analysis and machine intelligence | VOL. PP
Jing Liu, et. al.Jing Liu ... Jinhui Tang
01 Jan 2024
IEEE transactions on pattern analysis and machine intelligence | VOL. PP

Salient Co-Speech Gesture Synthesizing with Discrete Motion Representation
Zijie Ye ... Junliang Xing
-
Zijie Ye, et. al.Zijie Ye ... Junliang Xing
04 Jun 2023
04 Jun 2023

Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation
Xian Liu ... Bolei Zhou
-
Xian Liu, et. al.Xian Liu ... Bolei Zhou
01 Jun 2022
01 Jun 2022

Gesticulator: A framework for semantically-aware speech-driven gesture generation
Taras Kucherenko ... Patrik Jonell
-
Taras Kucherenko, et. al.Taras Kucherenko ... Patrik Jonell
21 Oct 2020
21 Oct 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation

Abstract

Talk to us

Similar Papers