Video-realistic expressive audio-visual speech synthesis for the Greek language

Panagiotis Paraskevas Filntisis,Athanasios Katsamanis,Pirros Tsiakoulis,Petros Maragos

doi:10.1016/j.specom.2017.08.011

Panagiotis Paraskevas Filntisis, Athanasios Katsamanis + Show 2 more

https://doi.org/10.1016/j.specom.2017.08.011

Copy DOI

Abstract

High quality expressive speech synthesis has been a long-standing goal towards natural human-computer interaction. Generating a talking head which is both realistic and expressive appears to be a considerable challenge, due to both the high complexity in the acoustic and visual streams and the large non-discrete number of emotional states we would like the talking head to be able to express. In order to cover all the desired emotions, a significant amount of data is required, which poses an additional time-consuming data collection challenge. In this paper we attempt to address the aforementioned problems in an audio-visual context. Towards this goal, we propose two deep neural network (DNN) architectures for Video-realistic Expressive Audio-Visual Text-To-Speech synthesis (EAVTTS) and evaluate them by comparing them directly both to traditional hidden Markov model (HMM) based EAVTTS, as well as a concatenative unit selection EAVTTS approach, both on the realism and the expressiveness of the generated talking head. Next, we investigate adaptation and interpolation techniques to address the problem of covering the large emotional space. We use HMM interpolation in order to generate different levels of intensity for an emotion, as well as investigate whether it is possible to generate speech with intermediate speaking styles between two emotions. In addition, we employ HMM adaptation to adapt an HMM-based system to another emotion using only a limited amount of adaptation data from the target emotion. We performed an extensive experimental evaluation on a medium sized audio-visual corpus covering three emotions, namely anger, sadness and happiness, as well as neutral reading style. Our results show that DNN-based models outperform HMMs and unit selection on both the realism and expressiveness of the generated talking heads, while in terms of adaptation we can successfully adapt an audio-visual HMM set trained on a neutral speaking style database to a target emotion. Finally, we show that HMM interpolation can indeed generate different levels of intensity for EAVTTS by interpolating an emotion with the neutral reading style, as well as in some cases, generate audio-visual speech with intermediate expressions between two emotions.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Video-realistic expressive audio-visual speech synthesis for the Greek language

Abstract

Talk to us

Similar Papers

More From: Speech Communication

Lead the way for us

Journal: Speech Communication	Publication Date: Aug 31, 2017
Citations: 9

Similar Papers

Photorealistic adaptation and interpolation of facial expressions using HMMS and AAMS for audio-visual speech synthesis
Panagiotis P Filntisis ... Petros Maragos
-
Panagiotis P Filntisis, et. al.Panagiotis P Filntisis ... Petros Maragos
01 Sep 2017
01 Sep 2017

Demonstration of an HMM-based photorealistic expressive audio-visual speech synthesis system
Panagiotis Paraskevas Filntisis ... Athanasios Katsamanis
-
Panagiotis Paraskevas Filntisis, et. al.Panagiotis Paraskevas Filntisis ... Athanasios Katsamanis
01 Sep 2017
01 Sep 2017

Dialogue act based expressive speech synthesis in limited domain for the Czech language
Martin Grůber ... Jindřich Matoušek
Informatica | VOL. 44
Martin Grůber, et. al.Martin Grůber ... Jindřich Matoušek
15 Jun 2020
Informatica | VOL. 44

Factored Maximum Penalized Likelihood Kernel Regression for HMM-Based Style-Adaptive Speech Synthesis
June Sig Sung ... Doo Hwa Hong
IEEE Journal of Selected Topics in Signal Processing | VOL. 8
June Sig Sung, et. al.June Sig Sung ... Doo Hwa Hong
01 Apr 2014
IEEE Journal of Selected Topics in Signal Processing | VOL. 8

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Video-realistic expressive audio-visual speech synthesis for the Greek language

Abstract

Talk to us

Similar Papers

More From: Speech Communication