Normal-to-Lombard adaptation of speech synthesis using long short-term memory recurrent neural networks

Bajibabu Bollepalli,Lauri Juvela,Manu Airaksinen,Cassia Valentini-Botinhao,Paavo Alku

doi:10.1016/j.specom.2019.04.008

Bajibabu Bollepalli, Lauri Juvela + Show 3 more

Open Access

https://doi.org/10.1016/j.specom.2019.04.008

Copy DOI

Journal: Speech Communication	Publication Date: Apr 18, 2019
Citations: 10	License type: other-oa

Affiliation: Aalto University, University of Edinburgh

Abstract

In this article, three adaptation methods are compared based on how well they change the speaking style of a neural network based text-to-speech (TTS) voice. The speaking style conversion adopted here is from normal to Lombard speech. The selected adaptation methods are: auxiliary features (AF), learning hidden unit contribution (LHUC), and fine-tuning (FT). Furthermore, four state-of-the-art TTS vocoders are compared in the same context. The evaluated vocoders are: GlottHMM, GlottDNN, STRAIGHT, and pulse model in log-domain (PML). Objective and subjective evaluations were conducted to study the performance of both the adaptation methods and the vocoders. In the subjective evaluations, speaking style similarity and speech intelligibility were assessed. In addition to acoustic model adaptation, phoneme durations were also adapted from normal to Lombard with the FT adaptation method. In objective evaluations and speaking style similarity tests, we found that the FT method outperformed the other two adaptation methods. In speech intelligibility tests, we found that there were no significant differences between vocoders although the PML vocoder showed slightly better performance compared to the three other vocoders.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Normal-to-Lombard adaptation of speech synthesis using long short-term memory recurrent neural networks

Abstract

Talk to us

Similar Papers

More From: Speech Communication

Lead the way for us

Similar Papers

Lombard speech synthesis using long short-term memory recurrent neural networks
Bajibabu Bollepalli ... Paavo Alku
-
Bajibabu Bollepalli, et. al.Bajibabu Bollepalli ... Paavo Alku
01 Mar 2017
01 Mar 2017

Correlation between subjective and objective evaluation of peri‐implant soft tissue color
Gianluca Paniz ... Eugenio Romeo
Clinical Oral Implants Research | VOL. 25
Gianluca Paniz, et. al.Gianluca Paniz ... Eugenio Romeo
10 Jun 2013
Clinical Oral Implants Research | VOL. 25

Objective and subjective intelligibility evaluations of noise-reduction algorithms in Mandarin
Junfeng Li ... Yonghong Yan
The Journal of the Acoustical Society of America | VOL. 131
Junfeng Li, et. al.Junfeng Li ... Yonghong Yan
01 Apr 2012
The Journal of the Acoustical Society of America | VOL. 131

Non-Invasive Early Detection of Oral Cancers Using Fluorescence Visualization with Optical Instruments.
Takamichi Morikawa ... Akira Katakura
Cancers | VOL. 12
Takamichi Morikawa, et. al.Takamichi Morikawa ... Akira Katakura
27 Sep 2020
Cancers | VOL. 12

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Normal-to-Lombard adaptation of speech synthesis using long short-term memory recurrent neural networks

Abstract

Talk to us

Similar Papers

More From: Speech Communication