Using Mandarin Training Corpus to Realize a Mandarin-Tibetan Cross-Lingual Emotional Speech Synthesis

Peiwen Wu,Hongwu Yang,Zhenye Gan

doi:10.1007/978-981-10-8111-8_11

Peiwen Wu, Hongwu Yang + Show 1 more

https://doi.org/10.1007/978-981-10-8111-8_11

Copy DOI

Export

Save

Cite

Publication Date: Jan 1, 2018

Affiliation: Northwest Normal University

Abstract
Full-Text
Similar Papers

Abstract

Listen

This paper presents a hidden Markov model (HMM)-based Mandarin-Tibetan cross-lingual emotional speech synthesis by using an emotional Mandarin speech corpus with speaker adaptation. We firstly train a set of average acoustic models by speaker adaptive training with a one-speaker neutral Tibetan corpus and a multi-speaker neutral Mandarin corpus. Then we train a set of speaker dependent acoustic models of target emotion, which are used to synthesize emotional Tibetan or Mandarin speech, by speaker adaptation with the target emotional Mandarin corpus. Subjective evaluations and objective tests show that the method can synthesize both emotional Mandarin speech and emotional Tibetan speech with high naturalness and emotional similarity. Therefore, the method can be adopted to realizing an emotional speech synthesis with exiting emotional training corpus for languages lacking emotional speech resources.

Full Text