Learning HMM State Sequences from Phonemes for Speech Synthesis

Giorgio Biagetti,Paolo Crippa,Laura Falaschetti,Simone Orcioni,Claudio Turchetti

doi:10.1016/j.procs.2016.08.206

Giorgio Biagetti, Paolo Crippa + Show 3 more

Open Access

https://doi.org/10.1016/j.procs.2016.08.206

Copy DOI

Journal: Procedia Computer Science	Publication Date: Jan 1, 2016
Citations: 3	License type: cc-by-nc-nd

Affiliation: Engineering (Italy)

Abstract

This paper presents a technique for learning hidden Markov model (HMM) state sequences from phonemes, that combined with modified discrete cosine transform (MDCT), is useful for speech synthesis. Mel-cepstral spectral parameters, currently adopted in the conventional methods as features for HMM acoustic modeling, do not ensure direct speech waveforms reconstruction. In contrast to these approaches, we use an analysis/synthesis technique based on MDCT that guarantees a perfect reconstruction of the signal frame feature vectors and allows for a 50% overlap between frames without increasing the data rate. Experimental results show that the spectrograms achieved with the suggested technique behave very closely to the original spectrograms, and the quality of synthesized speech is conveniently evaluated using the well known Itakura-Saito measure.

Full Text