Unit selection speech synthesis using multiple speech units at non-adjacent segments for prosody and waveform generation

Masatsune Tamura,Masami Akamine,Takehiko Kagoshima,Norbert Braunschweiler

doi:10.1109/icassp.2010.5495151

Abstract

In this paper, we propose a speech synthesis method that combines a natural waveform concatenation based speech synthesis method and our baseline plural unit selection and fusion method. Two main features of the proposed method are (i) prosody regeneration from selected speech units and (ii) using multiple speech units at non-adjacent segments. The nonadjacent segments is the segment that the previous or following speech units in the optimum speech unit sequence are not adjacent in the database. By using the prosody of selected speech units, the original prosodic expressions and sounds of recorded speech are retained, while discontinuities are reduced by using multiple speech units at non-adjacent segments. MOS evaluations showed that the proposed method provides a clear improvement against the conventional unit selection method and our baseline method.

Full Text