Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Objective evaluation of synthesized speech is critical for advancing speech generation systems, yet existing metrics for intelligibility and prosody remain limited in scope and weakly correlated with human perception. Word Error Rate (WER) provides only a coarse text-based measure of intelligibility, while F0-RMSE and related pitch-based metrics offer a narrow, reference-dependent view of prosody. To address these limitations, we propose <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">TTScore</i>, a targeted and reference-free evaluation framework based on conditional prediction of discrete speech tokens. TTScore employs two sequence-to-sequence predictors conditioned on input text: <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">TTScore-int</i>, which measures intelligibility through content tokens, and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">TTScore-pro</i>, which evaluates prosody from the perspective of pitch, through prosody tokens. For each synthesized utterance, the predictors compute the likelihood of the corresponding token sequences, yielding interpretable scores that capture alignment with intended linguistic content and prosodic structure. Experiments on the SOMOS, VoiceMOS, and TTSArena benchmarks demonstrate that TTScore-int and TTScore-pro provide reliable, aspect-specific evaluation and achieve stronger correlations with human judgments of overall quality than existing intelligibility and prosody-focused metrics.

Similar Papers
  • Conference Article
  • Cite Count Icon 69
  • 10.1145/3394486.3403331
LRSpeech
  • Aug 20, 2020
  • Jin Xu + 6 more

Speech synthesis (text to speech, TTS) and recognition (automatic speech recognition, ASR) are important speech tasks, and require a large amount of text and speech pairs for model training. However, there are more than 6,000 languages in the world and most languages are lack of speech training data, which poses significant challenges when building TTS and ASR systems for extremely low-resource languages. In this paper, we develop LRSpeech, a TTS and ASR system under the extremely low-resource setting, which can support rare languages with low data cost. LRSpeech consists of three key techniques: 1) pre-training on rich-resource languages and fine-tuning on low-resource languages; 2) dual transformation between TTS and ASR to iteratively boost the accuracy of each other; 3) knowledge distillation to customize the TTS model on a high-quality target-speaker voice and improve the ASR model on multiple voices. We conduct experiments on an experimental language (English) and a truly low-resource language (Lithuanian) to verify the effectiveness of LRSpeech. Experimental results show that LRSpeech 1) achieves high quality for TTS in terms of both intelligibility (more than $98%$ intelligibility rate) and naturalness (above 3.5 mean opinion score (MOS)) of the synthesized speech, which satisfy the requirements for industrial deployment, 2) achieves promising recognition accuracy for ASR, and 3) last but not least, uses extremely low-resource training data. We also conduct comprehensive analyses on LRSpeech with different amounts of data resources, and provide valuable insights and guidances for industrial deployment. We are currently deploying LRSpeech into a commercialized cloud speech service to support TTS on more rare languages.

  • Conference Article
  • Cite Count Icon 13
  • 10.21437/interspeech.2010-602
Synthesis of fast speech with interpolation of adapted HSMMs and its evaluation by blind and sighted listeners
  • Sep 26, 2010
  • Michael Pucher + 2 more

In this paper we evaluate a method for generating synthetic speech at high speaking rates based on the interpolation of hidden semi-Markov models (HSMMs) trained on speech data recorded at normal and fast speaking rates. The subjective evaluation was carried out with both blind listeners, who are used to very fast speaking rates, and sighted listeners. We show that we can achieve a better intelligibility rate and higher voice quality with this method compared to standard HSMM-based duration modeling. We also evaluate duration modeling with the interpolation of all the acoustic features including not only duration but also spectral and F0 models. An analysis of the mean squared error (MSE) of standard HSMM-based duration modeling for fast speech identifies problematic linguistic contexts for duration modeling. Index Terms: speech synthesis, fast speech, hidden semiMarkov model

  • Conference Article
  • 10.1109/cvidliccea56201.2022.9824329
Lightweight convolution-based Chinese Speech Synthesis Method
  • May 20, 2022
  • Ruotong Yang + 5 more

Speech synthesis technology is one of the key technologies for human-computer speech interaction and is widely used in various fields such as audiobooks and information broadcasting. In this paper, we propose a speech synthesis model based on lightweight convolution (Lightweight Convolution-Tacotron2, LConv-T) to address the problems of severe long-distance information loss and slow inference in Tacotron2 speech synthesis model using recurrent neural networks. The encoder uses multiple lightweight convolutional modules connected in a densely connected manner to obtain the contextual information of the input text over long distances. To address the shortcomings of the proposed LConv-T model for speech synthesis, this paper further proposes a Tacotron2-based feature fusion speech synthesis model (Dynamic Lightweight Convolution-Tacotron2, DLConv-T), which can improve the stability of speech synthesis by using the Bi-LSTM and dynamic lightweight convolution modules respectively. The text feature extraction effectively improves the speech synthesis effect. The experimental results show that compared with the Tacotron2 model, the LConv-T and DLConv-T models reduce the objective evaluation MCD values by 0.15db and 0.42db, and improve the subjective evaluation MOS by 0.15 and 0.47, respectively.

  • Research Article
  • Cite Count Icon 30
  • 10.1016/j.sigpro.2006.02.039
Design, implementation and evaluation of the Czech realistic audio-visual speech synthesis
  • May 24, 2006
  • Signal Processing
  • Miloš Železný + 3 more

Design, implementation and evaluation of the Czech realistic audio-visual speech synthesis

  • Research Article
  • Cite Count Icon 20
  • 10.1016/j.csl.2024.101747
Refining the evaluation of speech synthesis: A summary of the Blizzard Challenge 2023
  • Nov 8, 2024
  • Computer Speech & Language
  • Olivier Perrotin + 4 more

Refining the evaluation of speech synthesis: A summary of the Blizzard Challenge 2023

  • Single Report
  • Cite Count Icon 32
  • 10.21236/ada252015
Intelligibility and Acceptability Testing for Speech Technology
  • May 22, 1992
  • Astrid Schmidt-Nielsen

: The evaluation of speech intelligibility and acceptability is an important aspect of the use, development, and selection of voice communication devices-telephone systems, digital voice systems, speech synthesis by rule, speech in noise, and the effects of noise stripping. Standard test procedures can provide highly reliable measures of speech intelligibility, and subjective acceptability tests can be used to evaluate voice quality. These tests are often highly correlated with other measures of communication performance and can be used to predict performance in many situations. However, when the speech signal is severely degraded or highly processed. a more complete evaluation of speech quality is needed-one that takes into account the many different sources of information that contribute to how we understand speech.

  • Book Chapter
  • Cite Count Icon 12
  • 10.1007/978-3-319-66429-3_27
Deep Recurrent Neural Networks in Speech Synthesis Using a Continuous Vocoder
  • Jan 1, 2017
  • Mohammed Salah Al-Radhi + 2 more

In our earlier work in statistical parametric speech synthesis, we proposed a vocoder using continuous F0 in combination with Maximum Voiced Frequency (MVF), which was successfully used with a feed-forward deep neural network (DNN). The advantage of a continuous vocoder in this scenario is that vocoder parameters are simpler to model than traditional vocoders with discontinuous F0. However, DNNs have a lack of sequence modeling which might degrade the quality of synthesized speech. In order to avoid this problem, we propose the use of sequence-to-sequence modeling with recurrent neural networks (RNNs). In this paper, four neural network architectures (long short-term memory (LSTM), bidirectional LSTM (BLSTM), gated recurrent network (GRU), and standard RNN) are investigated and applied using this continuous vocoder to model F0, MVF, and Mel-Generalized Cepstrum (MGC) for more natural sounding speech synthesis. Experimental results from objective and subjective evaluations have shown that the proposed framework converges faster and gives state-of-the-art speech synthesis performance while outperforming the conventional feed-forward DNN.

  • Conference Article
  • 10.1109/icsp.2016.7877818
Modeling of fundamental frequency contours for HMM-based speech synthesis: Representation of fundamental frequency contours for statistical speech synthesis
  • Nov 1, 2016
  • Keikichi Hirose

Statistical parametric speech synthesis technologies, such as HMM-based and DNN-based ones, gain special attention from researchers because of their ability in generating speech in various voice qualities and styles. In these methods, all acoustic parameters (except durational ones) are handled in a frame-by-frame manner, which is not appropriate for prosodic features. Although relation of adjacent frames is viewed, it is not enough. Prosodic features are related to words, phrases, sentences, and even to paragraphs, and should be viewed in a wider time span. One possible way to handle the features well in speech synthesis process is to model fundamental frequency (F 0 ) movements and to apply its constraints. Among several models of F 0 contours, the generation process model of F 0 contours is ideal for the purpose, since it can well represent hierarchical structure of prosody as superposition of phrase and accent components, keeping a clear relationship between model commands and linguistic information. A method is developed which decomposes F 0 contours into three layers based on the model, and handles them as different streams in the HMM-based speech synthesis process. Advantage of the method is confirmed through objective and subjective evaluations. Issues of flexible control of prosody are also addressed.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/icct.2006.341940
Speech bandwidth extension method using speech recognition and speech synthesis
  • Nov 1, 2006
  • Masashi Takashina + 3 more

In this paper, we propose two kinds of novel speech bandwidth extension methods based on a speech recognition, a speech synthesis and a speech signal processing technique. In the proposed methods, the lost speech component by bandlimiting is generated with the technique of a speech recognition, a speech synthesis, and a speech signal processing by using the bandlimited speech information. For evaluating both proposed methods, we conducted the subjective and objective evaluation experiments. The experimental results shows the effectiveness of both proposed methods about the extension feeling of bandwidth. However, the evaluation of the speech quality became a not good result. Our problem is an improvement of the speech quality of the proposed method in the future.

  • Research Article
  • Cite Count Icon 3
  • 10.1109/taslp.2016.2537982
Candidate Expansion and Prosody Adjustment for Natural Speech Synthesis Using a Small Corpus
  • Jun 1, 2016
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Yan-You Chen + 4 more

This study proposes a hybrid approach to natural-sounding speech synthesis based on candidate expansion, unit selection, and prosody adjustment using a small corpus. The proposed method is more specific to tonal language, in particular Mandarin. In conventional speech synthesis studies, the quality of synthesized speech depends heavily on the size of the speech corpus. However, it is highly time-consuming and labor-intensive to prepare a large labeled corpus. In this work, candidate expansion is proposed to retrieve potential candidates that are unlikely to be retrieved using only linguistic features. The optimal unit sequence is then obtained from the expanded candidates by using the proposed unit selection mechanism at the phoneme and prosodic word levels. Finally, a prosodic word-level prosody adjustment is proposed to improve the continuity and smoothness of the prosody of the synthesized speech. To evaluate the proposed method, the Tsing-Hua corpus of speech synthesis was adopted. The results of an objective evaluation demonstrate the effectiveness of candidate expansion and the improvement of the continuity and smoothness of the prosody of the synthesized speech. The results of a subjective evaluation also show the proposed system could synthesize the speech with improved quality and naturalness, in particular for a small-sized or resource-limited corpus.

  • Conference Article
  • Cite Count Icon 3
  • 10.21437/interspeech.2016-258
Probabilistic Amplitude Demodulation Features in Speech Synthesis for Improving Prosody
  • Sep 8, 2016
  • Alexandros Lazaridis + 2 more

Amplitude demodulation (AM) is a signal decomposition technique by which a signal can be decomposed to a product of two signals, i.e, a quickly varying carrier and a slowly varying modulator. In this work, the probabilistic amplitude demodulation (PAD) features are used to improve prosody in speech synthesis. The PAD is applied iteratively for generating syllable and stress amplitude modulations in a cascade manner. The PAD features are used as a secondary input scheme along with the standard text-based input features in statistical parametric speech synthesis. Specifically, deep neural network (DNN)-based speech synthesis is used to evaluate the importance of these features. Objective evaluation has shown that the proposed system using the PAD features has improved mainly prosody modelling; it outperforms the baseline system by approximately 5% in terms of relative reduction in root mean square error (RMSE) of the fundamental frequency (F0). The significance of this improvement is validated by subjective evaluation of the overall speech quality, achieving 38.6% over 19.5% preference score in respect to the baseline system, in an ABX test.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 6
  • 10.1186/1687-6180-2015-2
Soft context clustering for F0 modeling in HMM-based speech synthesis
  • Jan 9, 2015
  • EURASIP Journal on Advances in Signal Processing
  • Soheil Khorram + 2 more

This paper proposes the use of a new binary decision tree, which we call a soft decision tree, to improve generalization performance compared to the conventional ‘hard’ decision tree method that is used to cluster context-dependent model parameters in statistical parametric speech synthesis. We apply the method to improve the modeling of fundamental frequency, which is an important factor in synthesizing natural-sounding high-quality speech. Conventionally, hard decision tree-clustered hidden Markov models (HMMs) are used, in which each model parameter is assigned to a single leaf node. However, this ‘divide-and-conquer’ approach leads to data sparsity, with the consequence that it suffers from poor generalization, meaning that it is unable to accurately predict parameters for models of unseen contexts: the hard decision tree is a weak function approximator. To alleviate this, we propose the soft decision tree, which is a binary decision tree with soft decisions at the internal nodes. In this soft clustering method, internal nodes select both their children with certain membership degrees; therefore, each node can be viewed as a fuzzy set with a context-dependent membership function. The soft decision tree improves model generalization and provides a superior function approximator because it is able to assign each context to several overlapped leaves. In order to use such a soft decision tree to predict the parameters of the HMM output probability distribution, we derive the smoothest (maximum entropy) distribution which captures all partial first-order moments and a global second-order moment of the training samples. Employing such a soft decision tree architecture with maximum entropy distributions, a novel speech synthesis system is trained using maximum likelihood (ML) parameter re-estimation and synthesis is achieved via maximum output probability parameter generation. In addition, a soft decision tree construction algorithm optimizing a log-likelihood measure is developed. Both subjective and objective evaluations were conducted and indicate a considerable improvement over the conventional method.

  • Research Article
  • 10.1121/1.427287
Comparative study of F0 extractors for high-quality speech synthesis
  • Oct 1, 1999
  • The Journal of the Acoustical Society of America
  • Hideki Kawahara + 1 more

Performance of a new F0 extraction algorithm based on fixed point analysis of filter center frequency to output instantaneous frequency [Kawahara et al., Eurospeech’99] was compared with numbers of F0 extraction algorithms based on different definitions of fundamental frequency. The proposed method uses partial derivatives of the mapping at fixed points to estimate carrier to noise ratio of F0 information. It also enables integration of distributed F0 cues among harmonic components to provide a reliable F0 estimate. Objective evaluations were conducted using simulations and a speech database with simultaneous EGG (electroglottograph) recording. Subjective evaluations were based on reproduced speech quality assessment by a high-quality speech analysis/modification/synthesis method STRAIGHT [Kawahara et al., Speech Commun. 27, 187–207]. Discussions about the relevance of various F0 definitions for high-quality speech synthesis will be presented based on these test results. It was also indicated that the proposed method is tunable to specific needs.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 5
  • 10.3390/app13095724
Semi-Supervised Learning for Robust Emotional Speech Synthesis with Limited Data
  • May 6, 2023
  • Applied Sciences
  • Jialin Zhang + 3 more

Emotional speech synthesis is an important branch of human–computer interaction technology that aims to generate emotionally expressive and comprehensible speech based on the input text. With the rapid development of speech synthesis technology based on deep learning, the research of affective speech synthesis has gradually attracted the attention of scholars. However, due to the lack of quality emotional speech synthesis corpus, emotional speech synthesis research under low-resource conditions is prone to overfitting, exposure error, catastrophic forgetting and other problems leading to unsatisfactory generated speech results. In this paper, we proposed an emotional speech synthesis method that integrates migration learning, semi-supervised training and robust attention mechanism to achieve better adaptation to the emotional style of the speech data during fine-tuning. By adopting an appropriate fine-tuning strategy, trade-off parameter configuration and pseudo-labels in the form of loss functions, we efficiently guided the learning of the regularized synthesis of emotional speech. The proposed SMAL-ET2 method outperforms the baseline methods in both subjective and objective evaluations. It is demonstrated that our training strategy with stepwise monotonic attention and semi-supervised loss method can alleviate the overfitting phenomenon and improve the generalization ability of the text-to-speech model. Our method can also enable the model to successfully synthesize different categories of emotional speech with better naturalness and emotion similarity.

  • Research Article
  • Cite Count Icon 57
  • 10.1155/2010/926951
Automatic Speech Recognition Systems for the Evaluation of Voice and Speech Disorders in Head and Neck Cancer
  • Aug 19, 2009
  • EURASIP Journal on Audio, Speech, and Music Processing
  • Andreas Maier + 7 more

In patients suffering from head and neck cancer, speech intelligibility is often restricted. For assessment and outcome measurements, automatic speech recognition systems have previously been shown to be appropriate for objective and quick evaluation of intelligibility. In this study we investigate the applicability of the method to speech disorders caused by head and neck cancer. Intelligibility was quantified by speech recognition on recordings of a standard text read by 41 German laryngectomized patients with cancer of the larynx or hypopharynx and 49 German patients who had suffered from oral cancer. The speech recognition provides the percentage of correctly recognized words of a sequence, that is, the word recognition rate. Automatic evaluation was compared to perceptual ratings by a panel of experts and to an age-matched control group. Both patient groups showed significantly lower word recognition rates than the control group. Automatic speech recognition yielded word recognition rates which complied with experts' evaluation of intelligibility on a significant level. Automatic speech recognition serves as a good means with low effort to objectify and quantify the most important aspect of pathologic speech--the intelligibility. The system was successfully applied to voice and speech disorders.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant