Four-Features Evaluation of Text to Speech Systems for Three Social Robots

Fernando Alonso Martin,María Malfaz,Miguel Ángel Salichs,Álvaro Castro-González,José Carlos Castillo

doi:10.3390/electronics9020267

Fernando Alonso Martin, María Malfaz + Show 3 more

Open Access

https://doi.org/10.3390/electronics9020267

Copy DOI

Journal: Electronics	Publication Date: Feb 5, 2020
Citations: 11	License type: CC BY 4.0

Affiliation: Carlos III University of Madrid

Abstract

The success of social robotics is directly linked to their ability of interacting with people. Humans possess verbal and non-verbal communication skills, and, therefore, both are essential for social robots to get a natural human–robot interaction. This work focuses on the first of them since the majority of social robots implement an interaction system endowed with verbal capacities. In order to do this implementation, we must equip social robots with an artificial voice system. In robotics, a Text to Speech (TTS) system is the most common speech synthesizer technique. The performance of a speech synthesizer is mainly evaluated by its similarity to the human voice in relation to its intelligibility and expressiveness. In this paper, we present a comparative study of eight off-the-shelf TTS systems used in social robots. In order to carry out the study, 125 participants evaluated the performance of the following TTS systems: Google, Microsoft, Ivona, Loquendo, Espeak, Pico, AT&T, and Nuance. The evaluation was performed after observing videos where a social robot communicates verbally using one TTS system. The participants completed a questionnaire to rate each TTS system in relation to four features: intelligibility, expressiveness, artificiality, and suitability. In this study, four research questions were posed to determine whether it is possible to present a ranking of TTS systems in relation to each evaluated feature, or, on the contrary, there are no significant differences between them. Our study shows that participants found differences between the TTS systems evaluated in terms of intelligibility, expressiveness, and artificiality. The experiments also indicated that there was a relationship between the physical appearance of the robots (embodiment) and the suitability of TTS systems.

Highlights

Social robots are intended to “live” around humans to help and/or entertain them
We considered the scores given to each Text To Speech (TTS) system, our independent variables, considering all the research questions, our dependent measures: the mean and the standard deviation values were calculated and are presented
The results show that Espeak was perceived as the most artificial TTS system, with a significant difference with respect to the other systems evaluated

Summary

Introduction

Social robots are intended to “live” around humans to help and/or entertain them. In this regard, the speech is probably the richest and the preferred way for humans to communicate, making the software that allows the robot to generate an artificial voice a crucial element during human–robot interaction. The speech is probably the richest and the preferred way for humans to communicate, making the software that allows the robot to generate an artificial voice a crucial element during human–robot interaction These systems, commonly known as Text To Speech (TTS) systems, can convert text to artificial voice. There are several definitions of what a TTS system is.

Methods

Results

Conclusion