Speaker-adapted confidence measures for speech recognition of video lectures

Isaias Sanchez-Cortina,Jesús Andrés-Ferrer,Alberto Sanchis,Alfons Juan

doi:10.1016/j.csl.2015.10.003

Isaias Sanchez-Cortina, Jesús Andrés-Ferrer + Show 2 more

Open Access

https://doi.org/10.1016/j.csl.2015.10.003

Copy DOI

Abstract

Automatic speech recognition applications can benefit from a confidence measure (CM) to predict the reliability of the output. Previous works showed that a word-dependent naïve Bayes (NB) classifier outperforms the conventional word posterior probability as a CM. However, a discriminative formulation usually renders improved performance due to the available training techniques.Taking this into account, we propose a logistic regression (LR) classifier defined with simple input functions to approximate to the NB behaviour. Additionally, as a main contribution, we propose to adapt the CM to the speaker in cases in which it is possible to identify the speakers, such as online lecture repositories.The experiments have shown that speaker-adapted models outperform their non-adapted counterparts on two difficult tasks from English (videoLectures.net) and Spanish (poliMedia) educational lectures. They have also shown that the NB model is clearly superseded by the proposed LR classifier.

Full Text