Reverberant speech recognition combining deep neural networks and deep autoencoders augmented with a phone-class feature

Masato Mimura,Tatsuya Kawahara,Shinsuke Sakai

doi:10.1186/s13634-015-0246-6

Abstract

We propose an approach to reverberant speech recognition adopting deep learning in the front-end as well as b a c k-e n d o f a r e v e r b e r a n t s p e e c h r e c o g n i t i o n s y s t e m, a n d a n o v e l m e t h o d t o i m p r o v e t h e d e r e v e r b e r a t i o n p e r f o r m a n c e of the front-end network using phone-class information. At the front-end, we adopt a deep autoencoder (DAE) for enhancing the speech feature parameters, and speech recognition is performed in the back-end using DNN-HMM acoustic models trained on multi-condition data. The system was evaluated through the ASR task in the Reverb Challenge 2014. The DNN-HMM system trained on the multi-condition training set achieved a conspicuously higher word accuracy compared to the MLLR-adapted GMM-HMM system trained on the same data. Furthermore, feature enhancement with the deep autoencoder contributed to the improvement of recognition accuracy especially in the more adverse conditions. While the mapping between reverberant and clean speech in DAE-based dereverberation is conventionally conducted only with the acoustic information, we presume the mapping is also dependent on the phone information. Therefore, we propose a new scheme (pDAE), which augments a phone-class feature to the standard acoustic features as input. Two types of the phone-class feature are investigated. One is the hard recognition result of monophones, and the other is a soft representation derived from the posterior outputs of monophone DNN. The augmented feature in either type results in a significant improvement (7–8 % relative) from the standard DAE.

Highlights

In recent years, the automatic speech recognition (ASR) technology based on statistical techniques achieved a remarkable progress supported by the ever increasing training data and the improvements in the computing resources
4.2.2 Importance of delta feature To confirm the importance of delta and acceleration parameters in deep neural networks (DNN)-based acoustic modeling for reverberant speech recognition, we evaluated the DNN-hidden Markov models (HMM) system trained with only the static part of the acoustic feature of the multi-condition data
5 Conclusions In this paper, we investigated an approach to reverberant speech recognition adopting deep learning in the front-end as well as back-end of the system and evaluated it through the ASR task of the Reverb Challenge 2014

Summary

Introduction

The automatic speech recognition (ASR) technology based on statistical techniques achieved a remarkable progress supported by the ever increasing training data and the improvements in the computing resources. Following the great success of deep neural networks (DNN), speech dereverberation by deep autoencoders (DAE) has been investigated [9,10,11,12,13] In these works, DAEs are trained using reverberant speech features as input and the clean speech features as target so that they recover the clean speech from corrupted speech in the recognition stage. We propose to use deep learning both in the front-end (DAE-based dereverberation) and backend (DNN-HMM acoustic model) in a reverberant speech-recognition system. Recognition of reverberant speech is performed combining “standard” DNN-HMM [14] decoding and a feature enhancement through deep autoencoder (DAE) [9, 10, 15]. One of the advantages of the DNN-HMM is that they are suited for handling multiple frames, which is vital especially for reverberant speech recognition where we need to handle long-term artifacts

DNN for reverberant speech recognition

Combination of DAE front-end and DNN-HMM back-end

DAE augmented with a phone-class feature

Phone-class features

Experimental evaluations

Conclusions

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: EURASIP Journal on Advances in Signal Processing	Publication Date: Jul 23, 2015
Citations: 26	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Reverberant speech recognition combining deep neural networks and deep autoencoders augmented with a phone-class feature

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: EURASIP Journal on Advances in Signal Processing

Lead the way for us

Similar Papers

Deep autoencoders augmented with phone-class feature for reverberant speech recognition
Masato Mimura ... Tatsuya Kawahara
-
Masato Mimura, et. al.Masato Mimura ... Tatsuya Kawahara
01 Apr 2015
01 Apr 2015

Exploring deep neural networks and deep autoencoders in reverberant speech recognition
Masato Mimura ... Shinsuke Sakai
-
Masato Mimura, et. al.Masato Mimura ... Shinsuke Sakai
01 May 2014
01 May 2014

An identification of speaker-dependence in reverberant-robust speech recognition
Takahiro Fukumori ... Takanobu Nishiura
The Journal of the Acoustical Society of America | VOL. 131
Takahiro Fukumori, et. al.Takahiro Fukumori ... Takanobu Nishiura
01 Apr 2012
The Journal of the Acoustical Society of America | VOL. 131

Suitable Reverberation Criteria for Distant-talking Speech Recognition
Takanobu Nishiura ... Takahiro Fukumori
-
Takanobu Nishiura, et. al.Takanobu Nishiura ... Takahiro Fukumori
23 Jun 2011
23 Jun 2011

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Reverberant speech recognition combining deep neural networks and deep autoencoders augmented with a phone-class feature

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: EURASIP Journal on Advances in Signal Processing