Improving BLSTM RNN based Mandarin speech recognition using accent dependent bottleneck features

Jiangyan Yi,Jianhua Tao,Zhengqi Wen,Hao Ni

doi:10.1109/apsipa.2016.7820723

Abstract

This paper proposes an approach to perform accent adaptation by using accent dependent bottleneck (BN) features to improve the performance of multi-accent Mandarin speech recognition system. The architecture of the adaptation uses two neural networks. First, deep neural network (DNN) acoustic model acts as a feature extractor which is used to extract accent dependent BN (BN-DNN) features. The input features of the BN-DNN model are MFCC features appended with i-vectors features. Second, bidirectional long short term memory (BLSTM) recurrent neural network (RNN) based acoustic model is used to perform accent-specific adaptation. The input features of the BLSTM RNN model are accent dependent BN features appended with MFCC features. Experiments on RASC863 and CASIA regional accent speech corpus show that the proposed method obtains obvious improvement compared with the BLSTM RNN baseline model.

Full Text