Development of a Mandarin-English Bilingual Speech Recognition System for Real World Music Retrieval

Q Zhang,Y Yan,Y Lin,J Shao,J Pan

doi:10.1093/ietisy/e91-d.3.514

Abstract

In recent decades, there has been a great deal of research into the problem of bilingual speech recognition – to develop a recognizer that can handle inter-and intra-sentential language switching between two languages. This paper presents our recent work on the development of a grammar-constrained, Mandarin-English bilingual Speech Recognition System (MESRS) for real world music retrieval. Two of the main difficult issues in handling the bilingual speech recognition systems for real world applications are tackled in this paper. One is to balance the performance and the complexity of the bilingual speech recognition system; the other is to effectively deal with the matrix language accents in embedded language. In order to process the intra-sentential language switching and reduce the amount of data required to robustly estimate statistical models, a compact single set of bilingual acoustic models derived by phone set merging and clustering is developed instead of using two separate monolingual models for each language. In our study, a novel Two-pass phone clustering method based on Confusion Matrix (TCM) is presented and compared with the log-likelihood measure method. Experiments testify that TCM can achieve better performance. Since potential system users' native language is Mandarin which is regarded as a matrix language in our application, their pronunciations of English as the embedded language usually contain Mandarin accents. In order to deal with the matrix language accents in embedded language, different non-native adaptation approaches are investigated. Experiments show that model retraining method outperforms the other common adaptation methods such as Maximum A Posteriori (MAP). With the effective incorporation of approaches on phone clustering and non-native adaptation, the Phrase Error Rate (PER) of MESRS for English utterances was reduced by 24.47% relatively compared to the baseline monolingual English system while the PER on Mandarin utterances was comparable to that of the baseline monolingual Mandarin system. The performance for bilingual utterances achieved 22.37% relative PER reduction.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IEICE Transactions on Information and Systems	Publication Date: Mar 1, 2008
Citations: 5	License type: free

R Discovery Prime

R Discovery Prime

Development of a Mandarin-English Bilingual Speech Recognition System for Real World Music Retrieval

Abstract

Talk to us

Similar Papers

More From: IEICE Transactions on Information and Systems

Lead the way for us

Similar Papers

Mandarin-English bilingual Speech Recognition for real world music retrieval
Qingqing Zhang ... Yonghong Yan
-
Qingqing Zhang, et. al. Qingqing Zhang ... Yonghong Yan
01 Mar 2008
01 Mar 2008

Multi-pronounciation dictionary construction for Mandarin-English bilingual phrase speech recognition system
C Wang ... W Shi
-
C Wang, et. al.C Wang ... W Shi
01 Jul 2015
01 Jul 2015

Model Generation of Accented Speech using Model Transformation and Verification for Bilingual Speech Recognition
Han-Ping Shen ... Pei-Shan Tsai
ACM Transactions on Asian and Low-Resource Language Information Processing | VOL. 14
Han-Ping Shen, et. al.Han-Ping Shen ... Pei-Shan Tsai
20 Apr 2015
ACM Transactions on Asian and Low-Resource Language Information Processing | VOL. 14

Phone modeling and combining discriminative training for mandarinenglish bilingual speech recognition
Yanmin Qian ... Jia Liu
-
Yanmin Qian, et. al.Yanmin Qian ... Jia Liu
01 Jan 2009
01 Jan 2009

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Development of a Mandarin-English Bilingual Speech Recognition System for Real World Music Retrieval

Abstract

Talk to us

Similar Papers

More From: IEICE Transactions on Information and Systems