A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese

Shiyu Zhou,Linhao Dong,Bo Xu,Shuang Xu

doi:10.1007/978-3-030-04221-9_19

Abstract

The choice of modeling units is critical to automatic speech recognition (ASR) tasks. Conventional ASR systems typically choose context-dependent states (CD-states) or context-dependent phonemes (CD-phonemes) as their modeling units. However, it has been challenged by sequence-to-sequence attention-based models. On English ASR tasks, previous attempts have already shown that the modeling unit of graphemes can outperform that of phonemes by sequence-to-sequence attention-based model. In this paper, we are concerned with modeling units on Mandarin Chinese ASR tasks using sequence-to-sequence attention-based models with the Transformer. Five modeling units are explored including context-independent phonemes (CI-phonemes), syllables, words, sub-words and characters. Experiments on HKUST datasets demonstrate that the lexicon free modeling units can outperform lexicon related modeling units in terms of character error rate (CER). Among five modeling units, character based model performs best and establishes a new state-of-the-art CER of \(26.64\%\) on HKUST datasets.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese
Shiyu Zhou ... Linhao Dong
-
Shiyu Zhou, et. al.Shiyu Zhou ... Linhao Dong
02 Sep 2018
02 Sep 2018

The use of discrete distributions with a very large codebook for automatic speech recognition and speaker verification
Guoli Ye
-
Guoli YeGuoli Ye
23 Dec 2014
23 Dec 2014

Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition
Qingyu Wang ... Bo Xu
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 37
Qingyu Wang, et. al.Qingyu Wang ... Bo Xu
26 Jun 2023
Proceedings of the AAAI Conference on Artificial Intelligence | VOL. 37

Using Auxiliary Sources of Knowledge for Automatic Speech Recognition

-

01 Jan 2004
01 Jan 2004

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese

Abstract

Talk to us

Similar Papers