Multi-Stream End-to-End Speech Recognition

Ruizhi Li,Xiaofei Wang,Sri Harish Mallidi,Takaaki Hori,Shinji Watanabe,Hynek Hermansky

doi:10.1109/taslp.2019.2959721

Abstract

Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end (E2E) Automatic Speech Recognition (ASR). The joint CTC/Attention model has achieved great success by utilizing both architectures during multi-task training and joint decoding. In this article, we present a multi-stream framework based on joint CTC/Attention E2E ASR with parallel streams represented by separate encoders aiming to capture diverse information. On top of the regular attention networks, the Hierarchical Attention Network (HAN) is introduced to steer the decoder toward the most informative encoders. A separate CTC network is assigned to each stream to force monotonic alignments. Two representative framework have been proposed and discussed, which are Multi-Encoder Multi-Resolution (MEM-Res) framework and Multi-Encoder Multi-Array (MEM-Array) framework, respectively. In MEM-Res framework, two heterogeneous encoders with different architectures, temporal resolutions and separate CTC networks work in parallel to extract complementary information from same acoustics. Experiments are conducted on Wall Street Journal (WSJ) and CHiME-4, resulting in relative Word Error Rate (WER) reduction of $\text{18.0}\!-\!\text{32.1}\%$ and the best WER of $\text{3.6}\%$ in the WSJ eval92 test set. The MEM-Array framework aims at improving the far-field ASR robustness using multiple microphone arrays which are activated by separate encoders. Compared with the best single-array results, the proposed framework has achieved relative WER reduction of $\text{3.7}\%$ and $\text{9.7}\%$ in AMI and DIRHA multi-array corpora, respectively, which also outperforms conventional fusion strategies.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IEEE/ACM Transactions on Audio, Speech, and Language Processing	Publication Date: Dec 26, 2019
Citations: 48	License type: publisher-specific-oa

R Discovery Prime

R Discovery Prime

Multi-Stream End-to-End Speech Recognition

Abstract

Talk to us

Similar Papers

More From: IEEE/ACM Transactions on Audio, Speech, and Language Processing

Lead the way for us

Similar Papers

Improving RNN Transducer Modeling for End-to-End Speech Recognition
Jinyu Li ... Hu Hu
-
Jinyu Li, et. al.Jinyu Li ... Hu Hu
13 Oct 2019
13 Oct 2019

Stream Attention-based Multi-array End-to-end Speech Recognition
Xiaofei Wang ... Takaaki Hori
-
Xiaofei Wang, et. al.Xiaofei Wang ... Takaaki Hori
01 May 2019
01 May 2019

Subband Temporal Envelope Features and Data Augmentation for End-to-end Recognition of Distant Conversational Speech
Cong-Thanh Do
-
Cong-Thanh DoCong-Thanh Do
01 May 2019
01 May 2019

Multiple-Hypothesis CTC-Based Semi-Supervised Adaptation of End-to-End Speech Recognition
Cong-Thanh Do ... Thomas Hain
-
Cong-Thanh Do, et. al.Cong-Thanh Do ... Thomas Hain
06 Jun 2021
06 Jun 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Multi-Stream End-to-End Speech Recognition

Abstract

Talk to us

Similar Papers

More From: IEEE/ACM Transactions on Audio, Speech, and Language Processing