Mutual-optimization Towards Generative Adversarial Networks For Robust Speech Recognition

Ke Ding,Yanyan Xu,Dengfeng Ke,Ne Luo,Kaile Su

doi:10.1109/icpr.2018.8546090

Abstract

In the context of Automatic Speech Recognition (ASR), improving the noise robustness remains an intractable task. Speech enhancement, combined with Generative Adversarial Networks (GAN), such as SEGAN, has effective performance in denoising raw waveform speech signals. Instead of waveforms, using Mel filterbank spectra in GAN is proposed, which has better performance in the task of ASR. However, these techniques will still miss useful information when GAN is used in them. In this paper, we investigate to protect the useful information in GAN, and propose a novel model, called Discriminator Generator Classifier-GAN (DGC-GAN). While normal GAN combining just two networks will lead the model to denoising rather than recognition, DGC-GAN has another network called classifier, which is an ASR system that will tune GAN to be recognized easier. By adding a classifier into previous GAN to get DGC-GAN, we achieve 29.1% Phone Error Rate (PER) relative improvement in a tiny dataset and 47.4% PER relative improvement in a large dataset.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Mutual-optimization Towards Generative Adversarial Networks For Robust Speech Recognition

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition
Chris Donahue ... Bo Li
-
Chris Donahue, et. al.Chris Donahue ... Bo Li
01 Apr 2018
01 Apr 2018

Adversarial joint training with self-attention mechanism for robust end-to-end speech recognition
Lujun Li ... Ludwig Kürzinger
EURASIP Journal on Audio, Speech, and Music Processing | VOL. 2021
Lujun Li, et. al.Lujun Li ... Ludwig Kürzinger
05 Jul 2021
EURASIP Journal on Audio, Speech, and Music Processing | VOL. 2021

Using Auxiliary Sources of Knowledge for Automatic Speech Recognition

-

01 Jan 2004
01 Jan 2004

Data augmentation using generative adversarial networks for robust speech recognition
Yanmin Qian ... Tian Tan
Speech Communication | VOL. 114
Yanmin Qian, et. al.Yanmin Qian ... Tian Tan
19 Aug 2019
Speech Communication | VOL. 114

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Mutual-optimization Towards Generative Adversarial Networks For Robust Speech Recognition

Abstract

Talk to us

Similar Papers