Audio-visual speech recognition with background music using single-channel source separation

Emad M Grais,Ibrahim Saygin Topkaya,Hakan Erdogan

doi:10.1109/siu.2012.6204436

Abstract

In this paper, we consider audio-visual speech recognition with background music. The proposed algorithm is an integration of audio-visual speech recognition and single channel source separation (SCSS). We apply the proposed algorithm to recognize spoken speech that is mixed with music signals. First, the SCSS algorithm based on nonnegative matrix factorization (NMF) and spectral masks is used to separate the audio speech signal from the background music in magnitude spectral domain. After speech audio is separated from music, regular audio-visual speech recognition (AVSR) is employed using multi-stream hidden Markov models. Employing two approaches together, we try to improve recognition accuracy by both processing the audio signal with SCSS and supporting the recognition task with visual information. Experimental results show that combining audio-visual speech recognition with source separation gives remarkable improvements in the accuracy of the speech recognition system.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Audio-visual speech recognition with background music using single-channel source separation

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Audio Visual Speech Recognition and Segmentation Based on DBN Models
Dongmei Jiang ... Hichem Sahli
-
Dongmei Jiang, et. al.Dongmei Jiang ... Hichem Sahli
01 Jun 2007
01 Jun 2007

Measuring the effect of high-speed video data on the audio-visual speech recognition accuracy
D V Ivanko ... D A Ryumin
Information and Control Systems | VOL. -
D V Ivanko, et. al.D V Ivanko ... D A Ryumin
19 Apr 2019
Information and Control Systems | VOL. -

Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices.
Dmitry Ryumin ... Denis Ivanko
Sensors | VOL. 23
Dmitry Ryumin, et. al.Dmitry Ryumin ... Denis Ivanko
17 Feb 2023
Sensors | VOL. 23

A probabilistic principal component analysis based hidden Markov model for audio-visual speech recognition
Zhanyu Ma ... Arne Leijon
-
Zhanyu Ma, et. al.Zhanyu Ma ... Arne Leijon
01 Oct 2008
01 Oct 2008

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Audio-visual speech recognition with background music using single-channel source separation

Abstract

Talk to us

Similar Papers