Integration of deep learning with expectation maximization for spatial cue-based speech separation in reverberant conditions

Sania Gul,Muhammad Salman Khan,Syed Waqar Shah

doi:10.1016/j.apacoust.2021.108048

Abstract

In this paper, we formulate a blind source separation (BSS) framework, which allows integrating U-Net based deep learning source separation network with probabilistic spatial machine learning expectation maximization (EM) algorithm for separating speech in reverberant conditions. Our proposed model uses a pre-trained deep learning convolutional neural network, U-Net, for clustering the interaural level difference (ILD) cues and machine learning expectation maximization (EM) algorithm for clustering the interaural phase difference (IPD) cues. The integrated model exploits the complementary strengths of the two approaches to BSS: the strong modeling power of supervised neural networks and the ease of unsupervised machine learning algorithms, whose few parameters can be estimated on as little as a single segment of an audio mixture. The results show an average improvement of 4.3 dB in signal to distortion ratio (SDR) and 4.3% in short time speech intelligibility (STOI) over the EM based source separation algorithm MESSL-GS (model-based expectation–maximization source separation and localization with garbage source) and 4.5 dB in SDR and 8% in STOI over deep learning convolutional neural network (U-Net) based speech separation algorithm SONET under the reverberant conditions ranging from anechoic to those mostly encountered in the real world.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Integration of deep learning with expectation maximization for spatial cue-based speech separation in reverberant conditions

Abstract

Talk to us

Similar Papers

More From: Applied Acoustics

Lead the way for us

Journal: Applied Acoustics	Publication Date: Mar 27, 2021
Citations: 9

Similar Papers

Clustering of spatial cues by semantic segmentation for anechoic binaural source separation
Sania Gul ... Syed Waqar Shah
Applied Acoustics | VOL. 171
Sania Gul, et. al.Sania Gul ... Syed Waqar Shah
13 Aug 2020
Applied Acoustics | VOL. 171

Deep learning for speech separation.

-

30 Apr 2020
30 Apr 2020

Clot burden of acute pulmonary thromboembolism: comparison of two deep learning algorithms, Qanadli score, and Mastora score.
Hongxia Zhang ... Xiaojuan Guo
Quantitative Imaging in Medicine and Surgery | VOL. 12
Hongxia Zhang, et. al.Hongxia Zhang ... Xiaojuan Guo
01 Jan 2021
Quantitative Imaging in Medicine and Surgery | VOL. 12

Two-stage audio-visual speech dereverberation and separation based on models of the interaural spatial cues and spatial covariance
Muhammad Salman Khan ... Jonathon Chambers
-
Muhammad Salman Khan, et. al.Muhammad Salman Khan ... Jonathon Chambers
01 Jul 2013
01 Jul 2013

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Integration of deep learning with expectation maximization for spatial cue-based speech separation in reverberant conditions

Abstract

Talk to us

Similar Papers

More From: Applied Acoustics