Effective Monoaural Speech Separation through Convolutional Top-Down Multi-View Network

Aye Nyein Aung,Che-Wei Liao,Jeih-Weih Hung

doi:10.3390/fi16050151

Abstract

Speech separation, sometimes known as the “cocktail party problem”, is the process of separating individual speech signals from an audio mixture that includes ambient noises and several speakers. The goal is to extract the target speech in this complicated sound scenario and either make it easier to understand or increase its quality so that it may be used in subsequent processing. Speech separation on overlapping audio data is important for many speech-processing tasks, including natural language processing, automatic speech recognition, and intelligent personal assistants. New speech separation algorithms are often built on a deep neural network (DNN) structure, which seeks to learn the complex relationship between the speech mixture and any specific speech source of interest. DNN-based speech separation algorithms outperform conventional statistics-based methods, although they typically need a lot of processing and/or a larger model size. This study presents a new end-to-end speech separation network called ESC-MASD-Net (effective speaker separation through convolutional multi-view attention and SuDoRM-RF network), which has relatively fewer model parameters compared with the state-of-the-art speech separation architectures. The network is partly inspired by the SuDoRM-RF++ network, which uses multiple time-resolution features with downsampling and resampling for effective speech separation. ESC-MASD-Net incorporates the multi-view attention and residual conformer modules into SuDoRM-RF++. Additionally, the U-Convolutional block in ESC-MASD-Net is refined with a conformer layer. Experiments conducted on the WHAM! dataset show that ESC-MASD-Net outperforms SuDoRM-RF++ significantly in the SI-SDRi metric. Furthermore, the use of the conformer layer has also improved the performance of ESC-MASD-Net.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Future Internet	Publication Date: Apr 28, 2024
Citations: 1	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Effective Monoaural Speech Separation through Convolutional Top-Down Multi-View Network

Abstract

Talk to us

Similar Papers

More From: Future Internet

Lead the way for us

Similar Papers

Research on Speech Separation and Recognition Algorithm Based on Deep Learning
Sarah Wan
-
Sarah WanSarah Wan
29 Jul 2021
29 Jul 2021

Deep Neural Network-based Speech Separation Combining with MVDR Beamformer for Automatic Speech Recognition System
Bong-Ki Lee ... Jaewoong Jeong
-
Bong-Ki Lee, et. al.Bong-Ki Lee ... Jaewoong Jeong
01 Jan 2019
01 Jan 2019

A convolutional recurrent neural network with attention framework for speech separation in monaural recordings
Chao Sun ... Junhong Lu
Scientific Reports | VOL. 11
Chao Sun, et. al.Chao Sun ... Junhong Lu
14 Jan 2021
Scientific Reports | VOL. 11

Research on speech separation technology based on deep learning
Yan Zhou ... Xinyu Pan
Cluster Computing | VOL. 22
Yan Zhou, et. al.Yan Zhou ... Xinyu Pan
14 Feb 2018
Cluster Computing | VOL. 22

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Effective Monoaural Speech Separation through Convolutional Top-Down Multi-View Network

Abstract

Talk to us

Similar Papers

More From: Future Internet