Consistency-Aware Multi-Channel Speech Enhancement Using Deep Neural Networks

Yoshiki Masuyama,Tatsuya Komatsu,Masahito Togami

doi:10.1109/icassp40776.2020.9053501

Abstract

This paper proposes a deep neural network (DNN)–based multichannel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often conducted in the time-frequency (T-F) domain because spatial filtering can be efficiently implemented in the T-F domain. In such a case, ordinary objective functions are computed on the estimated T-F mask or spectrogram. However, the estimated spectrogram is often inconsistent, and its amplitude and phase may change when the spectrogram is converted back to the time-domain. That is, the objective function does not evaluate the enhanced time-domain signal properly. To address this problem, we propose to use an objective function defined on the reconstructed time-domain signal. Specifically, speech enhancement is conducted by multi-channel Wiener filtering in the T-F domain, and its result is converted back to the time-domain. We propose two objective functions computed on the reconstructed signal where the first one is defined in the time-domain, and the other one is defined in the T-F domain. Our experiment demonstrates the effectiveness of the proposed system comparing to T-F masking and mask-based beamforming.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Consistency-Aware Multi-Channel Speech Enhancement Using Deep Neural Networks

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Insights Into Deep Non-Linear Filters for Improved Multi-Channel Speech Enhancement
Kristina Tesch ... Timo Gerkmann
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 31
Kristina Tesch, et. al.Kristina Tesch ... Timo Gerkmann
01 Jan 2023
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 31

Robust Speaker Recognition Based on Single-Channel and Multi-Channel Speech Enhancement
Hassan Taherian ... Deliang Wang
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 28
Hassan Taherian, et. al.Hassan Taherian ... Deliang Wang
01 Jan 2020
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 28

Tensor-To-Vector Regression for Multi-Channel Speech Enhancement Based on Tensor-Train Network
Jun Qi ... Chao-Han Huck Yang
-
Jun Qi, et. al.Jun Qi ... Chao-Han Huck Yang
01 May 2020
01 May 2020

CNN-based noise reduction for multi-channel speech enhancement system with discrete wavelet transform (DWT) preprocessing.
Pavani Cherukuru ... Mumtaz Begum Mustafa
PeerJ. Computer science | VOL. 10
Pavani Cherukuru, et. al.Pavani Cherukuru ... Mumtaz Begum Mustafa
28 Feb 2024
PeerJ. Computer science | VOL. 10

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Consistency-Aware Multi-Channel Speech Enhancement Using Deep Neural Networks

Abstract

Talk to us

Similar Papers