HMM-based Speech Enhancement Research Articles

Hidden Markov model (HMM)-based minimum mean square error speech enhancement method in Mel-frequency domain is focused on and a parallel cepstral and spectral (PCS) modeling is proposed. Both Mel-frequency spectral (MFS) and Mel-frequency cepstral (MFC) features are studied and experimented for speech enhancement. To estimate clean speech waveform from a noisy signal, an inversion from the Mel-frequency domain to the spectral domain is required which introduces distortion artifacts in the spectrum estimation and the filtering. To reduce the corrupting effects of the inversion, the PCS modeling is proposed. This method performs concurrent modeling in both cepstral and magnitude spectral domains. In addition to the spectrum estimator, magnitude spectrum, log-magnitude spectrum and power spectrum estimators are also studied and evaluated in the HMM-based speech enhancement framework.The performances of the proposed methods are evaluated in the presence of five noise types with different SNR levels and the results are compared with several established speech enhancement methods especially auto-regressive HMM-based speech enhancement. The experimental results for both subjective and objective tests confirm the superiority of the proposed methods in the Mel-frequency domain over the reference methods, particularly for non-stationary noises.

Read full abstract

A new voice activity detection (VAD) algorithm with soft decision output in Mel-frequency domain is developed based on hidden Markov model (HMM) and is incorporated in an HMM-based speech enhancement system. The proposed VAD uses a two-state ergodic HMM representing speech presence and speech absence. The states are constructed from noisy speech and noise HMMs used in the speech enhancement system. This composite model provides a robust detection of speech segments in the presence of noise and obviates the need for extra modeling in HMM-based speech enhancement applications. As the main purpose of the proposed VAD is to detect speech segments accurately, a hang-over mechanism is proposed and is applied on the output of the VAD to improve the speech detection rate. The VAD is integrated in the HMM-based speech enhancement system in Mel-frequency spectral (MFS) and cepstral (MFC) domains. The performance of the proposed VAD, the effectiveness of the hang-over mechanism and the performance of the VAD-integrated speech enhancement system are evaluated on four noise types at different SNR levels. The experimental results confirm the superiority of the proposed VAD compared to the reference methods particularly for speech detection rate at the dominant noisy conditions.

Read full abstract

HMM-based Speech Enhancement Research Articles

Articles published on HMM-based Speech Enhancement

Speech enhancement using hidden Markov models in Mel-frequency domain

Hidden-Markov-model-based voice activity detector with high speech detection rate for speech enhancement

HMM-based strategies for enhancement of speech signals embedded in nonstationary noise

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

HMM-based Speech Enhancement Research Articles

Articles published on HMM-based Speech Enhancement

Speech enhancement using hidden Markov models in Mel-frequency domain

Hidden-Markov-model-based voice activity detector with high speech detection rate for speech enhancement

HMM-based strategies for enhancement of speech signals embedded in nonstationary noise