An improved voice activity detection method based on spectral features and neural network

Liu Ting,Luo Xinwei

doi:10.3397/in-2021-2747

Abstract

The recognition accuracy of speech signal and noise signal is greatly affected under low signal-to-noise ratio. The neural network with parameters obtained from the training set can achieve good results in the existing data, but is poor for the samples with different the environmental noises. This method firstly extracts the features based on the physical characteristics of the speech signal, which have good robustness. It takes the 3-second data as samples, judges whether there is speech component in the data under low signal-to-noise ratios, and gives a decision tag for the data. If a reasonable trajectory which is like the trajectory of speech is found, it is judged that there is a speech segment in the 3-second data. Then, the dynamic double threshold processing is used for preliminary detection, and then the global double threshold value is obtained by K-means clustering. Finally, the detection results are obtained by sequential decision. This method has the advantages of low complexity, strong robustness, and adaptability to multi-national languages. The experimental results show that the performance of the method is better than that of traditional methods under various signal-to-noise ratios, and it has good adaptability to multi language.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An improved voice activity detection method based on spectral features and neural network

Abstract

Talk to us

Similar Papers

More From: INTER-NOISE and NOISE-CON Congress and Conference Proceedings

Lead the way for us

Similar Papers

Analysis of very low quality speech for mask-based enhancement

-

01 Sep 2013
01 Sep 2013

Parallelized Convolutional Recurrent Neural Network With Spectral Features for Speech Emotion Recognition
Pengxu Jiang ... Hongliang Fu
IEEE Access | VOL. 7
Pengxu Jiang, et. al.Pengxu Jiang ... Hongliang Fu
01 Jan 2019
IEEE Access | VOL. 7

Individual Thresholding of Voxel-based Functional Connectivity Maps. Estimation of Random Errors by Means of Surrogate Time Series.
M Clerici ... L Griffanti
Methods of information in medicine | VOL. 54
M Clerici, et. al.M Clerici ... L Griffanti
01 Jan 2015
Methods of information in medicine | VOL. 54

Mobile radio set comprising a speech processing arrangement
Walter Kellermann
The Journal of the Acoustical Society of America | VOL. 102
Walter KellermannWalter Kellermann
01 Jan 1997
The Journal of the Acoustical Society of America | VOL. 102

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An improved voice activity detection method based on spectral features and neural network

Abstract

Talk to us

Similar Papers

More From: INTER-NOISE and NOISE-CON Congress and Conference Proceedings