Tracking a moving speaker using excitation source information

Vikas C Raykar,S.R Mahadeva Prasanna,Ramani Duraiswami,B Yegnanarayana

doi:10.21437/eurospeech.2003-18

Abstract

Microphone arrays are widely used to detect, locate, and track a stationary or moving speaker. The first step is to estimate the time delay, between the speech signals received by a pair of microphones. Conventional methods like generalized crosscorrelation are based on the spectral content of the vocal tract system in the speech signal. The spectral content of the speech signalisaffectedduetodegradationsinthespeechsignalcaused by noise and reverberation. However, features corresponding to the excitation source of speech are less affected by such degradations. This paper proposes a novel method to estimate the time delays using the excitation source information in speech. The estimated delays are used to get the position of the moving speaker. The proposed method is compared with the spectrumbased approach using real data from a microphone array setup.

Full Text