Enhancing the Feature Extraction Process for Automatic Speech Recognition with Fractal Dimensions

Aitzol Ezeiza,Karmele López De Ipiña,Carmen Hernández,Nora Barroso

doi:10.1007/s12559-012-9165-0

Enhancing the Feature Extraction Process for Automatic Speech Recognition with Fractal Dimensions

Aitzol Ezeiza, Karmele López De Ipiña + Show 2 more

https://doi.org/10.1007/s12559-012-9165-0

Copy DOI

Journal: Cognitive Computation	Publication Date: Jul 24, 2012
Citations: 25

Affiliation: University of the Basque Country

#Traditional Mel Frequency Cepstral Coefficients #Automatic Speech Recognition Task + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

Mel frequency cepstral coefficients (MFCCs) are a standard tool for automatic speech recognition (ASR), but they fail to capture part of the dynamics of speech. The nonlinear nature of speech suggests that extra information provided by some nonlinear features could be especially useful when training data are scarce or when the ASR task is very complex. In this paper, the Fractal Dimension of the observed time series is combined with the traditional MFCCs in the feature vector in order to enhance the performance of two different ASR systems. The first is a simple system of digit recognition in Chinese, with very few training examples, and the second is a large vocabulary ASR system for Broadcast News in Spanish.

Full Text