Is Attention Always Needed? A Case Study on Language Identification from Speech

Atanu Mandal,Sudip Kumar Naskar,Mahidas Bhattacharya,Indranil Dutta,Santanu Pal

doi:10.2139/ssrn.4186504

Abstract

Language Identification (LID), a recommended initial step to Automatic Speech Recognition (ASR), is used to detect a spoken language from audio specimens. In state-of-the-art systems capable of multilingual speech processing, however, users have to explicitly set one or more languages before using them. LID, therefore, plays a very important role in situations where ASR based systems cannot parse the uttered language in multilingual contexts causing failure in speech recognition. We propose an attention based convolutional recurrent neural network (CRNN with Attention) that works on Mel-frequency Cepstral Coefficient (MFCC) features of audio specimens. Additionally, we reproduce some state-of-the-art approaches, namely Convolutional Neural Network (CNN) and Convolutional Recurrent Neural Network (CRNN), and compare them to our proposed method. We performed extensive evaluation on thirteen different Indian languages and our model achieves classification accuracy over 98%. Our LID model is robust to noise and provides 91.2% accuracy in a noisy scenario. The proposed model is easily extensible to new languages.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Is Attention Always Needed? A Case Study on Language Identification from Speech

Abstract

Talk to us

Similar Papers

More From: SSRN Electronic Journal

Lead the way for us

Similar Papers

Identification of Indian classical languages using Convolutional Recurrent Neural Networks
Ibin Oommen ... Anu George
-
Ibin Oommen, et. al.Ibin Oommen ... Anu George
18 Dec 2020
18 Dec 2020

Spoken Language Identification System Using Convolutional Recurrent Neural Network
Adal A Alashban ... Ali H Meftah
Applied Sciences | VOL. 12
Adal A Alashban, et. al.Adal A Alashban ... Ali H Meftah
13 Sep 2022
Applied Sciences | VOL. 12

Speaker-based language identification for Ethio-Semitic languages using CRNN and hybrid features
Malefia Demilie Melese ... Ibrahim Gashaw Kasa
Network: Computation in Neural Systems | VOL. ahead-of-print
Malefia Demilie Melese, et. al.Malefia Demilie Melese ... Ibrahim Gashaw Kasa
06 Jun 2024
Network: Computation in Neural Systems | VOL. ahead-of-print

Deep Convolution Neural Network for Thai Classical Music Instruments Sound Recognition
Apichai Huaysrijan ... Sunee Pongpinigpinyo
-
Apichai Huaysrijan, et. al.Apichai Huaysrijan ... Sunee Pongpinigpinyo
18 Nov 2021
18 Nov 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Is Attention Always Needed? A Case Study on Language Identification from Speech

Abstract

Talk to us

Similar Papers

More From: SSRN Electronic Journal