Significance of GMM-UBM based Modelling for Indian Language Identification

V Ravi Kumar,Anil Kumar Vuppala,Hari Krishna Vydana

doi:10.1016/j.procs.2015.06.027

V Ravi Kumar, Anil Kumar Vuppala + Show 1 more

Open Access

https://doi.org/10.1016/j.procs.2015.06.027

Copy DOI

Export

Save

Cite

Journal: Procedia Computer Science	Publication Date: Jan 1, 2015
Citations: 15

Abstract
Full-Text
Similar Papers

Abstract

Listen

Abstract Most of the Indian languages are originated from Devanagari, the script of the Sanskrit language. In-spite of similarity in phoneme sets, every language its own influence on the phonotactic constraints of speech in that language. A modelling technique that is capable of capturing the slightest variations imparted by the language is a pre-requisite for developing a language identification system (LID). Use of Gaussian mixture modelling technique with a large number of mixture components demands a large training data for each language class, which is hard to collect and handle. In this work, phonotactic variations imparted by the different languages are modelled using Gaussian mixture modelling with a universal background model (GMM-UBM) technique. In GMM-UBM based modelling certain amount of data from all the language classes is pooled to develop a universal background model (UBM) and the model is adapted to each class. Spectral features (MFCC) are employed to represent the language specific phonotactic information of speech in different languages. During the present study, LID systems are developed using the speech samples from IITKGP-MLILSC. In this work, performance of the proposed GMM-UBM based LID system is compared with conventional GMM based LID system. An average improvement of 7–8% is observed due to the use of UBM-based modelling of developing a LID system.

Full Text

Published Version

View

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

Significance of GMM-UBM based Modelling for Indian Language Identification

Abstract

Published Version

Talk to us

Similar Papers

More From: Procedia Computer Science

Lead the way for us

Similar Papers

GMM-UBM Based Modeling for Language Identification using New Feature Vectors
Dr A Nagesh* ... Dr M Sadanandam
International Journal of Innovative Technology and Exploring Engineering | VOL. 4
Dr A Nagesh*, et. al.Dr A Nagesh* ... Dr M Sadanandam
28 Feb 2020
International Journal of Innovative Technology and Exploring Engineering | VOL. 4

Dialect Identification of Assamese Language using Spectral Features
Tanvira Ismail ... L Joyprakash Singh
Indian Journal of Science and Technology | VOL. 10
Tanvira Ismail, et. al.Tanvira Ismail ... L Joyprakash Singh
01 Feb 2017
Indian Journal of Science and Technology | VOL. 10

Multilingual Speech Corpus in Low-Resource Eastern and Northeastern Indian Languages for Speaker and Language Identification
Joyanta Basu ... Soma Khan
Circuits, Systems, and Signal Processing | VOL. 40
Joyanta Basu, et. al.Joyanta Basu ... Soma Khan
20 Apr 2021
Circuits, Systems, and Signal Processing | VOL. 40

Identification of Kamrupi dialect and similar languages
Tanvira Ismail ... Sushanta Kabir Dutta
-
Tanvira Ismail, et. al.Tanvira Ismail ... Sushanta Kabir Dutta
01 Feb 2017
01 Feb 2017

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

Significance of GMM-UBM based Modelling for Indian Language Identification

Abstract

Published Version

Talk to us

Similar Papers

More From: Procedia Computer Science