Language modeling for automatic turkish broadcast news transcription

Ebru Arısoy,Murat Saraçlar,Haşim Sak

doi:10.21437/interspeech.2007-273

Language modeling for automatic turkish broadcast news transcription

Ebru Arısoy, Murat Saraçlar + Show 1 more

https://doi.org/10.21437/interspeech.2007-273

Copy DOI

Publication Date: Aug 27, 2007

Citations: 26

Affiliation: Boğaziçi University

#Turkish Broadcast News #Sub-word Language Models + Show 8 more

Abstract
Full-Text
Similar Papers

Abstract

The aim of this study is to develop a speech recognition system for Turkish broadcast news. State-of-the-art speech recognition systems utilize statistical models. A large amount of data is required to reliably estimate these models. For this study, a large Turkish Broadcast News database, consisting of the speech signal and corresponding transcriptions, is being collected. In this paper, information about this database and experiments performed using the system developed on the collected data are presented. In addition to the baseline system, various sub-word language models are investigated. Lexical stem-endings are proposed as a novel unit for language modeling and are shown to perform better than surface stem-endings and morphs. Currently, our best systems have lower than 20% error on clean speech.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.