Corpus-based dialectometry with topic models

Olli Kuparinen,Yves Scherrer

doi:10.1017/jlg.2024.6

Abstract

Abstract This paper presents a topic modeling approach to corpus-based dialectometry. Topic models are most often used in text mining to find latent structure in a collection of documents. They are based on the idea that frequently co-occurring words present the same underlying topic. In this study, topic models are used on interview transcriptions containing dialectal speech directly, without any annotations or preselected features. The transcriptions are modeled on complete words, on character n-grams, and after automatical segmentation. Data from three languages, Finnish, Norwegian, and Swiss German, are scrutinized. The proposed method is capable of discovering clear dialectal differences in all three datasets, while reflecting the differences between them. The method provides a significant simplification of the dialectometric workflow, simultaneously saving time and increasing objectivity. Using the method on non-normalized data could also benefit text mining, which is the traditional field of topic modeling.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Corpus-based dialectometry with topic models

Abstract

Talk to us

Similar Papers

More From: Journal of Linguistic Geography

Lead the way for us

Journal: Journal of Linguistic Geography	Publication Date: May 20, 2024
License type: CC BY 4.0

Similar Papers

Text Mining of Open-Ended Questions in Self-Assessment of University Teachers: An LDA Topic Modeling Approach
Diego Buenano-Fernandez ... Mario Gonzalez
IEEE Access | VOL. 8
Diego Buenano-Fernandez, et. al.Diego Buenano-Fernandez ... Mario Gonzalez
01 Jan 2020
IEEE Access | VOL. 8

Implementing Topic Modelling For Document Clustering
Jai Golechha ... Madhulika Bhadauria
-
Jai Golechha, et. al.Jai Golechha ... Madhulika Bhadauria
19 Jan 2023
19 Jan 2023

An Efficient Topic Modeling Approach for Text Mining and Information Retrieval through K-means Clustering
Junaid Rashid ... Aun Irtaza
Mehran University Research Journal of Engineering and Technology | VOL. 39
Junaid Rashid, et. al.Junaid Rashid ... Aun Irtaza
01 Jan 2020
Mehran University Research Journal of Engineering and Technology | VOL. 39

A Survey of Topic Modeling in Text Mining
Rubayyi Alghamdi ... Khalid Alfalqi
International Journal of Advanced Computer Science and Applications | VOL. 6
Rubayyi Alghamdi, et. al.Rubayyi Alghamdi ... Khalid Alfalqi
01 Jan 2015
International Journal of Advanced Computer Science and Applications | VOL. 6

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Corpus-based dialectometry with topic models

Abstract

Talk to us

Similar Papers

More From: Journal of Linguistic Geography