Does pronunciation distance predict accent categorization? Evidence for the respective contributions of segment distance and suprasegmentals
Does pronunciation distance predict accent categorization? Evidence for the respective contributions of segment distance and suprasegmentals
- Research Article
19
- 10.1371/journal.pone.0075734
- Jan 9, 2014
- PLoS ONE
In this study we develop pronunciation distances based on naive discriminative learning (NDL). Measures of pronunciation distance are used in several subfields of linguistics, including psycholinguistics, dialectology and typology. In contrast to the commonly used Levenshtein algorithm, NDL is grounded in cognitive theory of competitive reinforcement learning and is able to generate asymmetrical pronunciation distances. In a first study, we validated the NDL-based pronunciation distances by comparing them to a large set of native-likeness ratings given by native American English speakers when presented with accented English speech. In a second study, the NDL-based pronunciation distances were validated on the basis of perceptual dialect distances of Norwegian speakers. Results indicated that the NDL-based pronunciation distances matched perceptual distances reasonably well with correlations ranging between 0.7 and 0.8. While the correlations were comparable to those obtained using the Levenshtein distance, the NDL-based approach is more flexible as it is also able to incorporate acoustic information other than sound segments.
- Book Chapter
43
- 10.1007/978-3-319-69830-4_5
- Jan 1, 2018
In this study, we investigate which factors influence the linguistic distance of Catalan dialectal pronunciations from standard Catalan. We use pronunciations from three regions where the northwestern variety of the Catalan language is spoken (Catalonia, Aragon and Andorra). In contrast to Aragon, Catalan has an official status in both Catalonia and Andorra, which likely influences standardization. Because we are interested in the potentially large range of differences that standardization might promote, we examine 357 words in Catalan varieties and in particular their pronunciation distances with respect to the standard. In order to be sensitive to differences among the words, we fit a generalized additive mixed-effects regression model to this data. This allows us to examine simultaneously the general (i.e. aggregate) patterns in pronunciation distance and to detect those words that diverge substantially from the general pattern. The results reveal higher pronunciation distances from standard Catalan in Aragon than in the other regions. Furthermore, speakers in Catalonia and Andorra, but not in Aragon, show a clear standardization pattern, with younger speakers having dialectal pronunciations closer to the standard than older speakers. This clearly indicates the presence of a border effect within a single country with respect to word pronunciation distances. Since a great deal of scholarship focuses on single segment changes, we compare our analysis to the analysis of three segment changes that have been discussed in the literature on Catalan. This comparison shows that the pattern observed at the word pronunciation level is supported by two of the three cases examined. As not all individual cases conform to the general pattern, the aggregate approach is necessary to detect global standardization patterns.
- Conference Article
2
- 10.1109/icist.2014.6920396
- Apr 1, 2014
English is the only language available for international communication and is used by approximately 1.5 billions of speakers. It is also known to have a large diversity of pronunciation partly due to the influence of the speakers' mother tongue, called accents. Our project aims at creating a global and individual-basis map of English pronunciations to be used in teaching and learning World Englishes (WE) as well as research studies of WE [1], [2]. Creating the map mathematically requires a distance matrix in terms of pronunciation differences among all the speakers considered, and technically requires a method of predicting the pronunciation distance between any pair of the speakers. Our previous but very recent study [3] combined invariant pronunciation structure analysis [4], [5], [6], [7] and Support Vector Regression (SVR) effectively to predict the interspeaker pronunciation distances. In [3], very high correlation of 0.903 was observed between reference IPA-based pronunciation distances and the distances predicted by our proposed method. In this paper, after explaining our proposed method, some new results of analytical investigation of the method are described.
- Research Article
16
- 10.1121/10.0008930
- Dec 1, 2021
- The Journal of the Acoustical Society of America
Although unfamiliar accents can pose word identification challenges for children and adults, few studies have directly compared perception of multiple nonnative and regional accents or quantified how the extent of deviation from the ambient accent impacts word identification accuracy across development. To address these gaps, 5- to 7-year-old children's and adults' word identification accuracy with native (Midland American, British, Scottish), nonnative (German-, Mandarin-, Japanese-accented English) and bilingual (Hindi-English) varieties (one talker per accent) was tested in quiet and noise. Talkers' pronunciation distance from the ambient dialect was quantified at the phoneme level using a Levenshtein algorithm adaptation. Whereas performance was worse on all non-ambient dialects than the ambient one, there were only interactions between talker and age (child vs adult or across age for the children) for a subset of talkers, which did not fall along the native/nonnative divide. Levenshtein distances significantly predicted word recognition accuracy for adults and children in both listening environments with similar impacts in quiet. In noise, children had more difficulty overcoming pronunciations that substantially deviated from ambient dialect norms than adults. Future work should continue investigating how pronunciation distance impacts word recognition accuracy by incorporating distance metrics at other levels of analysis (e.g., phonetic, suprasegmental).
- Conference Article
2
- 10.1109/asru.2013.6707733
- Dec 1, 2013
English is the only language available for global communication. Due to the influence of speakers' mother tongue, however, those from different regions inevitably have different accents in their pronunciation of English. The ultimate goal of our project is creating a global pronunciation map of World Englishes on an individual basis, for speakers to use to locate similar English pronunciations. If the speaker is a learner, he can also know how his pronunciation compares to other varieties. Creating the map mathematically requires a matrix of pronunciation distances among all the speakers considered. This paper investigates invariant pronunciation structure analysis and Support Vector Regression (SVR) to predict the inter-speaker pronunciation distances. In experiments, the Speech Accent Archive (SAA), which contains speech data of worldwide accented English, is used as training and testing samples. IPA narrow transcriptions in the archive are used to prepare reference pronunciation distances, which are then predicted based on structural analysis and SVR, not with IPA transcriptions. Correlation between the reference distances and the predicted distances is calculated. Experimental results show very promising results and our proposed method outperforms by far a baseline system developed using an HMM-based phoneme recognizer.
- Conference Article
5
- 10.3115/1626516.1626521
- Jan 1, 2007
In this paper we use the Reeks Nederlandse Dialectatlassen as a source for the reconstruction of a 'proto-language' of Dutch dialects. We used 360 dialects from locations in the Netherlands, the northern part of Belgium and French-Flanders. The density of dialect locations is about the same everywhere. For each dialect we reconstructed 85 words. For the reconstruction of vowels we used knowledge of Dutch history, and for the reconstruction of consonants we used well-known tendencies found in most textbooks about historical linguistics. We validated results by comparing the reconstructed forms with pronunciations according to a proto-Germanic dictionary (Kobler, 2003). For 46% of the words we reconstructed the same vowel or the closest possible vowel when the vowel to be reconstructed was not found in the dialect material. For 52% of the words all consonants we reconstructed were the same. For 42% of the words, only one consonant was differently reconstructed. We measured the divergence of Dutch dialects from their 'proto-language'. We measured pronunciation distances to the proto-language we reconstructed ourselves and correlated them with pronunciation distances we measured to proto-Germanic based on the dictionary. Pronunciation distances were measured using Levenshtein distance, a string edit distance measure. We found a relatively strong correlation (r=0.87).
- Research Article
8
- 10.5842/47-0-649
- Jul 1, 2015
- Stellenbosch Papers in Linguistics Plus
Following Den Besten’s (2009) desiderata for historical linguistics of Afrikaans, this article aims to contribute some modern evidence to the debate regarding the founding dialects of Afrikaans. From an applied perspective (i.e. human language technology), we aim to determine which West Germanic language(s) and/or dialect(s) would be best suited for the purposes of recycling speech resources for the benefit of developing speech technologies for Afrikaans. Being recognised as a West Germanic language, Afrikaans is first compared to Standard Dutch, Standard Frisian and Standard German. Pronunciation distances are measured by means of Levenshtein distances. Afrikaans is found to be closest to Standard Dutch. Secondly, Afrikaans is compared to 361 Dutch dialectal varieties in the Netherlands and North-Belgium, using material from the Reeks Nederlandse Dialectatlassen , a series of dialect atlases compiled by Blancquaert and Pee in the period 1925-1982 which cover the Dutch dialect area. Afrikaans is found to be closest to the South-Holland dialectal variety of Zoetermeer; this largely agrees with the findings of Kloeke (1950). No speech resources are available for Zoetermeer, but such resources are available for Standard Dutch. Although the dialect of Zoetermeer is significantly closer to Afrikaans than Standard Dutch is, Standard Dutch speech resources might be a good substitute.
- Research Article
174
- 10.1371/journal.pone.0023613
- Sep 1, 2011
- PLoS ONE
In this study we examine linguistic variation and its dependence on both social and geographic factors. We follow dialectometry in applying a quantitative methodology and focusing on dialect distances, and social dialectology in the choice of factors we examine in building a model to predict word pronunciation distances from the standard Dutch language to 424 Dutch dialects. We combine linear mixed-effects regression modeling with generalized additive modeling to predict the pronunciation distance of 559 words. Although geographical position is the dominant predictor, several other factors emerged as significant. The model predicts a greater distance from the standard for smaller communities, for communities with a higher average age, for nouns (as contrasted with verbs and adjectives), for more frequent words, and for words with relatively many vowels. The impact of the demographic variables, however, varied from word to word. For a majority of words, larger, richer and younger communities are moving towards the standard. For a smaller minority of words, larger, richer and younger communities emerge as driving a change away from the standard. Similarly, the strength of the effects of word frequency and word category varied geographically. The peripheral areas of the Netherlands showed a greater distance from the standard for nouns (as opposed to verbs and adjectives) as well as for high-frequency words, compared to the more central areas. Our findings indicate that changes in pronunciation have been spreading (in particular for low-frequency words) from the Hollandic center of economic power to the peripheral areas of the country, meeting resistance that is stronger wherever, for well-documented historical reasons, the political influence of Holland was reduced. Our results are also consistent with the theory of lexical diffusion, in that distances from the Hollandic norm vary systematically and predictably on a word by word basis.
- Research Article
5
- 10.1353/jsl.2013.0013
- Jun 1, 2013
- Journal of Slavic Linguistics
The calculation of aggregate linguistic distances can compensate for some of the drawbacks inherent to the isogloss bundling method used in traditional dialectology to identify dialect areas. Synchronic aggregate analysis can also point out differences with respect to a diachronically based classification of dialects. In this study the Levenshtein algorithm is applied for the first time to obtain an aggregate analysis of the linguistic distances among 88 diatopic varieties of Croatian spoken along the Eastern Adriatic coast and in the Italian province of Molise. We also measured lexical differences among these varieties, which are traditionally grouped into Čakavian, Štokavian, and transitional Čakavian-Štokavian varieties. The lexical and pronunciation distances are subsequently projected onto multidimensional cartographic representations. Both kinds of analyses confirm that linguistic discontinuity is characteristic of the whole region, and that discontinuities are more pronounced in the northern Adriatic area than in the south. We also show that the geographic lines are in many cases the most decisive factor contributing to linguistic cohesion, and that the internal heterogeneity within Čakavian is often greater than the differences between Čakavian and Štokavian varieties. This holds both for pronunciation and lexicon.
- Conference Article
2
- 10.1109/icsda.2014.7051437
- Sep 1, 2014
English is the only language available for global communication and is known to have a large diversity of pronunciations due to the influence of speakers' mother tongue, called accents. Our previous studies [1], [2] made an attempt to do speaker-basis clustering of those pronunciations, where every speaker was assumed to speak with his own accent. The clustering procedure required a distance matrix only in terms of pronunciation differences among speakers and [1], [2] proposed a method to predict the pronunciation distance between any pair of the speakers. A distance matrix is often visualized on a two-dimensional plane by using the Multi-Dimensional Scaling (MDS) or drawing a dendrogram. In this study, considering learners' perceptual characteristics, a new method is proposed for visualization. When a visualization result is fed back to a learner, his main interest will be in the relations from himself to the others, not those among the others. Then, by using only a part of the distance matrix and other kinds of information such as age and gender, the proposed method can visualize multiple kinds of diversity found in acoustics of English pronunciation from a speaker's self-centered viewpoint. Unlike the conventional methods, our proposal is guaranteed to cause no distortion at all in results of visualization.
- Research Article
2
- 10.1155/2022/8269007
- Sep 20, 2022
- Mathematical Problems in Engineering
In order to solve the problem that violin performance evaluation is too subjective, this paper proposes a violin performance evaluation system based on mobile terminal technology. Violin sound evaluation mainly includes two indicators such as pronunciation quality and distance transmission ability. LabVIEW 8.2 is used to collect data, and fast Fourier transform (FFT) is used to analyze the time-domain signal in frequency domain, and the pronunciation process of violin is discussed in depth. Using statistical methods to process the test data of pronunciation distance transmission, a parameter model representing the pronunciation distance transmission ability is proposed, and finally the WeChat applet is used to control the recording and display. The results show that the process of violin sound production is accompanied by the weakening of octave peaks and the development of overtone peaks. The stable stage is marked by the stability of each octave peak and overtone peak. The sound head initially presents a simple excitation response, with only one near octave overtone and three cycles. Its maximum overtone peak/dominant frequency peak ratio coefficient is k = 0.36; the back entry complex excitation response has about nine cycles, with k = 0.735. Then enter the quasi-stable excitation response, whose waveform is similar to the typical A-string empty string pull waveform in the later stage, with k = 0.36; the maximum overtone peak/main frequency peak proportional coefficient is k = 0.39 in the rising section of the start, k = 0.2 in the loudest section, and k = 0.26 in the stable section. Conclusion. The system has initially established a new instrumented evaluation method, which provides a way of thinking for improving the quantifiable evaluation system in the future.
- Research Article
- 10.1016/j.specom.2026.103379
- Apr 1, 2026
- Speech Communication
Dysarthria is a type of motor speech disorder that reflects abnormalities in motor movements required for speech production. In clinical practice, identifying characteristic signs and symptoms of the neuropathophysiology underlying a dysarthria is vital for diagnosis and management. The gold standard for dysarthria assessment is auditory-perceptual evaluation by a speech and language therapist for differential diagnosis and management decisions. As the process is time-consuming for clinicians, there is growing interest in automatic dysarthria assessment (ADA). Recent approaches to ADA primarily focus on the classification of broad intelligibility or speech severity labels. However, this does not have much clinical utility and the assessment of communication-relevant parameters do not distinguish between dysarthria types and pathomechanisms. Studies on the classification of dysarthria function or clinical test protocol scores focusing on aspects of dysarthric speech production (such as the Frenchay dysarthria assessment (FDA)) are limited. Therefore, this paper focuses on the preliminary steps towards clinically interpretable ADA, including automatic FDA assessment. The phoneme posteriorgram (PPG) is a time-varying categorical distribution over acoustic speech units, and recent work demonstrates interpretable speech pronunciation distance for downstream tasks, e.g. pronunciation reconstruction. This work extends recent advances in posterior-based phoneme research and mispronunciation models to dysarthria assessment, exploring the extent to which dysarthric speech features in the FDA (identified by auditory-perceptual evaluation in clinical practice) are captured by PPG information. To achieve this, FDA aspects are systematically evaluated. The results show that interpretable PPG probability can capture dysarthric speech features that are related to motor system dysfunction. • This paper introduces a framework for the evaluation of interpretable features (based on analysis of dysarthric speech characteristics across relevant Frenchay dysarthria assessment (FDA) aspects) to address the lack of research on the automation of FDA score prediction. • Previous approaches to dysarthria classification have primarily focused on broad scores based on the deviation between control and dysarthric speech, or the classification of intelligibility labels (or intelligibility as an indicator of speech severity). This falls short of what is done in current clinical gold-standard dysarthria assessment, where auditory-perceptual evaluation is routinely conducted to define speech features related to dysfunction of the motor system (and the prosodic consequences). • FDA aspects are systematically evaluated (including manual listening by a speech and language therapist (SLT) as appropriate). A detailed analysis of the auditory perceptual features relevant to FDA scoring in the TORGO have not been previously documented, as well as analysis of the dysarthric speech processes in context of an interpretable feature. As a preliminary step towards clinically interpretable ADA (including automatic FDA assessment), this paper shows that interpretable PPG probability can capture these dysarthric speech features, with the potential utility to classify speech production impairment and a dysarthria profile. • Multiple pre-processing steps are required to process the TORGO phoneme alignment data. The code for this work will be publicly released. Previous studies have used the TORGO phoneme aligned data ( Yue et al., 2022 ). However, to the authors’ knowledge, no publicly available code is available to process this data.
- Research Article
23
- 10.3389/frai.2020.00039
- May 29, 2020
- Frontiers in Artificial Intelligence
We present an acoustic distance measure for comparing pronunciations, and apply the measure to assess foreign accent strength in American-English by comparing speech of non-native American-English speakers to a collection of native American-English speakers. An acoustic-only measure is valuable as it does not require the time-consuming and error-prone process of phonetically transcribing speech samples which is necessary for current edit distance-based approaches. We minimize speaker variability in the data set by employing speaker-based cepstral mean and variance normalization, and compute word-based acoustic distances using the dynamic time warping algorithm. Our results indicate a strong correlation of r = −0.71 (p < 0.0001) between the acoustic distances and human judgments of native-likeness provided by more than 1,100 native American-English raters. Therefore, the convenient acoustic measure performs only slightly lower than the state-of-the-art transcription-based performance of r = −0.77. We also report the results of several small experiments which show that the acoustic measure is not only sensitive to segmental differences, but also to intonational differences and durational differences. However, it is not immune to unwanted differences caused by using a different recording device.
- Research Article
8
- 10.1007/s12193-009-0015-7
- Dec 1, 2008
- Journal on Multimodal User Interfaces
This paper presents a novel audio visual diviseme (viseme pair) instance selection and concatenation method for speech driven photo realistic mouth animation. Firstly, an audio visual diviseme database is built consisting of the audio feature sequences, intensity sequences and visual feature sequences of the instances. In the Viterbi based diviseme instance selection, we set the accumulative cost as the weighted sum of three items: 1) logarithm of concatenation smoothness of the synthesized mouth trajectory; 2) logarithm of the pronunciation distance; 3) logarithm of the audio intensity distance between the candidate diviseme instance and the target diviseme segment in the incoming speech. The selected diviseme instances are time warped and blended to construct the mouth animation. Objective and subjective evaluations on the synthesized mouth animations prove that the multimodal diviseme instance selection algorithm proposed in this paper outperforms the triphone unit selection algorithm in Video Rewrite. Clear, accurate, smooth mouth animations can be obtained matching well with the pronunciation and intensity changes in the incoming speech. Moreover, with the logarithm function in the accumulative cost, it is easy to set the weights to obtain optimal mouth animations.
- Research Article
- 10.1121/1.2212629
- Jan 1, 2006
- The Journal of the Acoustical Society of America
A computer-based detection (e.g., speech recognition) system combines a word decoder and subword decoder to detect words (or phrases) in a spoken input provided by a user into a speaker connected to the detection system. The word decoder detects words by comparing an input pattern (e.g., of hypothetical word matches) to reference patterns (e.g., words). The subword decoder compares an input pattern (e.g., hypothetical words matches based on subword or phoneme recognition) to reference patterns (e.g., words) based on a word pronunciation distance measure that indicates how close each input pattern is to matching each reference pattern. The subword decoder sorts the source set of reference patterns based on a closeness of each reference pattern to correctly matching the input pattern based on generated pattern comparisons. The word decoder and subword decoder each provide an N-best list of hypothetical matches to the spoken input. A list fusion module of the detection system selectively combines the two N-best lists to produce a final or combined N-best list. The final or combined list has a predefined number of matches.