Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A new kid on the block: Distributional semantics predicts the word-specific tone signatures of monosyllabic words in conversational Taiwan Mandarin speech

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

A new kid on the block: Distributional semantics predicts the word-specific tone signatures of monosyllabic words in conversational Taiwan Mandarin speech

Similar Papers
  • Research Article
  • 10.14456/nujst.2018.29
Speech Intelligibility Evaluation of Sound Coding Strategies in Noisy Environments for Thai Cochlear Implant Users
  • Nov 16, 2018
  • NRCT Data Center
  • Siriporn Dachasilaruk + 1 more

The paper presents the effect of sound coding strategies on speech intelligibility of Thai-speaking cochlear implant (CI) users. Two sound coding strategies, namely continuous interleaved sampling (CIS) and advanced combination encoder (ACE) strategies were evaluated for word recognition. The tested words consisting of monosyllabic and bisyllabic words were corrupted by speech shaped noise and babble noise at SNR levels of 0, 5, 10 and 15 dB. The vocoded speech of clean and noisy words were tested by twelve normal-hearing listeners. The experimental results showed that speech intelligibility of the ACE strategy had higher mean scores than that of the CIS strategy in all tested conditions. The ACE strategy provided a significant speech intelligibility at SNR levels of 5 and 10 dB for monosyllabic words with speech shaped noise and at SNR levels of 0 dB for bisyllabic words with speech shaped noise. In addition, the noisy bisyllabic words provided higher intelligibility performance than the noisy monosyllabic words in all tested conditions. Keywords: cochlear implant, sound coding, speech intelligibility, filter bank, hearing loss

  • Research Article
  • Cite Count Icon 305
  • 10.1016/j.jml.2006.03.008
Morphological influences on the recognition of monosyllabic monomorphemic words
  • Jul 13, 2006
  • Journal of Memory and Language
  • R.H Baayen + 2 more

Morphological influences on the recognition of monosyllabic monomorphemic words

  • Research Article
  • Cite Count Icon 1
  • 10.1121/1.4743085
Function word reduction: The role of adjacent stress
  • Nov 1, 2000
  • The Journal of the Acoustical Society of America
  • Lisa M Lavoie

Monosyllabic function words in speech are often reduced from their full citation forms. Researchers have shown that reduction is influenced by part of speech, predictability, prosody, and disfluency. This work isolates the role of an additional factor, preceding stress, in function word reduction by analyzing the realization of ‘‘for’’ in a highly controlled corpus. Five American English speakers read pairs of disyllabic words with initial or final stress (BOOty, bouTIQUE) in the frame ‘‘Please say [word] for me.’’ Two speakers show striking sensitivity to preceding stress, always producing strong–weak alternations, such that strong ‘‘for’’ follows a weak syllable and vice versa. The rimes in the strong versions average 90 ms while the weak are just 50 ms. The strong rimes consist of a clearly segmentable vowel plus /r/. In the weak, however, /r/ is not distinguishable in the acoustics as its own segment; rather, it influences the vowel F3 value. The other three speakers always realize ‘‘for’’ as weak, with an average duration of 39 ms after either stressed or unstressed syllables. Though limited, these results suggest that some speakers construct rhythmic groupings in speech production. Further work will determine whether similar effects are found in spontaneous speech.

  • Research Article
  • Cite Count Icon 15
  • 10.1016/0885-2308(89)90021-1
Tone recognition of polysyllabic words in Mandarin speech
  • Jul 1, 1989
  • Computer Speech & Language
  • Lih-Cherng Liu + 3 more

Tone recognition of polysyllabic words in Mandarin speech

  • Research Article
  • Cite Count Icon 14
  • 10.1159/000080066
Relationship between Speech and Swallowing Disorders in Patients with Neuromuscular Disease
  • Sep 17, 2004
  • Folia Phoniatrica et Logopaedica
  • Masaki Nishio + 1 more

The intelligibility of monosyllabic speech, word speech, and conversational speech was evaluated in 113 dysarthric speakers, and the presence and severity of swallowing disorders were evaluated using videofluoroscopic and bedside examinations. The results revealed a high correlation between swallowing function and all levels of speech intelligibility. Furthermore, the prevalence of concomitant dysphagia in dysarthric patients was quite high regardless of the primary etiology and time elapsed since the onset. However, the relationship between the two functions is more complex than is initially apparent. The prevalence and severity of dysphagia vary markedly according to the type of dysarthria. Patients in the flaccid, spastic, and mixed categories encompass a broad range of severity levels with many individuals being severely impaired, while patients in the ataxic, hypokinetic, and unilateral upper motor neuron categories seldom have severe concomitant swallowing problems. Furthermore, the correlation coefficient between conversational intelligibility and swallowing function varies considerably according to the type of dysarthria. The correlation was not significant in the flaccid, hypokinetic, and UUMN dysarthria groups. Based on these findings, we discuss herein the clinical management of dysarthric patients with dysphagia.

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/chinsl.2008.ecp.68
Recognition of Syllable-Contracted Words in Spontaneous Speech Using Word Expansion and Duration Information
  • Dec 1, 2008
  • Wei-Bin Liang + 2 more

This paper presents a graphical model-based approach to syllable-contracted (SC) word recognition of spontaneous Mandarin speech. Phone deletion and pronunciation reduction are two major effects for the syllable-contracted words in spontaneous speech. In this study, the syllable- contracted (SC) words selected from a collected corpus are used for pronunciation lexicon expansion to deal with the phone deletion problem. The duration information of SC words is then incorporated to cover the effect of pronunciation reduction. The graphical model is employed to rescore all possible word sequences, including the expanded SC words, to obtain the final word sequence. In the experimental results, the Mandarin Conversional Dialogue Corpus (MCDC) was used to evaluate the proposed method. Compared with the previous work, a satisfactory improvement on the performance of the proposed approach can be achieved.

  • Research Article
  • Cite Count Icon 32
  • 10.1518/001872007779598055
Age Differences in Identifying Words in Synthetic Speech
  • Feb 1, 2007
  • Human Factors: The Journal of the Human Factors and Ergonomics Society
  • Roy W Roring + 2 more

We investigated whether context or different speech rates could improve older adult performance on identification of synthetically generated words. Synthetic speech systems can potentially improve the daily functioning of older adults. However, research must determine whether older adults can effectively implement current text-to-speech technologies, which few studies have examined. Older adults' sensory and cognitive declines may cause difficulties in identifying words in synthetic speech. Ninety-six participants (young, middle-aged, and older adults) identified auditory monosyllabic words (half natural, half synthetic) presented in isolation or at the ends of sentences. Participants heard speech at either normal or slower rates. We found an interaction of age, context, and voice type and that slower speech rates worsened performance for all groups. Contrasts revealed that context reduced age differences, though only for natural speech. Hearing acuity was highly correlated with age and fully accounts for the interaction. Context improves performance for everyone in natural speech. However, whereas context improves performance for synthetic speech, it does not differentially reduce the age impairment for older adults. Slower speed generally impairs everyone's performance compared with the normal rate. Systems using synthetic speech should avoid presenting words in isolation, and rich contextual support should be consistently adopted. Synthetic speech fidelity must be improved significantly before becoming truly useful for older adult populations.

  • Research Article
  • Cite Count Icon 1
  • 10.12785/amis/081l25
Tone Enhancing Model for Disyllable Words in Chinese Mandarin Speech
  • Apr 1, 2014
  • Applied Mathematics & Information Sciences
  • Jianbo Jiang + 4 more

Tone recognition is the core function in Chinese speech perception. The tone perception ability of people with sensorineural hearing loss (SNHL) is often weaker than normal people. Automatically tone enhancement would be useful in helping them understand Chinese speech better. In this paper, we focus on the tone enhancing model for Chinese disyllable words. We first analyze the acoustic features related to tone perception. By agglomerative hierarchical clus tering method, the first and second syllables of disyllable words are clustered into 6 clusters respectively. Discriminative features of the se clusters are experimentally determined from a set of possible features related to tone perception, such as the pitch value, pitch range an d position of minimum pitch, etc. We further propose a practicable tone enhancing model with these discriminative features: 1) an input pitch contour is classified by calculating the distance between it and the centroid of each cluster, and 2) selecting the smallest dis tance, then the unclassified pitch contour belongs to this cluster, 3) the pitch contour is modified for tone enhancement with model p arameters corresponding to this cluster using TD-PSOLA. Both statistical and subjective experiments show that higher hit rate of tone recognition can be obtained after tone enhancement with the proposed model. Especially, the proposed enhancing model can also avoid traditional tone recognition, which is more convictive and less laborious.

  • Conference Article
  • Cite Count Icon 1
  • 10.1109/iscslp49672.2021.9362073
Tone Realization in Mandarin Speech: A Large Corpus Based Study of Disyllabic Words
  • Jan 24, 2021
  • Yaru Wu + 2 more

This study aims to increase our knowledge about tone realization in disyllabic words in continuous Mandarin speech. Automatic alignments of large speech corpora were carried out to enable the study of potential tone variants, with a special focus on variation factors such as prosodic position and right tonal context. The alignments without tone variants (V0, phonological representation) show that Tone 4 is more frequent in phrase-final position than in other prosodic positions, supporting the "declination line" pattern often observed in speech production. Tone 4 is also the most frequent lexical tone (>50%) in all prosodic positions. Alignments permitting tone variants (V1, phonetic realization) show an increase of Tone 1 in phrase-initial position, compared to V0. Tone realization is observed to be related not only to the prosodic position, but also to the within-word right tonal context. Unsurprisingly, the most notable change in tone realization happens for Tone 3 in the first syllable of disyllabic words when followed by another Tone 3 because of the well-known "tone sandhi rule" in which T3T3 disyllabic words become T2T3. Cross-word right tonal context is found to impact only Tone 3. However, the results in this study show that Tone 3 sandhi rule is more a tendency than an absolute rule.

  • Research Article
  • Cite Count Icon 1096
  • 10.1006/cogp.1995.1010
Infants′ Detection of the Sound Patterns of Words in Fluent Speech
  • Aug 1, 1995
  • Cognitive Psychology
  • P.W Jusczyk + 1 more

Infants′ Detection of the Sound Patterns of Words in Fluent Speech

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 4
  • 10.1108/sjls-07-2023-0030
Phonological metathesis phenomenon in the early speech of Arabic-speaking children
  • Nov 7, 2023
  • Saudi Journal of Language Studies
  • Fawaz Qasem

Purpose The purpose of this study is to investigate the phonological metathesis phenomenon in the early speech of Arabic-speaking children. Based on analysis of a longitudinal data from 11 children's speech, the study mainly aims at investigating (1) the characteristics/nature of the phonological metathesis process in early child speech and (2) how it is different from adult phonological metathesis. The study explores the causes behind the metathesis phonological process in child speech. Design/methodology/approach This paper explored the phonological metathesis phenomenon based on a longitudinal study of 11 monolingual Arabic speakers of Yemeni-Ibbi Dialect (YIA) and cross-linguistic data from various languages. Findings The data analysis showed that metathesis phenomenon in child speech has several characteristics. It occurs at an early age, at 2 years and decreases with age. It was found that metathesis occurs mostly in disyllabic and trisyllabic words or more complex syllabic words, but metathesis rarely occurs in monosyllabic words in children's speech. The results indicated that unlike metathesis in adult speech, metathesis in children's speech occurs in undeliberate, slow, fast speech. The study explored that adjacency of sounds with similar phonological features, the ease of pronunciation, or the sonority effect are the motivations that trigger metathesis phenomenon to occur in child speech. Originality/value In the literature, the metathesis phonological process has got a little attention from researchers, and this is due to the rare cases of metathesis and the inconsistency of cases and the occurrence of metathesis. However, there is no consensus among the researchers about the causes of emergence and the occurrences behind the metathesis process. However, this study argues that the metathesis process has unique characteristics in child speech in comparison to adult speech and that there are some causes for metathesis to occur in human speech particularly in children's data such as adjacency of sounds with similar phonological features or sonority effect.

  • Research Article
  • 10.3109/13682829109012017
Perception of pitch contours in single words in alaryngeal speech
  • Jan 1, 1991
  • International Journal of Language & Communication Disorders
  • Alison Bridges

The perception and production of pitch contours were investigated in single words produced by two groups of alaryngeal speakers: tracheo-oesophageal (TE) and oesophageal (E) speakers. High quality tape-recordings of three tonal patterns by four oesophageal and eight tracheo-oesophageal speakers in monosyllabic words were judged by a group of six speech and language therapy listeners. The results indicated that tonal patterns can be produced with a relatively high level of reliability for both speaker groups. Some individual speakers from both groups approached predicted normal levels. These findings emphasise the importance of providing the opportunity for patients to acquire either of these speech modes in alaryngeal rehabilitation, rather than simply being provided with an artificial larynx, particularly in countries where tone languages are used. The high variability between groups also suggests that other variables apart from alaryngeal speech mode may be relevant in determining ability to signal tonal patterns.

  • Research Article
  • Cite Count Icon 11
  • 10.1044/1092-4388(2012/11-0341)
Additive Effects of Lengthening on the Utterance-Final Word in Child-Directed Speech
  • May 31, 2012
  • Journal of Speech, Language, and Hearing Research
  • Eon-Suk Ko + 1 more

The authors investigated lengthening effects in child-directed speech (CDS) across the sentence, testing the additive effects on duration of Word Position, Register, Focus, and Sentence Mode (statement/question). Five theater students produced 6 sentences containing 5 monosyllabic words in a simulated dialogue, varying in Register, Focus, and Sentence Mode. The authors segmented a total of 1,800 sentences using forced-alignment tools, and they analyzed the duration of each word. The results show significant effects of Register, Word Position, and their interactions. The simple effect of Register was significant in all 5 word positions, indicating a global elongation effect in CDS. Interestingly, there was no proportional increase of the final word in CDS. In addition, the 3-way interactions Register × Word Position × Focus and Register × Word Position × Sentence Mode were significant, which converge to the conclusion that the utterance-final word in CDS is additively elongated when it is focused and in a statement. Elongation in CDS is a global effect, but the additive effects of duration demonstrated in the authors' data suggest that the effect of enhanced utterance-final lengthening in CDS in naturalistic samples may be a by-product of discourse characteristics of CDS.

  • Research Article
  • Cite Count Icon 29
  • 10.1044/1092-4388(2009/07-0223)
The Development of Distinct Speaking Styles in Preschool Children
  • Jan 1, 2009
  • Journal of Speech, Language, and Hearing Research
  • Melissa A Redford + 1 more

To examine when and how socially conditioned distinct speaking styles emerge in typically developing preschool children's speech. Thirty preschool children, ages 3, 4, and 5 years old, produced target monosyllabic words with monophthongal vowels in different social-functional contexts designed to elicit clear and casual speaking styles. Thirty adult listeners were used to assess whether and at what age style differences were perceptible. Children's speech was acoustically analyzed to evaluate how style-dependent differences were produced. The ratings indicated that listeners could not discern style differences in 3-year-olds' speech but could hear distinct styles in 4-year-olds' and especially in 5-year-olds' speech. The acoustic measurements were consistent with these results: Style-dependent differences in 4- and 5-year-olds' words included shorter vowel durations and lower fundamental frequency in clear compared with casual speech words. Five-year-olds' clear speech words also had more final stop releases and initial sibilants with higher spectral energy than did their casual speech words. Formant frequency measures showed no style-dependent differences in vowel production at any age nor any differences in initial stop voice onset times. Overall, the findings suggest that distinct styles develop slowly and that early style-dependent differences in children's speech are unlike those observed in adult clear and casual speech. Children may not develop adultlike styles until they have acquired expert articulatory control and the ability to highlight the internal structure of an articulatory plan for a listener.

  • Research Article
  • 10.13064/ksss.2013.5.1.099
Prosodic Modifications of the Internal Phonetic Structure of Monosyllabic CVC Words in Conversational Speech
  • Mar 31, 2013
  • Phonetics and Speech Sciences
  • Yoonsook Mo

Previous laboratory studies have shown that prosodic structures are encoded in the modulations of phonetic patterns of speech including suprasegmental as well as segmental features. In particular, effects of prosodic context on duration and intensity of syllables and words have been widely reported. Drawing on prosodically annotated large-scale speech data from the Buckeye corpus of conversational speech of American English, the current study attempted to examine whether and how prosodic prominence and phrase boundary of everyday conversational speech, as determined by a large group of ordinary listeners, are related to the phonetic realization of duration and intensity. The results showed that the patterns of word durations and intensities are influenced by prosodic structure. Closer examinations revealed, however, that the effects of prosodic prominence are not the same as those of prosodic phrase boundary. With regard to intensity measures, the results revealed the systematic changes in the patterns of overall RMS intensity near prosodic phrase boundary but the prominence effects are restricted to the nucleus. In terms of duration measures, both prosodic prominence and phrase boundary are the most closely related to the lengthening of the nucleus. Yet, prosodic prominence is more closely related to the lengthening of the onset while phrase boundary lengthens the coda duration more. The findings from the current study suggest that the phonetic realizations of prosodic prominence are different from those of prosodic phrase boundary, and speakers signal different prosodic structures through deliberate modulations of the internal phonetic structure of words and listeners attend to such phonetic variations.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant