Articles published on Spontaneous speech
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
4105 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.jml.2026.104763
- Aug 1, 2026
- Journal of memory and language
- Melissa A Redford + 2 more
Pausing to breathe and the speech-language relationship in production.
- Research Article
- 10.1044/2026_jslhr-25-00550
- Jun 29, 2026
- Journal of speech, language, and hearing research : JSLHR
- Lifang Qiu + 6 more
Poststroke aphasia (PSA) significantly impairs language and cognitive functioning, yet few interventions comprehensively address both domains. This study evaluated the efficacy of an intervention combining acupuncture and computer-assisted cognitive rehabilitation using the RehaCom cognitive training system (hereinafter, ACR) on language and nonverbal cognitive function in individuals with PSA. A single-center randomized controlled trial was conducted at a tertiary rehabilitation hospital in Fujian Province, China. Eighty patients with PSA were randomly assigned (1:1) to either the combined-intervention group (ACR plus standard speech and language therapy [SLT]) or the SLT group (SLT only). The intervention consisted of five 30-min sessions per week. Primary and secondary outcomes were assessed using the Western Aphasia Battery-Revised (WAB-R) and the Non-language-based Cognitive Assessment (NLCA) at baseline, Week 4, and Week 12 (post-intervention). Generalized estimating equations were used to evaluate group differences over time. Of the 80 enrolled participants, 71 (88.75%) completed the 12-week intervention. Compared with the SLT group, participants in the combined-intervention group demonstrated significantly greater improvements in the WAB-R Aphasia Quotient at both 4 weeks (B = 5.961, 95% CI [3.765, 8.158], p < .001, Cohen's d = 0.135) and 12 weeks (B = 11.806, 95% CI [8.050, 15.563], p < .001, Cohen's d = 0.460). Significant gains were also observed in WAB-R subdomains, including Spontaneous Speech, Auditory Comprehension, Repetition, and Naming, with small-to-moderate effect sizes. Moreover, the combined-intervention group exhibited substantial improvements in nonverbal cognitive function, as measured by the NLCA, with moderate-to-large effect sizes at both post-intervention points (Cohen's d = 0.434 at 4 weeks; Cohen's d = 0.847 at 12 weeks). The ACR intervention yielded significant and clinically meaningful improvements in both language and cognitive functions in patients with PSA. This multimodal approach represents a promising therapeutic strategy for comprehensive aphasia rehabilitation.
- Research Article
- 10.1044/2026_ajslp-25-00455
- Jun 29, 2026
- American journal of speech-language pathology
- Jeremy Wolfberg + 11 more
The Rehabilitation Treatment Specification System (RTSS) is a theory-driven framework for identifying the active ingredients of rehabilitation treatments and was used to create standard labels in voice therapy (RTSS-Voice). The RTSS and RTSS-Voice have high validity, but reliability has not been tested. This study evaluated if RTSS-Voice ingredients can be reliably coded in standard-of-care voice therapy videos and used to identify commonalities and differences in ingredient dosage across five expert voice clinicians from different institutions. Five speech-language pathologists from different institutions videotaped one voice therapy session for 36 patients with vocal hyperfunction. Two trained graduate students coded sessions using four RTSS-Voice ingredients: (a) opportunities to practice voicing, (b) feedback, (c) volition ingredients, and (d) vocal hygiene information. Twelve sessions were double coded for interrater reliability. Practice voicing codes achieved strong to almost perfect reliability (κs: .80-.95), while informational codes achieved moderate-to-strong reliability (κs: .61-.87). All clinicians spent more time on volition ingredients (14-34 min/52.4%-65.7% of the session) compared to practice voicing (3-8 min/6.3%-23.3% of the session) and feedback (3-11 min/10.8%-19.1% of the session). For practice voicing, clinicians incorporated modified airflow and resonance into over 70% of their practice and spent more time in nonspeech and structured speech compared to spontaneous speech practice. Clinicians differed in delivering volitional information and feedback directly versus indirectly. Multiple RTSS-Voice ingredients were reliable across raters, indicating that, with appropriate training, these are readily observable clinician actions that can aid in the study of dosage in voice therapy. The RTSS and RTSS-Voice also identified commonalities and differences in ingredient usage across multiple expert clinicians. Future work could use the RTSS and RTSS-Voice to investigate relationships between variations in therapy ingredients and individual patient outcomes. https://doi.org/10.23641/asha.32785938.
- Research Article
- 10.1016/j.msard.2026.107342
- Jun 25, 2026
- Multiple sclerosis and related disorders
- Maiara Laís Mallmann Kieling Peres + 3 more
Self-perception of speech impairment in MS: Associations with dysarthria severity and speech measures.
- Research Article
- 10.1016/j.jvoice.2026.06.002
- Jun 24, 2026
- Journal of voice : official journal of the Voice Foundation
- Micalle Carl + 2 more
Voice Disorders and Speech Intelligibility in Developmental Dysarthria: Influence of Speech Task.
- Research Article
- 10.1044/2026_jslhr-25-00507
- Jun 10, 2026
- Journal of speech, language, and hearing research : JSLHR
- Mélanie Canault + 3 more
This study reports age-related changes in articulation rate (syllables [SPS] and phones per second [PPS]) during spontaneous speech in French children aged 2-4 years and provides baseline values for this age range. In this cross-sectional study, spontaneous speech (4,361 utterances) from 91 French children was collected. The articulation rate was calculated in SPS and PPS and was observed as a function of seven age groups. The distribution of articulation rate values for utterances in SPS and PPS by percentiles is provided. In line with previous studies, the results confirm that the articulation rate increased with age. It increased from 3.14 to 3.88 SPS between 22 and 50 months, and from 6.06 to 8.31 PPS with significant changes emerging at 38-41 months. The results also indicate that the number of sounds per syllable increased significantly with age and that the growth in syllable structure complexity preceded that of articulation rate in SPS. This study adds to the already available benchmarks for articulation rate by providing new data for French preschoolers. Further studies are still needed to understand what other factors (e.g., cognitive, linguistic and spontaneous speech styles) may be involved in the growth of articulation rate during development.
- Research Article
- 10.1044/2026_jslhr-25-00332
- Jun 10, 2026
- Journal of speech, language, and hearing research : JSLHR
- Laura Giglio + 2 more
Verb argument structure (VAS) is often impaired in poststroke aphasia. Previous studies investigated VAS impairments using constrained experiments or less often narrative speech, which relies on time-consuming manual annotations of VAS in speech transcriptions. Here, we aimed to develop and validate an automated approach to quantify VAS use in narrative speech using natural language processing (a dependency parser) and to apply it to a new data set. First, we validated this approach by comparing manually annotated measures of VAS from a previous study, with automatically annotated VAS measures developed here. We then applied this approach to a new data set of participants with aphasia (n = 106), whose narrative discourse samples were compared to a control group from AphasiaBank to (a) replicate previous findings and (b) assess VAS impairments in agrammatism specifically. The analysis of the new data set using automatically annotated VAS revealed that participants with Broca's aphasia and, according to grammatical distinctions, agrammatic participants were more likely to select verbs used with fewer arguments, on average, and to produce fewer arguments than controls. This study reveals that dependency parsers are suitable to characterize VAS use in spontaneous speech with virtually absent manual labor, confirming the feasibility of developing clinically applicable tools for complex language analysis. The results of the study additionally show that participants with aphasia differ from controls in their VAS sensitivity and abilities.
- Research Article
- 10.1016/j.jvoice.2026.04.037
- Jun 8, 2026
- Journal of voice : official journal of the Voice Foundation
- Liese Bostyn + 5 more
Contextual Voice Training for Student Teachers: Exploring the Role of Virtual Reality.
- Research Article
- 10.1080/09588221.2026.2685745
- Jun 8, 2026
- Computer Assisted Language Learning
- Djemai Mahmoud Boulaares
Despite extensive development of computer-assisted language learning (CALL) tools and Intelligent Tutoring Systems (ITS) addressing grammar, learners often still struggle to apply complex morpho-syntactic rules in spontaneous speech. In particular, Arabic poses challenges in agreement (gender, number, case) that traditional instruction and rule-based feedback have only partially overcome. Moreover, recent reviews reveal that generative-AI research in language learning has focused overwhelmingly on written English tasks, leaving speaking and less-studied languages largely underexplored. To fill this gap, we conducted a longitudinal mixed-methods quasi-experiment comparing a GPT-4–based dialogic tutor to conventional explicit instruction for non-native Arabic learners’ oral mastery of agreement rules. Participants (N = 60 intermediate AFL learners) were assigned by intact class to either an AI-mediated speaking practice environment or traditional drill-based instruction. Pre-, post-, and delayed-post oral tests (elicited speech tasks targeting verb–subject, adjective–noun, and subject–predicate agreement) were analysed for accuracy in obligatory contexts, error density per 100 words, and fluency metrics (speech rate, pause ratio, response latency). System log data (feedback events, response times) and learner questionnaires (anxiety, perceived usefulness) provided additional insights. Results from ANCOVAs and mixed-effects models suggested that the AI group outperformed the control on agreement accuracy and showed greater improvements in fluency indices, with these gains largely maintained at a 4-week delay. Error density declined more sharply in the AI condition. Learning analytics indicated that log-derived features (e.g. accuracy, focus, and time-on-task) significantly predicted individual gains. Learners interacting with the AI tutor reported lower speaking anxiety and high technology acceptance, consistent with broader evidence that AI-mediated speaking support can enhance enjoyment and willingness to communicateLearners interacting with the AI tutor reported lower speaking anxiety and high technology acceptance, echoing Zhang et al. (2024) findings of boosted enjoyment and willingness to communicate under an AI speaking assistant. Qualitative comments further suggested that the dialogic agent may have functioned as a safe, personalised practice space. In sum, this study suggests that a generative-AI dialogic tutor may support the transition from explicit rule knowledge to more fluent use of Arabic agreement, thereby potentially scaffolding the proceduralisation of grammar. This work helps address the dearth of speaking-focused GenAI research, is consistent with cognitive load theory, and may extend skill-acquisition accounts by suggesting a possible role for GenAI in lowering the affective filter and enhancing noticing. Pedagogically, we frame our system as an Adaptive Learning Ecosystem that dynamically modulates task difficulty and feedback. We conclude with recommendations for curriculum designers on ethically integrating AI tutors – from scaffolding prompts to protecting student data – to harness GenAI affordances while preserving language-specific complexity.
- Research Article
- 10.1016/j.dib.2026.112932
- Jun 4, 2026
- Data in Brief
- Manuela Jaeger + 3 more
Dataset of audiovisual unscripted monologues with speech annotations
- Research Article
- 10.1044/2026_aja-25-00171
- Jun 2, 2026
- American journal of audiology
- Varsha Rallapalli + 2 more
This study aimed to compare the effects of independent dynamic range compression (independent DRC) and conventional mixed-signal dynamic range compression (conventional DRC) on the sound quality of speech and music, in an ideal condition where the individual sound sources are available to the compressor. The hypothesis is that independent DRC optimized for individual signals yields higher sound quality ratings than conventional DRC optimized for a single signal. Participants were 15 young adults with audiometrically normal hearing. Stimuli included a 10-s-long spontaneous speech sample spoken by a female talker mixed with a classical music excerpt. The speech level was fixed at 65 dBA, and the music level was varied to achieve three speech-to-background ratios (SBRs; -10, 0, and 10 dB). A hearing aid simulator applied frequency-specific gain for a standard mild-moderate hearing loss with three amplification strategies: (a) conventional DRC-speech and music were compressed with a 50-ms release time after the signals were mixed; (b) independent DRC-speech was compressed with a 50-ms release time, and music was independently compressed with a 4,000-ms release before the signals were mixed; and (c) linear amplification (control condition)-stimuli were processed with linear amplification, before or after mixing. The participants rated sound quality across four scales: overall sound quality, speech clarity, speech naturalness, and music naturalness. Independent DRC yielded higher ratings than conventional DRC and linear amplification for overall sound quality and speech clarity at -10 dB SBR. Both DRC strategies resulted in poorer speech clarity ratings than linear amplification at 10 dB SBR. There were no significant differences between the amplification strategies for speech naturalness and music naturalness scales across SBRs. The results are a proof-of-concept that, under ideal circumstances, independent compression of sound sources can improve speech clarity and overall sound quality, particularly when the speech is softer than background music. Generally, DRC may result in poorer speech clarity compared to linear amplification when speech is more intense than background music. The results provide a baseline for future work involving independent compression of sound sources. https://doi.org/10.23641/asha.31855201.
- Research Article
- 10.1016/j.laheal.2026.100081
- Jun 1, 2026
- Language and Health
- Andrea Juan De León + 2 more
Linguistic description of sentence production of Spanish speakers with Williams syndrome
- Research Article
- 10.1080/14992027.2026.2668502
- May 29, 2026
- International Journal of Audiology
- Yao Wang + 12 more
Objective To identify risk factors associated with psychological inflexibility among caregivers of children with hearing loss (CHL) and to propose targeted support strategies based on findings. Design Caregivers’ psychological inflexibility was assessed using the Acceptance and Action Questionnaire-Management of Child Hearing Loss. Children’s auditory functioning in everyday situations and the intelligibility of their spontaneous speech were evaluated using the Classification of Auditory Performance (CAP) and the Speech Intelligibility Rating (SIR), respectively. Binary logistic regression analyses were conducted to determine significant risk and protective factors associated with caregivers’ psychological inflexibility. Study sample A total of 248 caregivers and their children participated in the study. Results Older age of the child, severe to profound hearing loss, and older caregiver age were identified as significant risk factors for higher psychological inflexibility. In contrast, better speech intelligibility in children served as a protective factor. Conclusions Improving caregivers’ psychological flexibility requires the development of an integrated support framework that incorporates early risk identification, family empowerment, respite services, and the psychological integration of rehabilitation resources. Embedding this framework within existing rehabilitation systems may facilitate a shift from a child-centred rehabilitation model towards a more holistic, family-centered approach that promotes coordinated development of the entire family system.
- Research Article
- 10.1016/j.schres.2026.05.018
- May 28, 2026
- Schizophrenia research
- Jean-Pierre Lindenmayer + 7 more
Psychometric analysis of speech-based digital biomarkers of alogia symptoms in treatment resistant schizophrenia: Comparison with clinical ratings of expressive negative symptoms.
- Research Article
- 10.1016/j.scog.2026.100443
- May 26, 2026
- Schizophrenia Research: Cognition
- Perry Van Der Zande + 2 more
From thought to language: Comparing schizophrenia spectrum disorders and Wernicke\u2019s aphasia with machine learning and LLMs
- Research Article
- 10.1080/10228195.2026.2631467
- May 24, 2026
- Language Matters
- Lawren Smith + 1 more
This article examines extraposition in Paarl-Kaaps against the backdrop of research showing it to be a multifactorial, gradient phenomenon shaped by syntactic, prosodic, and processing factors. Using spontaneous speech data from the SEcoKa corpus, we conducted a series of binary univariate logistic regressions, followed by a multivariate analysis, to identify significant predictors of extraposition, and their interactions. The strongest overall predictor is grammatical category: extraposition occurs most frequently with relative clauses (RCs), followed by adpositional phrases (PPs), direct objects, and short adverbs. Further analysis reveals that PP-extraposition is robustly conditioned by constituent weight and argument status, while relative clause- extraposition shows an unexpected, inverted weight pattern, suggesting that structural constraints tied to head type or attachment site obscure weight effects. Despite category-based contrasts, light constituents (including adverbs) display non-negligible extraposition rates, suggesting that lightness is not, in itself, a strong inhibitor of extraposition.
- Research Article
- 10.1016/j.neuropsychologia.2026.109505
- May 22, 2026
- Neuropsychologia
- Derya Çokal + 8 more
Automated detection of referential features in schizophrenic speech using large language models.
- Research Article
- 10.1136/bmjopen-2025-112999
- May 20, 2026
- BMJ Open
- Luc\Xeda Fern\Xe1Ndez-Romero + 11 more
IntroductionPrimary progressive aphasia (PPA) is a neurodegenerative syndrome associated with Alzheimer’s disease and frontotemporal degeneration. Non-invasive brain stimulation (NIBS) is a promising treatment, especially associated with language therapy, but comparative efficacy and long-term effects between the different techniques (transcranial direct current stimulation (tDCS) and transcranial magnetic stimulation (TMS)) remain unknown. The present study aims to investigate the effects of non-invasive brain stimulation, alone or associated (tDCS/TMS/tDCS plus TMS) combined with language therapy delivered during a period of 6 months, in the progression of language impairment in PPA, compared with sham stimulation combined with language therapy.Methods and analysisThe study is a randomised, double-blinded, parallel, sham-controlled clinical trial. Patients with PPA in early stages (global Clinical Dementia Rating equal to or less than 1) are eligible. They are to be randomised to one of the four treatment arms of the study (active tDCS-active TMS, active tDCS-sham TMS, sham tDCS-active TMS, sham tDCS-sham TMS). All patients will receive language therapy immediately after each session of NIBS, for 6 months. The primary outcome is the Mini-Linguistic State Examination. The secondary outcomes are naming of trained items, Addenbrooke’s Cognitive Examination, Interview for Deterioration in Daily Living Activities, Clinical Dementia Rating including behaviour and language domains, Neuropsychiatric Inventory and regional brain metabolism. Exploratory substudies will be conducted including blood biomarkers, quantitative electroencephalography and spontaneous speech assessment.Ethics and disseminationThe study is registered (ClinicalTrials.gov: NCT07158216) and approved by the Ethics Committee of the Hospital Clinico San Carlos (code 25/309-IC_P_CE). Patients will be enrolled after signing an informed consent form. Study outcomes will be disseminated through presentations at scientific conferences, publications in peer-reviewed journals and other academic forums.Trial registration numberNCT07158216.
- Research Article
- 10.1007/s11357-026-02310-y
- May 18, 2026
- GeroScience
- Eloïse Da Cunha + 5 more
Geroscience needs biomarkers that capture the progressive decline of integrated biological systems with age. Physical capacity, a direct manifestation of systemic integrity, is a core pillar of biological aging but is typically assessed through discrete clinical tests. Speech production, a complex motor act requiring coordinated respiratory, laryngeal, and articulatory control, shares fundamental physiological pathways with global physical function and may therefore serve as an accessible digital biomarker of aging. In a longitudinal cohort of 464 community-dwelling older adults (mean age 79.6 ± 8.7years), we tested the hypothesis that changes in speech track specific changes in physical capacity. Participants underwent a 3-month adapted physical activity (APA) program. At baseline (T0) and post-intervention (T1), we performed a battery of ten objective physical tests (strength, power, endurance, gait, balance, flexibility, mobility, appendicular lean mass, fatigue) and recorded spontaneous speech during emotional autobiographical recall. Multi-layered acoustic, temporal, and linguistic features were automatically extracted. The longitudinal association was analyzed via Spearman correlations and univariate linear mixed-effects models. The APA intervention induced significant improvements in key physical domains, including mobility, gait speed, handgrip strength, and balance (all p < 0.01). These gains were specifically correlated with concurrent changes in speech features (|ρ| = 0.11-0.22). For instance, greater lower-limb strength correlated with reduced vocal shimmer and lexical diversity, while improved flexibility was associated with a lower spectral centroid and zero-crossing rate, indicating smoother phonation. Linear mixed models confirmed significant within-individual coupling between trajectories of physical function and speech dynamics. Emotional context systematically modulated these associations, revealing different speech-stress signatures under cognitive-affective load. This study provides novel longitudinal evidence that speech is a dynamic digital biomarker of domain-specific physical capacity, reflecting underlying functional integrity. The domain-specific speech-physical coupling suggests that speech analysis can serve as a novel, integrative tool for remote monitoring of aging trajectories and the functional efficacy of interventions targeting the age-related physiological decline.
- Research Article
1
- 10.2196/79411
- May 13, 2026
- JMIR Formative Research
- Jingyu Li + 5 more
BackgroundAlzheimer's disease and related dementias (ADRD) are progressive neurodegenerative conditions where early detection is critical for timely intervention and care planning. However, current diagnostic methods are often inaccessible, costly, and delayed, especially for underserved populations. There is a growing need for scalable, noninvasive tools that can support timely diagnosis. Spontaneous speech contains rich acoustic and linguistic markers that can serve as noninvasive behavioral markers for cognitive decline. Foundation models, pretrained on large-scale audio or text data, generate high-dimensional embeddings that encode rich contextual and acoustic information.ObjectiveThis study benchmarks open-source foundation language and speech models to evaluate their effectiveness in detecting ADRD from spontaneous speech as a potential solution for early, noninvasive, and scalable ADRD detection.MethodsIn this study, we used the Pioneering Research for Early Prediction of Alzheimer’s and Related Dementias EUREKA (PREPARE) Challenge dataset, which consists of audio recordings from over 1600 participants with 3 distinct categories of cognitive decline: healthy control (HC), mild cognitive impairment (MCI), and Alzheimer's disease (AD). We further excluded samples that are non-English, nonspontaneous speech, or of poor quality. Our final samples included 703 (59.13%) HC, 81 (6.81%) MCI, and 405 (34.06%) AD cases. We systematically benchmarked 18 open-source foundation speech and language models to classify cognitive status into 3 categories (HC, MCI, or AD). Post hoc interpretability analysis was performed for the best-performing model using Shapley additive explanations linking high-dimensional embeddings with explainable acoustic and linguistic markers.ResultsWhisper-medium model achieved the highest performance among speech models at 0.731 accuracy and 0.802 area under the curve, while Bidirectional Encoder Representations from Transformers with pause annotation achieved the top accuracy of 0.662 and 0.744 area under the curve among language models. Overall, ADRD detection based on state-of-the-art automatic speech recognition model-generated audio-embeddings outperformed other models, and the inclusion of nonsemantic information, such as pause patterns, consistently improved the classification performance of text-embedding–based models.ConclusionsOur work presents a comprehensive comparative evaluation of state-of-the-art speech and language models for AD and MCI detection on a large, clinically relevant dataset. Embeddings derived from acoustic models, which capture both semantic and acoustic information, show promising performance and highlight the potential for developing a more scalable, noninvasive, and cost-effective early detection tool for ADRD.