Articles published on Lexical density
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
828 Search results
Sort by Recency
- New
- Research Article
- 10.67050/ijee/v15i2/ijee262006
- Jun 30, 2026
- International Journal of English and Education
- Aziza Safarova
Lexical diversity is quite significant in reflecting language proficiency in high-stakes English tests that depict whether a candidate is capable of knowing and using various and appropriate vocabulary. However, there is very little comparative linguistic research on investigating the differences in lexical diversity of the major proficiency tests. This research fills this gap with a comparison and analysis of lexical diversity measurements of IELTS and TOEFL answers. It involves a corpus-based methodology, where with the help of computational tools, Type-Token Ratio (TTR), Measure of Textual Lexical Diversity (MTLD), Hypergeometric Distribution D (HD-D), and VOCD are measured. Results show that IELTS answers are more lexical than TOEFL answers, which are more controlled lexical due to integrated task designs. These results highlight the impacts of the test design on the production of language. Language test evaluation studies also utilize the research paper to offer information on how to improve test scores, scoring rubrics, and how a pedagogical approach to English language learners may be provided in the high-stakes testing environment.
- New
- Research Article
- 10.1038/s41598-026-59616-2
- Jun 23, 2026
- Scientific reports
- Jui-Hung Liu + 2 more
Retrieval-Augmented Generation (RAG) systems for specialized industrial domains face a persistent cold-start problem when high-quality, target-domain training data is scarce. We investigate cross-source transfer in this setting using wind-turbine technical documentation as an industrial testbed. A 3 × 4 factorial experiment over twelve embedding-model configurations reveals an asymmetric transfer pattern: embeddings fine-tuned on a semantically rich bilingual corpus from an alternative manufacturer outperform domain-matched models on the target manufacturer's own test set (MRR 0.7902 vs. 0.7012). A seven-category semantic analysis shows the effect is largest for safety-critical queries (+ 48.2%; n = 21, post-hoc power 0.91), and we report a data-dilution effect in which naive multi-source aggregation degrades performance. Because the two corpora differ along several correlated axes - monolingual lexical density, raw volume, and cross-lingual breadth - we characterize semantic richness with independent, pre-specified metrics (MATTR, MTLD, hapax, Jaccard) and present it as a multi-dimensional, testable construct rather than a single factor. To probe robustness, we evaluated five multi-source mixing protocols (naive concatenation, balanced down-sampling, inverse-frequency oversampling, similarity-based quality filtering, and K-means diversity-aware sampling) against three single-brand baselines on a fourth-manufacturer zero-shot test set (Test D, n = 99). No mixing protocol significantly outperformed the strongest single-brand fine-tune, confirming that the asymmetric transfer effect is robust to data-combination strategy and extends to out-of-distribution generalization. We interpret the observed asymmetry as an empirical pattern within this two-manufacturer benchmark whose broader applicability to other industrial domains remains to be validated.
- Research Article
- 10.1044/2026_lshss-25-00236
- Jun 11, 2026
- Language, speech, and hearing services in schools
- Ashley Ippolito + 1 more
Traditional writing assessments often prioritize surface-level correctness, limiting insight into students' underlying linguistic development. This research note introduces ScriptSense, a web-based linguistic analysis platform designed to provide accessible descriptions of student writing. Grounded in the Morphological Pathways Framework and the Lexical Quality Hypothesis, ScriptSense analyzes written language for indicators of morphological complexity, lexical diversity, and syntactic development. Using a cross-sectional pilot design, ScriptSense was applied to 50 third-grade narrative writing samples collected in general education classrooms. Automated analyses extracted descriptive indices of morphological use (e.g., inflectional and derivational morphemes), lexical diversity (e.g., type-token ratio [TTR], lexical density), and syntactic structure (e.g., clause-to-sentence ratio, mean length of utterance). Outputs were analyzed descriptively, with subset verification to support interpretive consistency. Students produced an average of 56 words per narrative (SD = 21), with moderate lexical diversity (M TTR = 0.68) and lexical density (M = 0.59). All students demonstrated productive use of inflectional morphology, and over two thirds produced at least one derivational morpheme. Variability was observed across syntactic measures (mean clauses per sentence = 4.30, SD = 0.77), reflecting differences in sentence elaboration beyond rubrics. By emphasizing descriptive linguistic patterns rather than error-based scoring, ScriptSense offers an accessible approach to examining children's written language. Findings illustrate the platform's capacity to generate multidimensional linguistic profiles.
- Research Article
- 10.1044/2026_jslhr-25-00499
- Jun 5, 2026
- Journal of speech, language, and hearing research : JSLHR
- Dirk B Den Ouden + 7 more
While aphasia treatment studies commonly use picture-naming performance as an outcome measure, narrative discourse better reflects functional language use. Discourse variables may also hold prognostic value for naming treatment response, but their predictive role remains underexplored. We analyzed baseline and posttreatment narrative discourse samples from 95 chronic stroke survivors with aphasia enrolled in a lexical retrieval intervention study. Participants received 3 weeks each of phonological and semantic naming therapy in a crossover design. The Cinderella story retells were analyzed for a range of discourse features: mean length of utterance, words per minute, verbs per utterance, propositional density, type-token ratio, core lexicon, main concepts analysis, and error ratios. We used univariate (generalized) binomial and linear mixed-effects modeling with multiple predictors to assess whether baseline discourse variables predicted naming gains on the Philadelphia Naming Test and whether discourse variables themselves changed following treatment. Higher aphasia severity and lower baseline propositional density predicted greater naming gains across time points and treatment types. Gains after phonological therapy were also predicted by higher baseline core lexicon production. Posttreatment discourse showed gains in mean length of utterance, words per minute, and core lexicon, which were maintained at 6 months, while phonological errors declined only at follow-up. Phonological treatment led to increases in words per minute. Discourse variables reflecting propositional efficiency and lexical appropriateness uniquely predict treatment response beyond general aphasia severity. Lexical retrieval therapy generalizes to improvements in narrative discourse, underscoring the clinical value of incorporating discourse-level measures in aphasia assessment and treatment outcome evaluation. https://doi.org/10.23641/asha.32568822.
- Research Article
- 10.1080/10494820.2026.2670531
- May 23, 2026
- Interactive Learning Environments
- Liangpeng Gao + 4 more
ABSTRACT Instructional language is central to knowledge construction in online learning, but its syntactic structure remains insufficiently quantified. This study developed a dependency-grammar-based framework to examine how linguistic diversity and complexity relate to instructional effectiveness in Massive Open Online Courses (MOOCs). A stratified random sample of 210 sessions from six transportation-engineering MOOCs was analyzed, including national-level “Excellent Courses” and standard courses. Audio recordings were transcribed with a speech recognition API and manually verified. HanLP was used to extract dependency relations and linguistic indicators, including Type-Token Ratio, dependency distance, word-frequency statistics, speech rate, and pause ratio. Binary logistic regression used national-level “Excellent Course” status as a proxy for instructional effectiveness. Greater linguistic diversity, especially higher proportions of nominal subjects, adjectival modifiers, and determiners, was associated with higher course quality. Excessive lexical and syntactic complexity, reflected by higher mean word frequency and longer dependency distances, was associated with lower effectiveness. Fewer prosodic pauses, moderate speech rate, and shorter instructional units were also linked to better MOOC performance. The framework provides an objective and scalable approach for evaluating and optimizing instructional language in online higher education.
- Research Article
- 10.1080/0907676x.2026.2665195
- May 19, 2026
- Perspectives
- Xinlei Jiang + 1 more
ABSTRACT Dependency distance, the linear distance between two syntactically related words, has shown promise in measuring syntactic complexity and processing load in interlanguage research, psycholinguistics, translation studies, and into-A interpreting. Drawing on a large-scale learner corpus of Chinese – English consecutive interpreting, the present study explored the potential of dependency distance as an indicator for cognitive load in learners’ into-B interpreting. A negative binomial regression was conducted to examine the relationship between cognitive load (measured by the frequency of disfluencies) and both dependency-based and traditional linguistic complexity indices (i.e. source-text mean dependency distance, target-text mean dependency distance, source-text mean clause length, target-text mean clause length, source-text lexical density and percentage of numerals). The results showed that: (1) longer target-text mean dependency distance significantly increased disfluency frequency, whereas source-text mean dependency distance did not; (2) mean clause length in both source and target texts had no significant effect; and (3) percentage of numerals was reaffirmed as significant predictors of disfluencies. These findings provide preliminary evidence that dependency distance is a more sensitive indicator for cognitive load in learners’ into-B interpreting than traditional syntactic complexity measures, aligning with decay – and interference-based accounts of dependency distance rather than expectation-based facilitation.
- Research Article
- 10.1016/j.cortex.2026.05.007
- May 15, 2026
- Cortex; a journal devoted to the study of the nervous system and behavior
- Amélie Brisebois + 5 more
Transcription-less analysis of five discourse tasks in Laurentian French persons with post-stroke aphasia: Adaptation and reliability.
- Research Article
- 10.36941/jesr-2026-0351
- May 5, 2026
- Journal of Educational and Social Research
- Teti Sobari + 1 more
The purpose of this study was to investigate the impact of GenAI-Automatic Corrective Feedback (ACF) integration in collaborative writing based on metacognitive instruction on argumentative essay writing skills. The method used in this study was a quasi-experimental study involving 250 students divided into two groups: an experimental group and a control group. The experimental group received the GenAI-Automatic Corrective Feedback (ACF) integration intervention in collaborative writing based on metacognitive instruction, while the control group used a genre-based writing approach. The instruments used included writing assessments covering lexical, accuracy, and fluency. Data analysis used ANOVA and regression analysis to investigate the impact of the intervention on writing skills. The research findings showed that the GenAI-Automatic Corrective Feedback (ACF) intervention in collaborative writing based on metacognitive instruction significantly improved argumentative essay writing skills compared to the genre-based writing intervention. Improvements in argumentative essay writing skills in the experimental group were evident in several aspects, including lexical complexity (lexical density, lexical sophistication, and lexical variety), accuracy, and fluency. Increased lexical complexity was evident in the use of complex vocabulary, complex sentence combinations, and more varied lexical use. Improved argumentative essay accuracy was evident in the accuracy of sentences and paragraphs and supporting ideas relevant to the problem of the argumentative essay. Improved argumentative essay fluency was evident in the use of sentences and grammar with minimal errors and no discordant sentences. The components of metacognitive instruction that significantly contributed to argumentative essay writing skills were metacognitive knowledge, including declarative, procedural, and conditional knowledge, and metacognitive regulation, including planning, monitoring, and evaluation, as well as information management and debugging strategies. Received: 09 December 2025 / Accepted: 13 April 2026 / Published: May 2026
- Research Article
- 10.1016/j.psychres.2026.117215
- May 1, 2026
- Psychiatry research
- Meltem Çınar Bozdağ + 7 more
Spoken language biomarkers in Turkish-speaking schizophrenia patients: Evidence from linguistic analysis and word embeddings.
- Research Article
- 10.26803/ijlter.25.4.22
- Apr 30, 2026
- International Journal of Learning, Teaching and Educational Research
- Librilianti Kurnia Yuki + 1 more
Most students struggle to write argumentative essays due to a lack of schemata, so resources are needed to enrich them. The purpose of this study was to investigate the impact of SVVR-assisted immersive teaching using a flipped classroom model on argumentative essay writing skills, class engagement, and student perceptions. A quasi-experimental method was used, involving 200 university students majoring in Indonesian language education. The participants were divided into two groups with the same number of 100 students each: the experimental group received the intervention of spherical video virtual reality in a flipped classroom model, while the control group received conventional teaching. The instruments used were a rubric for assessing argumentative writing skills, a class engagement questionnaire, a rubric for assessing the lexical complexity of argumentative essays and interview questions. The data analysis included the Wilcoxon signed-rank test, the Quade test, paired-sample t-test, and one-way analysis of covariance (ANCOVA). The findings indicated that SVVR-assisted immersive teaching with a flipped classroom model improved argumentative essay writing skills, class engagement, and student perceptions. Improved argumentative writing skills were evident in the use of evidence and data to support arguments, which were presented more robustly and scientifically, resulting from the observation of objects. Improvements in the quality of argumentative essays were also evident in the increased lexical complexity across all aspects of lexical density, lexical sophistication, and lexical variety. Increased class engagement was evident in cognitive engagement, behavioral engagement, and emotional engagement. Positive perceptions were evident in the experience, emotion, active motivation, and strategies used for improving writing learning. Thus, improvements in all competencies occurred because the intervention enhanced realistic and context-rich experiences, critical thinking skills, and evidence-based reasoning. This research suggests that the use of virtual reality technology can enhance the students' understanding of difficult concepts and create positive impressions of the learning process.
- Research Article
- 10.1044/2026_jslhr-25-00282
- Apr 22, 2026
- Journal of speech, language, and hearing research : JSLHR
- Merve Savaş + 1 more
This study examined how bilingualism influences early linguistic and pragmatic alterations in idiopathic Parkinson's disease (IPD), integrating group-based and factorial analyses to identify early communicative markers. Sixty-five participants (13 bilingual IPD, 14 monolingual IPD, 14 bilingual healthy, 24 monolingual healthy) produced Turkish narratives based on Frog, Where Are You? Group comparisons (Kruskal-Wallis H and Mann-Whitney U tests) were performed across four groups for microstructural indices (mean length of utterance in morphemes [MLU-M], type-token ratio [TTR], morphological errors, verbal fragmentations) and pragmatic markers (enrichment, exclamation, uncertainty, metaphor, emotional terms). Supplementary 2 × 2 factorial analyses (disease: IPD vs. healthy; bilingualism: bilingual vs. monolingual) were conducted to examine main and interaction effects, with acoustic parameters (fundamental frequency [F0] and intensity ranges) included for prosodic evaluation. Group comparisons revealed that bilingual IPD speakers exhibited the lowest MLU-M (p = .012), highest morphological error (p = .036), and greatest verbal fragmentation (p < .001). Pragmatically, they produced fewer enrichment expressions (p = .041) but more exclamations than monolingual IPD participants (p < .001). Acoustic analysis showed reduced but still broader F0 and intensity ranges in bilingual IPD speakers relative to monolingual IPD speakers (p = .012, p = .047). The 2 × 2 factorial analysis confirmed significant main effects of disease on MLU-M and TTR (p < .05) and Disease × Bilingualism interactions for morphological errors and enrichment (p < .05), demonstrating that bilingualism amplified morphosyntactic instability but mitigated prosodic flattening. Early-stage IPD involves concurrent microstructural and pragmatic decline, with bilingualism exerting both protective and burdening effects. Crucially, the reduction of enrichment expressions (p < .05) emerged as an early and sensitive indicator of pragmatic deterioration in bilingual Parkinson's disease, linking executive-control demands with sociopragmatic incompleteness. Discourse-level analyses combining group-based and factorial approaches thus provide a refined framework for identifying subclinical linguistic-pragmatic changes beyond conventional motor or lexical measures. https://doi.org/10.23641/asha.31999344.
- Research Article
- 10.54613/ku.v18i.1584
- Apr 17, 2026
- QO‘QON UNIVERSITETI XABARNOMASI
- Mukhammadrakhimkhon Juraev
This article examines the interplay between qualitative and quantitative measures of text complexity and their impact on reading comprehension and learner engagement in Indo-European language education. While quantitative metrics (e.g., readability scores, lexical density, sentence length) provide statistical indicators of textual difficulty, qualitative dimensions (e.g., thematic richness, syntactic variety, narrative depth, cultural relevance) capture cognitive and affective factors crucial for meaningful comprehension. Drawing on recent empirical studies, the paper demonstrates that relying solely on one type of measure often misaligns with learners' cognitive capacities and motivational profiles. An integrated approach that balances statistical indicators with qualitative contextualization proves more effective in scaffolding instruction, differentiating materials according to proficiency levels, and fostering deeper engagement. The discussion highlights how learner proficiency mediates interaction with text complexity, emphasizing the need for adaptive pedagogical frameworks. Practical implications for curriculum design, text selection, and differentiated instruction are outlined, alongside recommendations for future longitudinal and cross-linguistic research. Ultimately, the synthesis underscores that a holistic evaluation of text complexity is essential for optimizing reading instruction and supporting successful second language acquisition across Indo-European linguistic contexts.
- Research Article
- 10.1044/2026_persp-25-00204
- Apr 15, 2026
- Perspectives of the ASHA Special Interest Groups
- Jill R Potratz
Purpose: This study examined the differences between two language sample types, conversation and narrative, to determine whether there is an optimal type for assessing bilingual school-age children's syntactic complexity and lexical diversity with language sample analysis (LSA). Method: Conversation and narrative language samples were elicited from 24 typically developing school-age children divided into three groups: bilingual Spanish–English, bilingual Mam–English (Mam is an indigenous Mayan language), and monolingual English. Two measures of syntactic complexity (mean length of utterance in morphemes [MLUm] and clausal density [CD]) and two measures of lexical diversity (number of different words [NDW] and moving-average type–token ratio [MATTR]) were calculated to compare different language sample types. Results: All four LSA measures differed significantly between the two language sample types. The measures of syntactic complexity (MLUm and CD) were significantly greater in narratives compared to conversational samples, whereas the measures of lexical diversity (NDW and MATTR) were significantly greater in conversational samples compared to narratives. Group differences were found in both language sample types. Conclusions: The optimal language sample context for bilingual school-age children may best be determined by the suspected area of difficulty for each child. However, a clinician may choose to include both language sample types in their evaluation for a wide variety of evidence. Future research should examine other language sample types, such as play-based and expository samples.
- Research Article
- 10.5811/westjem.64056
- Mar 25, 2026
- Western Journal of Emergency Medicine
- Ryan Mckillip + 4 more
Background: Residency leaders increasingly rely on personal statements to select candidates.The availability of artificial intelligence (AI) writing tools raises concerns that personal statements may reflect AI-generated writing rather than authentic applicant voices.Objective: Assess the prevalence and impact of AIgenerated writing in EM residency personal statements submitted for the 2024 application cycle.Methods: This retrospective study analyzed personal statements submitted to the EM residency of a large academic medical center from 2017 to 2024.The primary outcome was the prevalence of 27 AI-associated target words identified in prior research, or 12 control words, compared between 2024 and 2017-2023 (pre-widespread release of AI writing tools) using one-sample t tests.Secondary outcomes included complexity (Flesch Reading Ease, word count), lexical diversity (type-token ratio), and personalization (first-and third-person pronoun frequency).Results: A total of 8,617 statements were studied (7,803 pre-2024, 814 in 2024).The proportion of statements with AI-associated words increased significantly from pre-2024 to 2024 (22.9% vs. 33.2%,P<0.001) (Figure 1).Control words were unchanged (84.4% vs. 84.3%,P=0.720).Words with the most significant absolute increases were "pivotal" (2.2% to 8.5%), "underscore" (0.5% to 3.9%), and "invaluable" (6.9% to 8.9%) (Figure 2).Word count decreased (686.5 vs. 674.3words, P=0.005).Flesch Reading Ease decreased (43.9 vs. 41.9,P<0.001) but remained at the college level.Type-token ratio increased (0.487 vs. 0.500, P<0.001), suggesting greater
- Research Article
- 10.1515/applirev-2025-0007
- Mar 19, 2026
- Applied Linguistics Review
- Xinye Zhang + 5 more
Abstract Recent studies with automatic text analyzers have explored linguistic measures for predicting writing quality, mostly in English texts by diverse learners. However, research on L2 writing in non-alphabetic languages among students with varied L1 backgrounds remains scarce. This study examines how lexical and syntactic complexity affect writing quality in Chinese-as-a-Second-Language (CSL) students in Hong Kong, using 340 samples from 115 secondary school students with diverse L1 backgrounds. Linear mixed-effects analysis reveals that linguistic indices, including lexical richness and syntactic complexity serve as strong predictors of writing quality, with the combination of logarithmic Type-Token Ratio (LTTR) and syntactic measures (i.e., noun phrase frequency, tree depth, and coordinate phrase usage) explaining 68.5% of the variance. Error analysis demonstrates that L1 word order significantly influences both linguistic complexity patterns and error distributions, with SVO-L1 students demonstrating superior performance compared to other groups. This study extends understanding of linguistic complexity and writing quality relationships to non-alphabetic L2 languages while highlighting the mediating role of L1 typological features in shaping measurable aspects of CSL writing development. Theoretical and pedagogical implications are discussed.
- Research Article
- 10.1057/s41599-026-06878-w
- Mar 15, 2026
- Humanities and Social Sciences Communications
- Liwei Yang + 1 more
The English translation of Chinese classics is crucial for cultural dissemination and East-West exchanges, but assessing their readability poses challenges due to linguistic and cultural differences. This study applies natural language processing and machine learning to evaluate the readability of five English translations of The Analects by D. C. Lau, James Legge, William Jennings, Edward Slingerland, and Burton Watson. Using a corpus with 4236 lines for model training and 2824 held-out lines for application and reporting, 114 readability features were analyzed. Two machine learning techniques, XGBoost and BP Neural Network, were employed to develop customized readability models. The results, evaluated using mean squared error (MSE) and R², with Pearson correlations reported for feature screening and descriptive analysis, reveal that both models achieve high predictive performance, with the BP Neural Network slightly outperforming XGBoost. Key factors influencing readability include sentence length, syntactic structure, lexical complexity, and lexical density (used as a proxy for information density). This study offers a novel method for assessing and improving the readability of multiple English translations of The Analects. While the approach may be potentially transferable to other Chinese classics, such generalization requires further validation.
- Research Article
- 10.18573/jcads.189
- Mar 12, 2026
- Journal of Corpora and Discourse Studies
- Justine Smith
Public trust in journalism is waning, yet research on disinformation has focused predominantly on non-institutional sources such as social media or partisan websites. Far less is known about how deception can emerge within mainstream newswriting that outwardly adheres to professional norms. This study addresses that gap through a corpus-assisted discourse analysis of fabricated and verified reporting by Stephen Glass, a former journalist for The New Republic whose fabrications were exposed in 1998. A purpose-built, matched-author corpus of twenty-three articles (twelve fabricated, eleven verified) was examined using frequency profiling, keyness, collocation, and concordance analysis informed by Appraisal Theory. The analysis identifies systematic linguistic contrasts between fabricated and verified journalism: fabricated texts display lower lexical density, heavier use of verbs and pronouns, and a more personalised narrative style. Evaluative items such as most, just, and not occur more frequently and perform rhetorical functions of emphasis, mitigation, and denial. Collocational and attributional evidence shows that fabricated articles embed stance more often in the journalist’s own voice, projecting confidence and sincerity while limiting alternative readings. By holding author, outlet, and register constant, the study isolates linguistic traces of deception from broader stylistic variation. Methodologically, it demonstrates how corpus tools can be integrated with discourse analysis to reveal how deception is enacted through patterned use of grammatical and evaluative resources. The findings contribute to ongoing work on disinformation by showing that credibility in fabricated journalism is linguistically performed rather than merely asserted, with implications for media literacy and computational detection of deceptive news.
- Research Article
- 10.32388/usaktb.3
- Feb 26, 2026
- Qeios
- Gabriel Ferraz Ferreira + 8 more
PURPOSE: This study aimed to assess bibliometric trends in orthopaedic research before and after the public release of ChatGPT. METHODS: A bibliometric analysis was conducted using PubMed data from January 2021 to December 2025, encompassing articles from ten high-impact orthopaedic journals. A seven-month washout period (December 2022 – June 2023) was applied to account for the typical lag between manuscript preparation and publication. Trends in daily publication frequency, number of co-authors per article, sentence length, lexical diversity, and sentiment were compared between the pre-ChatGPT period (January 2021 – June 2023) and the post-adoption period (July 2023 – December 2025). RESULTS: A total of 21,524 articles were analysed (10,403 before; 11,121 after). The mean number of publications per day increased from 11.42 ± 7.32 to 12.15 ± 8.26 (p = 0.044). After adjusting for monthly seasonality, the difference remained significant (adjusted increase: +1.15 publications/day; p = 0.002). The mean number of authors per article rose from 5.94 ± 3.95 to 6.28 ± 4.39 (p < 0.001). The average sentence length decreased from 14.81 ± 5.08 to 14.52 ± 5.06 words per sentence (p < 0.001), while lexical diversity increased (Type-Token Ratio: 0.51 ± 0.07 to 0.52 ± 0.07; p < 0.001). Mean sentiment scores rose from 3.28 ± 3.27 to 3.64 ± 3.47 (p < 0.001). CONCLUSION: Following the public release of ChatGPT, orthopaedic publications have exhibited a measurable rise in daily output, modest increases in co-authorship, and subtle changes in linguistic style. These temporal associations, while not evidence of causality, correspond with the widespread adoption of AI-assisted writing tools and warrant ongoing evaluation.
- Research Article
- 10.3390/educsci16030365
- Feb 26, 2026
- Education Sciences
- Nali Borrego Ramírez + 3 more
This study examines the relationship between writing performance, assessed with the Early Writing Alert System (SISAT), and linguistic patterns in student narratives from one public and one private university in northeastern Mexico. Variables such as lexical density and richness, text volume, and thematic progression were analyzed to explore how institutional context influences narrative writing and its assessment. A non-experimental, descriptive–comparative design with interpretive triangulation was employed. The corpus comprised 148 narratives produced over three academic periods, analyzed using automated linguistic tools alongside SISAT scores. Descriptive statistics, Spearman correlations, and Kruskal–Wallis tests were applied to examine differences between the two institutions and across periods. The results indicate intermediate performance at both universities, with differentiated patterns: at the public university, lexical richness and density positively correlated with SISAT scores, while greater text volume was negatively associated; at the private university, both text length and diversity were positively related, though excessive lexical density appeared counterproductive. No statistically significant differences were observed between periods or between the two universities. Our findings highlight that quantitative linguistic indicators complement normative assessment and underscore the role of institutional context in writing development. The study also emphasizes the formative and expressive functions of narrative writing, supporting pedagogical strategies that integrate automated assessment with qualitative analysis to foster self-regulation, symbolic expression, and ethical reflection.
- Research Article
- 10.33140/jepr.08.01.06
- Feb 19, 2026
- Journal of Educational & Psychological Research
- Richard Lamb + 3 more
Purpose: This study investigated whether real-time adaptation of science text complexity using neurocognitive data can improve reading comprehension and performance for students with dyslexia. Materials and Methods: One hundred participants (50 with dyslexia, 50 neurotypical) completed science reading tasks while functional near-infrared spectroscopy (fNIRS) recorded hemodynamic responses. A deep Convolutional Neural Network (CNN) classified cognitive demand into high, moderate, or low levels. Textual features such as reading level, lexical density, and complexity were dynamically adjusted based on these classifications. Results: The CNN achieved an accuracy of 0.86 in classifying cognitive demand. Adaptive text adjustments significantly improved comprehension scores for students with dyslexia (55% to 75%) and test performance (60% to 78%) (p < .001). Neurotypical students showed modest gains. The approach demonstrated that real-time adaptation based on cognitive load can reduce overload and enhance accessibility. Conclusions: Integrating neurocognitive data with adaptive AI systems offers a promising pathway to personalize science education for students with reading disabilities. This method improves comprehension and performance while supporting inclusive learning environments. Future research should explore multimodal supports and long-term impacts of adaptive AI technologies in diverse educational contexts.