Articles published on Lexical diversity
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
2113 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.bandl.2026.105768
- Jul 1, 2026
- Brain and language
- Sharlene D Newman + 3 more
Constructing and expressing the story world: neural correlates of narrative generation.
- New
- Research Article
- 10.67050/ijee/v15i2/ijee262006
- Jun 30, 2026
- International Journal of English and Education
- Aziza Safarova
Lexical diversity is quite significant in reflecting language proficiency in high-stakes English tests that depict whether a candidate is capable of knowing and using various and appropriate vocabulary. However, there is very little comparative linguistic research on investigating the differences in lexical diversity of the major proficiency tests. This research fills this gap with a comparison and analysis of lexical diversity measurements of IELTS and TOEFL answers. It involves a corpus-based methodology, where with the help of computational tools, Type-Token Ratio (TTR), Measure of Textual Lexical Diversity (MTLD), Hypergeometric Distribution D (HD-D), and VOCD are measured. Results show that IELTS answers are more lexical than TOEFL answers, which are more controlled lexical due to integrated task designs. These results highlight the impacts of the test design on the production of language. Language test evaluation studies also utilize the research paper to offer information on how to improve test scores, scoring rubrics, and how a pedagogical approach to English language learners may be provided in the high-stakes testing environment.
- Research Article
- 10.64898/2026.06.16.26355796
- Jun 18, 2026
- medRxiv : the preprint server for health sciences
- Maxwell Levis + 10 more
Veterans face an elevated risk of suicide compared to the general population, motivating national efforts to develop predictive models that can guide proactive care. Current models used by the U.S. Department of Veterans Affairs (VA) rely primarily on structured electronic health record (EHR) data, though clinical notes contain rich contextual information that can be quantified using natural language processing (NLP) to derive psychosocial variables that may improve risk detection. Machine learning methods, particularly classification and regression trees (CART), can also uncover interactions between clinical and psychosocial variables, enabling identification of patient characteristics that modify suicide risk factors. However, integrating structured and unstructured data presents challenges because NLP features often greatly outnumber traditional clinical variables, potentially biasing interaction discovery. In prior work, we addressed this imbalance by introducing a weighted CART framework that balances structured variables with NLP-derived psychosocial features from semantic lexicons (SÉANCE). While effective, semantic approaches summarize language into predefined constructs and may overlook important lexical variation present in clinical narratives. In this study, we extend that framework by replacing semantic features with a high-dimensional bag-of-words (BoW) representation of clinical notes and by evaluating models across cohorts defined by structured suicide risk stratification (low, medium, high) and varying temporal lookback windows. Using a cohort of 27,241 veterans, we analyzed clinical documentation collected up to 30, 90, or 270 days prior to death (or a matched index date for controls), enabling temporally flexible risk modeling. XGBoost models were trained to balance structured and unstructured features and identify cross-modal interactions between textual and clinical variables. When incorporated into generalized linear models, these interactions improved predictive performance, particularly among low- and medium-risk patients, and substantially reduced the performance gap between interpretable and more complex models. Notably, the BoW representation outperformed our prior semantic index-based approach. Together, these findings demonstrate the utility of interpretable NLP methods for uncovering clinically meaningful interactions between psychosocial and demographic factors in suicide risk and establish a strong benchmark for future deep learning approaches aimed at capturing richer contextual and temporal information from clinical narratives.
- Research Article
- 10.1044/2026_lshss-25-00236
- Jun 11, 2026
- Language, speech, and hearing services in schools
- Ashley Ippolito + 1 more
Traditional writing assessments often prioritize surface-level correctness, limiting insight into students' underlying linguistic development. This research note introduces ScriptSense, a web-based linguistic analysis platform designed to provide accessible descriptions of student writing. Grounded in the Morphological Pathways Framework and the Lexical Quality Hypothesis, ScriptSense analyzes written language for indicators of morphological complexity, lexical diversity, and syntactic development. Using a cross-sectional pilot design, ScriptSense was applied to 50 third-grade narrative writing samples collected in general education classrooms. Automated analyses extracted descriptive indices of morphological use (e.g., inflectional and derivational morphemes), lexical diversity (e.g., type-token ratio [TTR], lexical density), and syntactic structure (e.g., clause-to-sentence ratio, mean length of utterance). Outputs were analyzed descriptively, with subset verification to support interpretive consistency. Students produced an average of 56 words per narrative (SD = 21), with moderate lexical diversity (M TTR = 0.68) and lexical density (M = 0.59). All students demonstrated productive use of inflectional morphology, and over two thirds produced at least one derivational morpheme. Variability was observed across syntactic measures (mean clauses per sentence = 4.30, SD = 0.77), reflecting differences in sentence elaboration beyond rubrics. By emphasizing descriptive linguistic patterns rather than error-based scoring, ScriptSense offers an accessible approach to examining children's written language. Findings illustrate the platform's capacity to generate multidimensional linguistic profiles.
- Research Article
- 10.1093/chidev/aacag097
- Jun 2, 2026
- Child development
- Anna Brown + 5 more
We tested the extent to which mothers' speech contributes to the transmission of family background inequality in education. In 894 families (93.1% White), representative of the full range of Britain's socioeconomic conditions, we quantified mothers' vocabulary sophistication, lexical diversity, and grammatical complexity from 10-min-long audio-recorded interviews. Mothers' vocabulary sophistication significantly predicted children's (49% males) cognition, literacy, and educational achievement from ages 5 to 12 years, accounting for 2%-5% of the variance. After adjusting for mothers' education and household income, these effects reduced to 1% and 2% or became nonsignificant. Our findings suggest vocabulary sophistication contributes only modestly to the transmission of family background inequality in education.
- Research Article
- 10.1016/j.system.2026.104014
- Jun 1, 2026
- System
- Zeynep Köylü
Sojourn contexts have been favored due to authentic exposure to target language through meaningful interactions. Through a cluster analysis, this study explores patterns of written L2 development in advanced learner English during a semester-long study abroad (SA) program in an anglophone country. Addressing the criticism towards drawing conclusions through aggregated group means, the current study also accounts for issues of ergodicity to help clusters of participants with similar language learning behavior emerge in a data-driven approach. Focusing on syntactic and lexical complexity, as well as formulaicity through holistic and quantifiable measures, this study explores the developmental trajectories of 26 tertiary-level L1 Catalan/Spanish bilingual sojourners who produced a diary entry per week throughout their study abroad semester. The results of the analysis yielded two clusters. Cluster-1 improved holistic proficiency, lexical diversity, and formulaicity, while Cluster-2 had greater gains on syntactic complexity and quantifiable formulaic measures. Results indicate that sojourners with higher initial proficiency scores tended to refine lexical and formulaic aspects, whereas those with a lower proficiency indicated stronger development on syntactic properties of their written performance. These results provide insights into written L2 development after reaching a relatively advanced level of proficiency, which is often not assumed to allow for further growth. The study highlights the importance of designing targeted SA programs that accommodate the specific developmental needs of learners based on their initial proficiency levels.
- Research Article
- 10.1016/j.egyr.2026.109236
- Jun 1, 2026
- Energy Reports
- Qinglin Meng + 9 more
Integrated audit intelligence: A multi-task learning model for classification and opinion generation in power systems
- Research Article
1
- 10.1016/j.caeai.2025.100514
- Jun 1, 2026
- Computers and Education: Artificial Intelligence
- Peiwen Huang + 5 more
Despite the growing importance of English oral communication skills, traditional language learning approaches show limited effectiveness in simultaneously addressing psychological barriers and speaking proficiency among college students. While previous studies have explored anxiety reduction or speaking enhancement separately, a significant gap exists in research examining integrated approaches that tackle Public Speaking Anxiety (PSA), Interview Anxiety, and English-speaking proficiency improvement simultaneously. This study investigated whether an AI-integrated VR oral training application could effectively address these interconnected challenges. A quasi-experimental design was employed with 20 English major students from a mid-central university in Taiwan. Participants completed five training sessions using Meta Quest 2 headsets and an AI-integrated VR oral training application providing tailored feedback on pronunciation, grammar, and fluency based on IELTS standards. Pre- and post-intervention assessments utilized validated instruments including the Personal Report of Public Speaking Anxiety (PRPSA) and Measure of Anxiety in Selection Interviews (MASI), alongside comprehensive speaking proficiency measures. Results demonstrated significant improvements in English speaking proficiency, including increased sentence length and word count, with grammatical errors and incomplete sentences decreasing markedly (p < .001). Concurrently, significant reductions in both PRPSA and MASI scores (p < .05) were observed, though lexical diversity showed slight decline. VR-related motion-sickness symptoms were mildly alleviated, and participants' perceived control increased significantly (p < .05), while interest and attention levels remained stable. These findings suggest that AI-integrated VR oral training applications can effectively enhance English speaking proficiency while simultaneously reducing anxiety levels and improving self-efficacy among English learners. The study addresses a critical research gap by demonstrating the potential of integrated technological approaches to tackle multiple barriers to effective English oral communication, offering promising implications for language education and anxiety management in academic contexts.
- Research Article
- 10.1016/j.amper.2026.100258
- Jun 1, 2026
- Ampersand
- Fatemeh Etaat
The aim of this study is to conduct a linguistic analysis of English essays produced by intermediate-level second language (L2) learners compared to those generated by AI, across five linguistic dimensions: lexical richness, syntactic complexity, semantic similarity, discourse cohesion, and surface-level errors. A parallel corpus of 160 essays, 80 AI-generated and 80 learner-written, was collected and analyzed using Natural Language Processing (NLP) techniques. The results revealed that AI essays tend to be longer and syntactically more complex, with significantly higher lexical diversity and greater use of content words. While both types of essays share similar sentiment and cohesion patterns, the AI essays demonstrate more advanced sentence structures and deeper syntactic tree depths. Readability metrics show that the learners’ essays are simpler and more accessible. Error analysis revealed that the human essays contain four times more errors, particularly in spelling and stylistic choices. The study highlights how AI-generated language diverges from learner-produced writing and offers insights into how AI tools can be effectively leveraged to support language development at this proficiency level. • A comparative linguistic analysis of English essays produced by intermediate-level second language (L2) learners • A total of 160 essays, 80 AI-generated and 80 human-written, were analyzed using NLP tools • Focusing on five linguistic dimensions: lexical richness, syntactic complexity, semantic similarity, discourse cohesion, and surface-level errors • AI essays were longer, lexically denser, and syntactically more complex, with deeper parse trees and more frequent use of content words • The human-authored essays demonstrated simpler sentence structures and greater readability, containing 4 times more surface-level errors
- Research Article
- 10.1080/02687038.2026.2680447
- May 31, 2026
- Aphasiology
- Manaswita Dutta + 1 more
ABSTRACT Background Individuals with latent aphasia frequently experience persistent communication challenges despite scoring above the cut-off ontraditional aphasia assessments, often resulting in reduced life participation and lower self-confidence (Armstrong etal. 2013; Cavanaugh & Haley, 2020). Due to their high-level language deficits, people with latent aphasia often do not meet the criteria for aphasia services, resultingin a lack of essential treatment (Richardson et al. 2021). Effectively identifying language impairments in latent aphasia is crucial and discourse analysis has emerged as a key tool, with recent research highlighting a range of linguistic deficits in this population (Fromm et al. 2017; Stark et al. 2025). Continued research should prioritize identification of more sensitive discourse analysis tools and metrics to improve the detection and accurate characterization of language abilities in latent aphasia. Aim This study investigated narrative abilities of individualswith latent aphasia who score within normal limits on the Western Aphasia Battery-Revised (WAB-R; Kertesz, 2007) by comparing their performance to non-brain-injured control participants without aphasia. Specifically, we examined narrative coherence – a macrostructural feature of language – in relation to microlinguistic abilities. Method The study evaluated data from 118 age-, education-, andgender-matched participants from AphasiaBank: 59 with latent aphasia and 59 non-brain-injured control participants without aphasia. Cinderella story narratives of all participants were analyzed for microlinguistic features including productivity,lexical diversity, and verb production. Coherence was assessed using an adapted transcription-free version of the rubric developed by Linnik and colleagues. Group differences in discourse scores were examined using Mann-WhitneyU tests. Additionally, Spearman’s rank correlations were conducted to explore relationships among narrative variables. Results Significant group differences were noted on all microlinguistic measures. Importantly, both participant groups showed marked differences across all four coherence domains, with the latent aphasia group scoring consistently lower than controls. Correlational analyses revealed few, weak associations between microlinguistic and coherence variables in both the latent aphasia and control groups. Conclusions Discourse analysis is a valuable tool for examining language characteristics in latent aphasia. Microlinguistic measures capture structural language deficits; however, they may have limited reliability in distinguishing the subtle impairments associated with latent aphasia. In contrast, coherencevariables more consistently and effectively capture these subtle deficits. In individuals with latent aphasia, breakdowns in coherence are less likely to be explained solely by microlinguistic features. Therefore, a comprehensive approach that integrates both micro- and macro-linguistic measures is necessary to enhance the diagnostic sensitivity of language assessments for latent aphasia.
- Research Article
- 10.1080/09296174.2026.2670857
- May 23, 2026
- Journal of Quantitative Linguistics
- Trudie Strauss + 3 more
ABSTRACT Statistical laws have long been proposed to model word frequency distributions, the most famous being Zipf’s law (the inverse proportionality between a word’s rank and its frequency) and its extensions. This study evaluates whether a class of Zipfian models can adequately capture both the overall frequency spectra and linguistically interpretable properties (such as word repetition rate, vocabulary growth rate, and entropy) across 193 languages. Using a Bayesian framework, seven Zipfian formulations were compared by assessing their overall fit and their predictive performance on these linguistic measures. The results show that, for nearly all languages, at least one formulation provides a statistically plausible fit to the observed frequency spectrum. Yet, even though different models emerged as best fits for different languages, all can be closely approximated by a power-law function for sufficiently large frequency classes, indicating a strong presence of this law across languages. Furthermore, some models that fit the overall spectrum less well still captured individual linguistic measures effectively. A key methodological contribution of this study lies in expressing linguistic measures as posterior predictive distributions rather than point estimates, allowing uncertainty to be represented directly and offering a richer, probabilistic understanding of lexical structure and variation across languages.
- Research Article
- 10.3233/shti260329
- May 21, 2026
- Studies in health technology and informatics
- Baptiste Pras + 1 more
Biomedical Entity Linking (BEL) is essential for structuring knowledge from biomedical texts, yet global evaluation metrics often obscure systematic model weaknesses. We propose a fine-grained evaluation framework that analyzes performance across interpretable mention-level characteristics, including length, lexical variation, synonymy, homonymy, and training frequency. Using the BELB benchmark, we apply this analysis to neural and rule-based systems. Our results show that performance degradation stems from mention-level difficulty, with consistent drops across characteristics that reflect limited training coverage, and expose model weaknesses beyond aggregate scores in a unified benchmark setting.
- Research Article
- 10.1016/j.dib.2026.112875
- May 19, 2026
- Data in Brief
- Mohammad Abuoudeh
An acoustical dataset of storytelling and story reading in Jordanian Arabic
- Research Article
- 10.1007/s11357-026-02310-y
- May 18, 2026
- GeroScience
- Eloïse Da Cunha + 5 more
Geroscience needs biomarkers that capture the progressive decline of integrated biological systems with age. Physical capacity, a direct manifestation of systemic integrity, is a core pillar of biological aging but is typically assessed through discrete clinical tests. Speech production, a complex motor act requiring coordinated respiratory, laryngeal, and articulatory control, shares fundamental physiological pathways with global physical function and may therefore serve as an accessible digital biomarker of aging. In a longitudinal cohort of 464 community-dwelling older adults (mean age 79.6 ± 8.7years), we tested the hypothesis that changes in speech track specific changes in physical capacity. Participants underwent a 3-month adapted physical activity (APA) program. At baseline (T0) and post-intervention (T1), we performed a battery of ten objective physical tests (strength, power, endurance, gait, balance, flexibility, mobility, appendicular lean mass, fatigue) and recorded spontaneous speech during emotional autobiographical recall. Multi-layered acoustic, temporal, and linguistic features were automatically extracted. The longitudinal association was analyzed via Spearman correlations and univariate linear mixed-effects models. The APA intervention induced significant improvements in key physical domains, including mobility, gait speed, handgrip strength, and balance (all p < 0.01). These gains were specifically correlated with concurrent changes in speech features (|ρ| = 0.11-0.22). For instance, greater lower-limb strength correlated with reduced vocal shimmer and lexical diversity, while improved flexibility was associated with a lower spectral centroid and zero-crossing rate, indicating smoother phonation. Linear mixed models confirmed significant within-individual coupling between trajectories of physical function and speech dynamics. Emotional context systematically modulated these associations, revealing different speech-stress signatures under cognitive-affective load. This study provides novel longitudinal evidence that speech is a dynamic digital biomarker of domain-specific physical capacity, reflecting underlying functional integrity. The domain-specific speech-physical coupling suggests that speech analysis can serve as a novel, integrative tool for remote monitoring of aging trajectories and the functional efficacy of interventions targeting the age-related physiological decline.
- Research Article
- 10.69682/arti.2026.93(3).80-85
- May 18, 2026
- Scientific Works
- Səminə Abdullayeva
In the Azerbaijani language, there exists a group of words that do not have broad possibilities of usage in the lexicon. Such words either reflect a certain socio-cultural context or are used in the speech of specific specialists, including artists and professionals. These words belong to the general lexical layer of the language and reflect its rich lexical diversity. Words of this type, which do not acquire an active and general-usage character in the Azerbaijani lexicon, can be considered a group of relatively limited-use vocabulary, as they are still part of the overall lexical system. This group includes archaisms, dialectisms, and terminological vocabulary. Archaisms, within the synchronic layer of the language, have limited usage and reflect historical, social, everyday, and other ethnographic concepts, creating a national and ethnic color. This constitutes the general norm of archaisms. Dialectisms express both historical-diachronic and contemporary synchronic concepts. They also reflect local ethnic features of mentality. A certain part of dialectisms may enter the general vocabulary, as a result of which instability of norm is observed in their local and regional forms, whereas in the general vocabulary stable normative functioning is established. Terms related to specific fields of activity function in accordance with the norm of obligatory speech usage. Those that enter the general lexical fund may acquire a stable norm in both spoken and written language.
- Research Article
- 10.1007/s10339-026-01346-4
- May 9, 2026
- Cognitive processing
- Han Zhang + 1 more
Reading involves a remarkable coordination of perceptual and linguistic processes. This coordination is reflected by a close coupling between reading eye movements and lexical properties of words, such as word frequency and predictability, as well as morpho-syntactic and semantic regularities. Yet, reading is also subject to interference, and a common source of interference is when others are speaking nearby. Then, how does irrelevant speech affect reading as reflected in eye movements? Fifty-nine participants (mean age: 19, SD age: 1.22, 25 females, 34 males, primary language: English) read passages from the PROVO corpus Luke and Christianson (Behav Res Methods 50(2):826-833. https://doi.org/10.3758/s13428-017-0908-4 , 2018) under intelligible irrelevant speech or the silence condition, with eye movements recorded. We investigated how irrelevant speech influences the coupling between eye movements and several indicators of lexical difficulty, including word frequency, cloze probability (exact-word predictability), large language model-generated surprisal, morpho-syntactic predictability (part-of-speech and inflection), and semantic similarity (general meaning). Associations between lexical variables and word looking times remained robust during first-pass reading and were unaffected by irrelevant speech, suggesting intact lexical processing during early stages of reading. However, disruption emerged in later eye-movement measures (i.e., total viewing time), where low-frequency and semantically unpredictable words prompted increased re-reading under irrelevant speech. These results indicate that intelligible irrelevant speech selectively interferes with post-lexical, higher-order processing. The increased re-reading of rare and semantically unpredictable words suggests that readers compensate for speech interference by revisiting these words to better integrate their meanings into a coherent understanding of the text.
- Research Article
- 10.1016/j.jcomdis.2026.106650
- May 7, 2026
- Journal of communication disorders
- Iris Hindi + 1 more
Non-interactive and naturalistic language exposure in autism: An investigation of narrative production in bilingual children.
- Research Article
- 10.1044/2026_jslhr-25-00014
- May 7, 2026
- Journal of speech, language, and hearing research : JSLHR
- Laura Xiaoqian Guo + 1 more
This study investigated decontextualized talk produced by preschoolers with language impairment. Although literature has highlighted the pivotal transition from using "here and now" to using "there and then" language in typically developing children entering preschool age, there is limited understanding of this phenomenon in children with developmental disorders. This study analyzed play-based conversations between four preservice speech-language pathologists (SLPs) and two participant groups: children with autism spectrum disorder (ASD; n = 7) and children with developmental language disorder (DLD; n = 9). Data collection included video-recorded conversation samples and standardized assessments of child receptive language development. For dyadic measures, the study examined group differences in total conversation turns, turn-taking rates, and proportion of decontextualized turns. For child and adult language measures, the study analyzed their speech samples, including mean length of utterance (MLU) in words and number of different words (NDW) per minute, and preservice SLPs' use of language facilitation techniques during play-based conversations. Our findings revealed that dyads from both groups (DLD, ASD) engaged in decontextualized talk (narratives, explanations). Children's standardized language measures correlated with decontextualized talk and conversation turns, although no group differences were observed in these measures. Analysis of preservice SLPs' language behaviors showed equivalent linguistic input (syntactic length, lexical diversity) and comparable use of language facilitation techniques across diagnostic groups. Preservice SLPs' use of repetition showed strong positive correlations with children's immediate language outcomes (MLU in words, NDW per minute). When controlling for both child age and adult facilitation technique use, group differences in dyadic measures remained nonsignificant. This lack of group differences may be attributed to the unique features of language elicited from child-led free-play, a small sample size, and the heterogeneity of language profiles even within two diagnostic groups. Results provide new information about play-based verbal interactions of children with DLD and ASD, suggesting potential for clinicians to incorporate decontextualized talk into interventions. Future studies can examine the effects of decontextualized talk strategies, such as engaging children in narratives and explanations in more structured activities, on their language outcomes.
- Research Article
- 10.3389/fpsyg.2026.1704308
- May 5, 2026
- Frontiers in Psychology
- Cai Wang + 6 more
IntroductionDevelopmental Language Disorder (DLD) is a common neurodevelopmental condition that affects both the structural and pragmatic aspects of language. Conversational abilities—including topic initiation, maintenance, repair, turn-taking, and the integration of non-verbal behaviors—are essential for social communication and peer relationships. While a substantial body of research has described conversational difficulties in DLD among English-speaking populations, there is limited evidence on how these impairments manifest in Mandarin-speaking children. There is a critical need for culturally adapted assessment approaches; thus, this protocol primarily aims to evaluate the feasibility of a multimodal framework tailored specifically to the Mandarin-speaking context.MethodsThis protocol describes an age-stratified cross-sectional observational study with two phases. In Phase 1, an adapted multimodal profiling framework will be developed and piloted through semi-structured free conversation and role-play tasks, recorded with audio–video equipment. Transcription will follow Codes for the Human Analysis of Transcripts (CHAT) conventions within European Distributed Corpora (EUDICO) Linguistic Annotator (ELAN), with word segmentation supported by Jieba and manual correction. Annotation tiers will capture utterance-level features, communicative acts (Inventory of Communicative Acts-Abridged [INCA-A]), repairs, turn-taking, and gesture–speech alignment (stroke apex ±500 ms). Feasibility (session completion, recording length, and utterance yield) and reliability (Cohen’s κ, ICC, and drift checks) will be evaluated to refine the coding manual. In Phase 2, the finalized protocol will be applied to 90 children with DLD and 90 typically developing (TD) peers (aged 4–6, matched for age and sex). Data analysis will quantify primary outcome measures (topic maintenance ratio, successful repair rate, gesture–speech synchrony index) and secondary outcome measures (lexical diversity, initiation–maintenance balance, and turn-taking productivity), with exploratory analyses of demographic and home-environment moderators.DiscussionThis protocol advances methodological research by providing a transparent and reproducible framework for multimodal profiling of pragmatic language in Mandarin-speaking preschoolers. Its strengths include multimodal data integration and explicit feasibility benchmarks. Limitations include the cross-sectional design and a modest sample size. Future research may extend the protocol to longitudinal studies, contributing to the establishment of standardized procedural benchmarks and improved data comparability within this specific linguistic environment.
- Research Article
- 10.36941/jesr-2026-0351
- May 5, 2026
- Journal of Educational and Social Research
- Teti Sobari + 1 more
The purpose of this study was to investigate the impact of GenAI-Automatic Corrective Feedback (ACF) integration in collaborative writing based on metacognitive instruction on argumentative essay writing skills. The method used in this study was a quasi-experimental study involving 250 students divided into two groups: an experimental group and a control group. The experimental group received the GenAI-Automatic Corrective Feedback (ACF) integration intervention in collaborative writing based on metacognitive instruction, while the control group used a genre-based writing approach. The instruments used included writing assessments covering lexical, accuracy, and fluency. Data analysis used ANOVA and regression analysis to investigate the impact of the intervention on writing skills. The research findings showed that the GenAI-Automatic Corrective Feedback (ACF) intervention in collaborative writing based on metacognitive instruction significantly improved argumentative essay writing skills compared to the genre-based writing intervention. Improvements in argumentative essay writing skills in the experimental group were evident in several aspects, including lexical complexity (lexical density, lexical sophistication, and lexical variety), accuracy, and fluency. Increased lexical complexity was evident in the use of complex vocabulary, complex sentence combinations, and more varied lexical use. Improved argumentative essay accuracy was evident in the accuracy of sentences and paragraphs and supporting ideas relevant to the problem of the argumentative essay. Improved argumentative essay fluency was evident in the use of sentences and grammar with minimal errors and no discordant sentences. The components of metacognitive instruction that significantly contributed to argumentative essay writing skills were metacognitive knowledge, including declarative, procedural, and conditional knowledge, and metacognitive regulation, including planning, monitoring, and evaluation, as well as information management and debugging strategies. Received: 09 December 2025 / Accepted: 13 April 2026 / Published: May 2026