Articles published on Corpus Linguistics
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
3548 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.leaqua.2026.101961
- Jul 1, 2026
- The Leadership Quarterly
- Gerlinde Mautner + 2 more
Our aim in this paper is to add to the methodological toolbox of qualitative research in leadership by illustrating how corpus linguistics (CL) can be used to study large-scale datasets comprised of text or talk. CL differs from other approaches to analysing large-scale textual data, such as topic modeling and sentiment analysis, because it enables detailed quantitative and qualitative analysis of linguistic choices on the level of vocabulary and grammar. CL can be used to study text or talk produced by leaders (such as CEO speeches, letters to shareholders, or interviews) or written about leaders (such as newspaper or magazine articles, social media posts, or biographies). Whilst CL methods can be applied using R or Python, here we demonstrate how a user-friendly proprietary software, Sketch Engine, can be used. We illustrate the relative strengths of the method using a corpus of media texts comprising leader profiles published in The Times (UK) newspaper where senior executives (n = 733) answered the question “What does leadership mean to you?”. We conclude by discussing the potential that CL offers for informing future research and theory development, spanning positivist, interpretivist and social constructionist styles of theorizing. We also outline the practical benefits the method offers for improving leadership practice and for people involved in leadership teaching and training by providing robust evidence about concrete and learnable behaviors.
- New
- Research Article
- 10.1136/bmjopen-2026-122001
- Jun 29, 2026
- BMJ open
- Katharine Weetman + 5 more
Good communication is imperative to high-quality patient care, particularly for achieving positive outcomes in healthcare consultations. Evidence about communication in clinical consultations for people with motor neurone disease (MND) is limited, but the knowledge gap prevents possible improvements to communication skills that are unique to the challenges that people living with MND encounter. This research aims to better understand and improve communication experiences for patients with MND. A 24-month qualitative longitudinal study whereby data collection will explore 10-15 cases of persons living with MND. For each case, we will recruit one patient, one close person/carer and one healthcare professional for separate interviews alongside an observed clinical consultation. The patient will be interviewed three times at approximately 4-6-month intervals over 12 months to explore the changing experiences, communication challenges and impact on quality of life. Close persons and healthcare professionals will be interviewed once. Data will be analysed to draw out patterns within and between cases, incorporating thematic analysis, corpus linguistics and conversation analysis techniques. Ethical approval was granted by the Health Research Authority (ref: 26/WM/0026).A co-design workshop will develop a framework for an educational toolkit for healthcare professionals to improve communication skills with MND patients. Findings will be disseminated to the academic community, healthcare staff and the public to increase reach and impact (academic journals, seminars, conferences, newsletters, community groups and engagement events). We will make recommendations for improvements to policy and practice. ISRCTN15571034.
- New
- Research Article
- 10.1186/s40862-026-00412-w
- Jun 22, 2026
- Asian-Pacific Journal of Second and Foreign Language Education
- Wei Lun Wong + 3 more
Abstract Corpus linguistics (CL) now forms the empirical backbone of many artificial intelligence (AI) language technologies. However, the reciprocal influence of these two domains remains critically under-described in studies from emerging regions. To address this gap, this present study offers a global bibliometric map of open-access research situated at the CL–AI intersection between 2020 and mid-2025. It aims to expose thematic evolution, geographical leadership and conceptual blind spots. Adopting a quantitative bibliometric approach, we sequentially combined performance indicators, network visualisations and keyword clustering. The initial Scopus search yielded 3214 records. After applying the inclusion criteria, removing duplicates and screening for relevance and open-access status, the final dataset comprised 162 documents, which were then normalised and analysed through trend counting, co-citation analysis and co-word mapping. Findings reveal a sharp post-2022 surge in output, accompanied by a thematic pivot from earlier data-mining paradigms to transformer-driven studies. China and the United States dominate both productivity and citations. Nevertheless, early citation visibility is emerging from Bangladesh, Egypt and Malaysia. Bridging keywords such as language models and semantics occupy high betweenness positions. It signals conceptual chokepoints where linguistic theory and AI innovation converge. The results imply that transformer architectures reward curated, linguistically representative corpora, while open-access dissemination and South–North collaborations broaden global impact. Funding bodies are therefore encouraged to incentivise cross-regional partnerships and journals to sponsor special issues that unite computational and pedagogical strands. Future research should extend coverage to paywalled databases, incorporate qualitative expert interviews and pursue longitudinal institutional case studies to gauge how synthetic corpora and collaborative funding reshape the evolving CL–AI landscape.
- Research Article
- 10.1080/00220272.2026.2689998
- Jun 20, 2026
- Journal of Curriculum Studies
- Magnus Hultén + 2 more
ABSTRACT While late-twentieth-century educational policy shifts have been widely studied, research has often relied on selective empirical material, privileging policy documents—the stories of the victors—while obscuring messier political processes and broader developments. This study broadens educational policy research by analyzing public discourse, specifically newspapers. It examines the historical emergence and meaning of the Swedish term målstyrning, commonly associated with the introduction of New Public Management in Swedish education. Drawing on conceptual history and corpus linguistics, the study analyses 1,470 newspaper articles using the term. The findings show that målstyrning emerged in Swedish newspaper discourse from the 1960s, coinciding with longer-term shifts towards instrumental and outcome-oriented logics in education and public administration. We argue that målstyrning filled a conceptual void by providing a linguistic and ideological framework for these changes, and that NPM reforms in Swedish education should be understood as part of much longer historical developments. The study highlights the risks of overreliance on policy texts, which may exaggerate the novelty and radicalism of policy change.
- Research Article
- 10.1080/00014788.2026.2672412
- Jun 9, 2026
- Accounting and Business Research
- Renata Stenka + 2 more
Integrated Reporting (IR) is a globally adopted framework, yet little is known about the actual reporting practices it entails. To address the ongoing tension over whether IR genuinely adopts a broader, more inclusive stakeholder orientation or remains primarily, or even exclusively, shareholder-focused, we examine how stakeholders in general, and shareholders in particular, are discursively constructed in best practice integrated reports. We use transitivity analysis, operationalised through Corpus Linguistics (CL) tools, to examine how grammatical and lexical choices construct meaning by revealing who is represented as ‘acting’, or ‘being acted upon’, in the text. We find that stakeholders are portrayed as passive, in need of guidance, yet also appreciative of corporate actions, while shareholders are presented as rights-bearing, and influential. We argue that this convenient construction allows companies to appear responsive to stakeholders while maintaining a primary focus on shareholders, thereby further reinforcing the financialisation of sustainability.
- Research Article
- 10.62536/sjehss.2026.v4.i6.pp16-22
- Jun 7, 2026
- Sciental Journal of Education Humanities and Social Sciences
- Kunduz Ibodullayeva
This study analyzes the linguostatistical and functional features of academic discourse in the Uzbek language using corpus linguistics methods. An electronic corpus, compiled from academic articles in the "Language and Literature Education" journal, served as the research material. During the research, frequency, morphological, and discourse analyses were conducted using the Frequency, Concordance, and Concordance Plot functions of the AntConc software. The study confirms the effectiveness of corpus linguistics methods in describing academic discourse on an objective and empirical basis. It also establishes a theoretical and practical foundation for modeling the academic style in Uzbek and for advancing automated analysis systems for academic texts.
- Research Article
- 10.1007/s10936-026-10239-8
- Jun 1, 2026
- Journal of psycholinguistic research
- Pengfei Bao
This study integrates experimental psycholinguistics and usage-based corpus linguistics to investigate how pragmatic and prosodic cues constrain syntactic parsing in natural discourse. Focusing on the resolution of prepositional phrase (PP) attachment ambiguity (e.g., "I saw the man with the telescope"), we examine the role of information structure. Using the large-scale OpenSubtitles corpus, we manually disambiguated and annotated thousands of instances for the givenness/newness of referents and the presence of commas as orthographic prosodic boundaries. Mixed-effects logistic regression analyses reveal that both information status and commas are robust predictors of attachment preferences. A contextually Given NP within the PP significantly promotes NP-attachment, while a comma preceding the PP strongly favors VP-attachment. These findings extend laboratory-based results by demonstrating the potent role of pragmatic and orthographic prosodic cues in ecologically valid language processing, underscoring the necessity of integrating discourse context into models of the syntax-pragmatics interface.
- Research Article
- 10.21009/stairs.7.1.2
- May 30, 2026
- STAIRS: English Language Education Journal
- Rizdika Mardiana + 5 more
This study examines student discourse in an English classroom at a senior high school in Papua, Indonesia, through an integrated framework combining corpus linguistics and Conversation Analysis (CA). The research aims to (1) characterize students’ spoken output in active learning settings, (2) determine the extent to which their utterances reflect authentic and communicative English use, and (3) provide corpus-informed pedagogical insights tailored to learners’ linguistic needs. Data were collected via classroom observations, audio–video recordings, transcriptions of teacher–student interactions, and semi-structured interviews. A mini spoken corpus was developed to analyze lexical frequency, collocational tendencies, and formulaic expressions, while CA was employed to examine sequential organization, turn-taking, adjacency pairs, and repair mechanisms. The analysis revealed that student utterances were largely short, formulaic, and lexically concentrated around thematic items such as invitation, party, and birthday. Frequent formulaic sequences, including I would like… and I hope you are coming to my birthday party, facilitated learners’ fluency and extended participation. CA findings further indicated that teacher-controlled turn-taking and repair sequences played a crucial role in scaffolding learner output and promoting interactional competence. Overall, the integration of corpus and CA methodologies illuminates the interplay between lexical development and interactional structure in classroom discourse. The study concludes that active learning environments, when coupled with data-driven insights, can enhance both communicative authenticity and linguistic accuracy, offering valuable implications for EFL pedagogy in under-researched contexts such as Papua.
- Research Article
- 10.29063/ajrh2026/v30i9s.2
- May 29, 2026
- African journal of reproductive health
- Licui Zhu + 1 more
This study examines how corpus-grounded artificial intelligence (AI) can strengthen Spanish reproductive health communication capacity within China's digital health ecosystem. A mixed-methods design was employed, combining corpus linguistics, AI-assisted message generation, expert-informed evaluation, and quantitative user assessment. A domain-specific Spanish reproductive health corpus was constructed from 3,052 documents, including clinical guidelines, patient education materials, and FAQs/user queries, yielding 1,065,110 tokens for linguistic analysis and AI grounding. Corpus-derived readability benchmarks, lexical simplification rules, and discourse patterns were integrated into an AI content generation pipeline to produce reproductive health messages, which were then compared with non-corpus-grounded AI outputs. The user evaluation phase was conducted among 240 Spanish-speaking or Spanish-proficient adults in selected Chinese metropolitan cities. Data were collected through a structured questionnaire measuring Corpus-Grounded AI Generation, Spanish Communication Quality, and Health Communication Capacity, and were analyzed using reliability testing, correlation analysis, and structural equation modeling. The findings showed that corpus-grounded AI significantly improved Spanish communication quality, while communication quality had the strongest effect on users' comprehension, confidence, and help-seeking intention. Mediation analysis further demonstrated that Spanish communication quality significantly mediated the relationship between corpus-grounded AI generation and health communication capacity. The study concludes that linguistically informed AI design can enhance the clarity, accessibility, and effectiveness of reproductive health education and offers a practical framework for multilingual digital health communication in sensitive healthcare contexts.
- Research Article
- 10.31861/gph2026.858-859.55-66
- May 24, 2026
- Germanic Philology Journal of Yuriy Fedkovych Chernivtsi National University
- Yuliia Demianchuk
The study investigates the referential–evaluative semantics of English word compounds containing the component “war”. The relevance stems from the need to understand how multi-component units ensure referential precision while shaping evaluative–imagistic perspectives in discourse. The aim is to differentiate the referential–descriptive and evaluative–imagistic layers of meaning in these compounds and identify their speech actualization mechanisms. The methodology integrates semantic–typological analysis, encompassing denotative–referential parameters, compositionality, valency features, taxonomic relations, and pragmatic–connotative markers. The research corpus includes official, publicistic, and literary texts exceeding 15,000 pages, yielding 8,310 target combinations via continuous sampling. Results show that the referential–descriptive layer links units to extralinguistic reality through intensional and extensional features, whereas the evaluative–imagistic layer relies on lexical intensifiers, metaphorical models, legal nominations, and target-oriented syntax. Grammatical collocation models (Adj+N, N+N, N+Prep+N) serve as precision tools to enhance interpretive clarity. The scientific novelty lies in conceptualizing this two-layer semantic organization and parameterizing functions that integrate naming and evaluation. The findings offer practical value for corpus linguistics, terminology, and translation studies.
- Research Article
- 10.1080/0309877x.2026.2670671
- May 22, 2026
- Journal of Further and Higher Education
- Anna Lindroos Cermakova + 1 more
ABSTRACT Universities are operating in a challenging and somewhat paradoxical situation in the domain of early childhood literacy education. Against a background of the very limited visibility of university research in the formulation of new policies and practices, there is a great deal of public discourse to which university researchers have contributed. Utilising corpus linguistics methods, the study explores two datasets collected in 2017–2022 from UK newspapers and Twitter, highlighting how university expertise is operationalised in these two types of public discourses on primary literacy. Findings show that while universities are referenced as key sources of expertise in both domains, their representation differs. Newspaper media frequently name specific universities as sources of research and authority, with some universities being cited relatively frequently. Selection and representation emerge from perceived alignments to the newspapers’ own news values and audiences. In Twitter social media interactions, universities play a broader role, with more involvement from overseas HEIs, and a broader range of content and style. There is a separation between academia and practitioners, suggesting limited influence and reach of academic Twitter. Highly effective academic institutional accounts communicate through varied strategies, sometimes heavily reliant on specific academic influencers. Our findings imply that there are opportunities for individuals, in universities or as alumni, and universities as a whole, to devise effective communication strategies in the domain of primary literacy education research, paying attention to both professional news and social media. This is likely to apply to other areas of university expertise.
- Research Article
- 10.1075/prag.25004.bou
- May 5, 2026
- Pragmatics
- Basma Bouziri + 1 more
Abstract Metadiscourse has been a major focus of research over the last twenty-five years, attracting methodological approaches from the areas of textual pragmatics and discourse studies, many of which are supported by corpus linguistics. A major challenge in corpus-based discourse studies, however, is subjectivity, which may affect their quality and undermine their methodological rigor. To reduce subjectivity and guarantee consistency, assessing reliability of coding is essential. This study advocates combining quantitative with qualitative approaches to reliability. We argue that this mixed-method approach will provide a better assessment of reliability. To this aim, this methodological synthesis surveyed research covering empirical corpus-based studies on metadiscourse published in indexed and peer-reviewed journals. One major finding is that most studies did not report conducting any reliability measure. Issues in reliability accounts were also identified for those that did. Another major finding is a pervasive lack of transparency and comprehensiveness in reliability reports. Recommendations for enhancing reliability are listed.
- Research Article
- 10.22158/eltls.v8n3p1
- May 4, 2026
- English Language Teaching and Linguistics Studies
- Shuo Geng
In the new media era, public educational discourse has tended toward popularization, with Zhang Xuefeng’s educational discourse serving as a typical representative. From the perspective of corpus linguistics, this study constructs two corpora: the Zhang Xuefeng Educational Corpus (ZXC) and the Expert Educational Corpus (EEC) of university scholars. With AntConc, NeoSCA and SPSS adopted, a comparative analysis is conducted on lexical richness and syntactic complexity. The results indicate that there is no significant difference in sentence length between the two types of discourse, while extremely significant differences exist in lexical richness, subordinate structures, coordinate structures and phrasal structures. Overall, Zhang Xuefeng’s discourse tends to be simplified and popular, whereas expert discourse is more rigorous and standardized. This study reveals the differentiation law of popularization and specialization in new media educational discourse, interprets the adaptive value of simplified expression in mass communication, and provides empirical references for optimizing the dissemination of public educational information and balancing professionalism and accessibility.
- Research Article
- 10.1080/19331681.2026.2652900
- May 1, 2026
- Journal of Information Technology & Politics
- Alessia D’Andrea + 3 more
ABSTRACT Disinformation plays a key role in contemporary hybrid warfare, particularly in the Russia – Ukraine conflict, where media narratives shape public perception. This study examines whether established linguistic indicators of disinformation appear in pro-Kremlin news discourse and identifies context-specific strategies. A corpus of English-language articles from the EUvsDisinfo database (2021–2023) was analyzed using Critical Discourse Analysis and corpus linguistics, integrating quantitative measures with qualitative interpretation. Critical Discourse Analysis allows identifying the features that express meaning, reveal ideologies, and influence discourse. On the other hand, corpus linguistics provides a quantitative basis by examining the frequency and the distribution of these features in the corpus. The findings confirm the presence of known disinformation markers designed to influence the reader’s perception through text, rhetoric, and semantic manipulation mechanisms. Disinformation thus emerges as a coherent, context-driven discursive practice rather than a fixed set of features.
- Research Article
- 10.32603/2412-8562-2026-12-2-200-210
- Apr 24, 2026
- Discourse
- I V Kononova + 1 more
Introduction . The article examines the evolution of phatic communication means in Englishlanguage film dialogue over nearly ninety years (1930–2018). The relevance of the study is determined by the need to investigate the dynamics of discursive practices reflecting socio-cultural changes in society. The scientific novelty lies in the application of diachronic corpus analysis to phatic markers in film speech to identify trends towards colloquialization and democratization of communicative norms conveyed through mass art. The aim of the research is to identify quantitative and qualitative changes in the use of phatic units based on the Movies corpus and to interpret them in the context of the evolution of communicative practices. Methodology and sources . The research is carried out within the framework of diachronic discourse studies and corpus linguistics. The material is the Movies corpus, containing dialogues from English-language films of 1930–2018 (total volume about 200 million words). The analysis was conducted in three stages: 1) compiling a list of phatic markers based on preliminary film viewing and theoretical works; 2) calculating the normalized frequency of the selected units by decade, followed by consolidation into three periods (1930–1960, 1961–1990, 1991–2018) and computing the weighted average frequency taking into account the size of subcorpora; 3) interpreting the obtained data. Results and discussions. The analysis of the dynamics of phatic markers revealed a steady trend towards colloquialization and democratization of film speech. This is most clearly manifested in the change of dominant greeting and farewell formulas: formal good morning and goodbye consistently give way to informal hi and bye . In the group of markers reflecting inquiries about the interlocutor's state, the ritualized formula how do you do? is being replaced by more colloquial equivalents how are you doing and what's up? . Discourse markers and fillers imitating speech spontaneity show a significant increase. Politeness formulas consistently decline, giving way to shorter and more direct analogues. The obtained data indicate that film dialogue gradually moves away from theatrical conventionality and approaches live conversational speech, reflecting real language changes and the transformation of communicative norms in society. Conclusion . The conducted research confirmed that phatic communication in Englishlanguage film dialogue has undergone significant evolution. From the formal, ritualized structures of the 1930s–1960s, film speech has shifted towards colloquial forms imitating spontaneous oral communication. Starting from the 1960s, a decline in the frequency of etiquette clichés and an increase in discourse markers, fillers, and reduced forms is recorded, reflecting the general trend towards language democratization and the enrichment of film speech with elements natural to real communication. Phatic means become an important tool for creating a credible speech portrait.
- Research Article
- 10.1080/14781700.2026.2628092
- Apr 22, 2026
- Translation Studies
- Paolo Canavese
ABSTRACT This article aims to show that corpus linguistics and corpus-based translation studies can be useful in informing historical research on institutional translation. To this end, it draws on a synopsis of a corpus study of Swiss federal legislation in Italian. By combining quantitative and qualitative corpus analysis, it is possible to systematically identify linguistic change at the lexical, syntactic, and textual levels. Based on the assumption that evolving translation policies impact the linguistic formulation of institutional texts, these diachronic linguistic tendencies can be correlated with contextual developments, such as organizational shifts, measures to improve text quality, growing professionalization within language services, and technological change. At the same time, the article reflects critically on the limitations of corpus data and emphasizes the need to triangulate corpus findings with other methods, including sociological and historiographical research, in order to move from correlation to hypotheses of causal relationships.
- Research Article
- 10.54254/2753-7064/2026.32819
- Apr 20, 2026
- Communications in Humanities Research
- Congling Liu + 1 more
This critical analysis of discourse based on a corpus carefully studies language forms and discourse uses of hidden doers in news reports about a diplomatic event between China and the US. This study focuses on two main grammar structures: passive voice and using nouns to express actions. It analyzes how news discourses purposely hide or play down people who should take responsibility in describing diplomatic events. The research uses a self-made similar parallel corpus, which includes 18 typical news reports from China Daily and 18 from The New York Times respectively. Real research results show that structures that remove doers are widely used in news texts from both news agencies. However, The New York Times uses these language ways to hide doers much more frequently. This study not only proves the practical uses and analysis value of corpus linguistics methods in finding hidden rules of discourse in news language. It also shows the importance of checking small grammar choices when we learn how international news media build stories about sensitive diplomatic arguments and international conflicts. So it provides a view from language to help people understand media prejudice and hidden ideas in international news reporting.
- Research Article
- 10.33806/ijaes1076
- Apr 16, 2026
- International Journal of Arabic-English Studies
- Victor Khachan + 1 more
The teaching/learning of English literature to university EFL/ESL students has been essential to the development of their critical thinking skills. Using a corpus linguistics approach, the present study investigates the extent to which learners’ critical awareness is apparent through their analysis and interpretations of the characters and themes in the novel, A Farewell to Arms (1929), by Ernest Hemingway. A literary learner corpus of 206 literary argumentative essays was compiled from a '20th Century American Novel' course at an English-medium university in Lebanon. The essays were quantitatively analyzed to investigate prevalent thematic trends. This study was able to draw thematic differences between Hemingway’s novel and the learners’ literary analytic interpretations. The main findings indicated that the learners’ collocational networks strongly supported Hemingway’s stylistic ‘iceberg principle’- a critical thinking-based expectation Hemingway had for his readers. Most importantly, this lexical analysis revealed a clear ‘genre’ shift from narration in Hemingway’s A Farewell to Arms Corpus to argumentation in the Learner Literary Corpus. Overall, this study underscores the importance of literature courses in developing students’ critical thinking skills and the role corpus linguistics can play in quantifying ‘critical thinking’ measures in literature.
- Research Article
- 10.54254/2753-7064/2026.ht32772
- Apr 13, 2026
- Communications in Humanities Research
- Linu Pei
This study uses corpus linguistics to study the appropriateness of communicative culture in the American Chinese textbook Chinese: Listening, Speaking, Reading & Writing 3. We find three communicative culture items: "yao fan", "pang", and "xiao jie". These items may cause cross-cultural conflicts. We use the Modern Chinese Corpus of the Ministry of Education to check if the textbook uses these words properly.We find that correct translation of word meanings and clear cultural notes are both important for proper language use. This study can help with textbook writing and Chinese teaching.
- Research Article
- 10.17977/um011v11i22023p95-103
- Apr 9, 2026
- Jurnal Pendidikan Humaniora
- Mochamad Nuruz Zaman + 3 more
This study investigates the performance of the translation correction tool, Grammarly, in assessing the accuracy of translations; with a descriptive-based qualitative method. The object of this research is the bilingual annual report of PT. Gudang Garam Tbk, which was officially released from the website. Data is collected by analyzing documents and generating data findings. The data is further analyzed by ethnographic methods using domain, taxonomy, componential, and cultural themes. The findings are patterned manually with the classification of translation techniques, which resulted in accurate, inaccurate, and inaccurate translations. This finding reviews the Grammarly evaluating system to assess accuracy so that there are differences between manual analysis and system assessment (translation correction tool). This research contributes to the dynamics of translation learning technology to prioritize manual editing because the system does not have the accuracy to meet industry needs.