Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A systematic review of human-LLM interactions in computational thinking empirical studies

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

ABSTRACT Background and Context Artificial intelligence-based tools have been explored in developing computational thinking (CT) competencies. More recently, the role of large language models (LLMs) in teaching and learning CT has been investigated. However, limited research has mapped the role of LLMs within the CT landscape, particularly regarding how humans interact with LLMs to develop or evaluate CT competencies. Objectives The present study addresses this gap in the literature by exploring how different human-LLM interaction modes (standard-prompting mode, user-interface mode, context-based mode, and agent-facilitator mode) can be applied in CT studies and identifying the challenges that may arise in such applications. Method A systematic review was conducted following the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) guidelines to analyze 18 empirical, peer-reviewed CT studies that involved human-LLM interactions. Findings The findings indicate that most studies: (1) sampled college students; (2) focused on computer science; (3) were situated in formal learning environments; (4) applied a context-based human-LLM interaction mode; (5) targeted CT practices; (6) used ChatGPT as the LLM-based tools; (7) employed human-LLM interaction for content generation; and (8) identified the variability and instability of LLM outputs as the biggest challenge. Contributions Theoretically, this review identifies current trends and gaps in CT research involving LLMs. Practically, it provides researchers and practitioners with insights into how to use LLMs in CT studies.

Similar Papers
  • Research Article
  • Cite Count Icon 13
  • 10.1177/07356331241312365
Can Students Make STEM Progress With the Large Language Models (LLMs)? An Empirical Study of LLMs Integration Within Middle School Science and Engineering Practice
  • Jan 6, 2025
  • Journal of Educational Computing Research
  • Qing Guo + 4 more

The rapid development of large language models (LLMs) presented opportunities for the transformation of science and STEM education. Research on LLMs was in the exploratory phase, characterized by discussions and observations rather than empirical investigations. This study presented a framework for incorporating LLMs into Science and Engineering Practice (SEP), utilizing a case study on submarine construction, followed by a four-week quasi-experimental validation. The research employed conditional cluster sampling, selecting two homogeneous natural classes from a middle school in China to serve as the experimental and control groups. The key experimental variable was the inclusion of LLMs in the SEP project. Various validated and self-developed assessment tools were used to measure students’ STEM learning outcomes. Statistical analyses, including pre- and post-test paired comparisons within classes and ANCOVA for between-class differences, were performed to evaluate the effects of LLM integration. The results showed that students participating in SEP integrated with LLMs significantly improved their mastery of scientific knowledge, attitudes towards science, perceived usefulness of technology, understanding of engineering, computational thinking skills, and problem-solving abilities. In contrast, students participating in traditional SEP exhibited weaker knowledge acquisition, differences in understanding engineering concepts, and lack of development in computational thinking and problem-solving skills. This study was a pioneering effort in integrating LLMs into science education and provided a framework and case reference for the deeper application of LLMs in the future.

  • Research Article
  • Cite Count Icon 39
  • 10.1101/2024.04.26.24306390
A Systematic Review of ChatGPT and Other Conversational Large Language Models in Healthcare
  • Apr 27, 2024
  • medRxiv
  • Leyao Wang + 7 more

Background:The launch of the Chat Generative Pre-trained Transformer (ChatGPT) in November 2022 has attracted public attention and academic interest to large language models (LLMs), facilitating the emergence of many other innovative LLMs. These LLMs have been applied in various fields, including healthcare. Numerous studies have since been conducted regarding how to employ state-of-the-art LLMs in health-related scenarios to assist patients, doctors, and public health administrators.Objective:This review aims to summarize the applications and concerns of applying conversational LLMs in healthcare and provide an agenda for future research on LLMs in healthcare.Methods:We utilized PubMed, ACM, and IEEE digital libraries as primary sources for this review. We followed the guidance of Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRIMSA) to screen and select peer-reviewed research articles that (1) were related to both healthcare applications and conversational LLMs and (2) were published before September 1st, 2023, the date when we started paper collection and screening. We investigated these papers and classified them according to their applications and concerns.Results:Our search initially identified 820 papers according to targeted keywords, out of which 65 papers met our criteria and were included in the review. The most popular conversational LLM was ChatGPT from OpenAI (60), followed by Bard from Google (1), Large Language Model Meta AI (LLaMA) from Meta (1), and other LLMs (5). These papers were classified into four categories in terms of their applications: 1) summarization, 2) medical knowledge inquiry, 3) prediction, and 4) administration, and four categories of concerns: 1) reliability, 2) bias, 3) privacy, and 4) public acceptability. There are 49 (75%) research papers using LLMs for summarization and/or medical knowledge inquiry, and 58 (89%) research papers expressing concerns about reliability and/or bias. We found that conversational LLMs exhibit promising results in summarization and providing medical knowledge to patients with a relatively high accuracy. However, conversational LLMs like ChatGPT are not able to provide reliable answers to complex health-related tasks that require specialized domain expertise. Additionally, no experiments in our reviewed papers have been conducted to thoughtfully examine how conversational LLMs lead to bias or privacy issues in healthcare research.Conclusions:Future studies should focus on improving the reliability of LLM applications in complex health-related tasks, as well as investigating the mechanisms of how LLM applications brought bias and privacy issues. Considering the vast accessibility of LLMs, legal, social, and technical efforts are all needed to address concerns about LLMs to promote, improve, and regularize the application of LLMs in healthcare.

  • Research Article
  • 10.3390/ai7010025
Pedagogical Transformation Using Large Language Models in a Cybersecurity Course
  • Jan 13, 2026
  • AI
  • Rodolfo Ostos + 9 more

Large Language Models (LLMs) are increasingly used in higher education, but their pedagogical role in fields like cybersecurity remains under-investigated. This research explores integrating LLMs into a university cybersecurity course using a designed pedagogical approach based on active learning, problem-based learning (PBL), and computational thinking (CT). Instead of viewing LLMs as definitive sources of knowledge, the framework sees them as cognitive tools that support reasoning, clarify ideas, and assist technical problem-solving while maintaining human judgment and verification. The study uses a qualitative, practice-based case study over three semesters. It features four activities focusing on understanding concepts, installing and configuring tools, automating procedures, and clarifying terminology, all incorporating LLM use in individual and group work. Data collection involved classroom observations, team reflections, and iterative improvements guided by action research. Results show that LLMs can provide valuable, customized support when students actively engage in refining, validating, and solving problems through iteration. LLMs are especially helpful for clarifying concepts and explaining procedures during moments of doubt or failure. Still, common issues like incomplete instructions, mismatched context, and occasional errors highlight the importance of verifying LLM outputs with trusted sources. Interestingly, these limitations often act as teaching opportunities, encouraging critical thinking crucial in cybersecurity. Ultimately, this study offers empirical evidence of human–AI collaboration in education, demonstrating how LLMs can enrich active learning.

  • Supplementary Content
  • 10.2196/89862
Large Language Models in Colorectal Cancer Care and Clinical Decision Support: Systematic Review
  • May 21, 2026
  • Journal of Medical Internet Research
  • Jinglei Tian + 5 more

BackgroundColorectal cancer (CRC) is a leading cause of cancer morbidity and mortality worldwide. The complexity of guideline-concordant care and unstructured clinical data has driven demand for decision-support tools. Large language models (LLMs) show promise for processing clinical data and patient–provider communication, yet evidence is fragmented, and a CRC-specific synthesis across the full care continuum is lacking.ObjectiveThis systematic review evaluates the current applications, performance determinants, and clinical implications of LLMs across the continuum of CRC care.MethodsFollowing PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses), we searched 6 databases (PubMed, Embase, Web of Science, Scopus, CINAHL, Cochrane) through April 1, 2026. Eligible studies were peer-reviewed original investigations of LLMs on CRC tasks with extractable outcomes; reviews, editorials, and abstracts were excluded. Two reviewers assessed quality with QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies-2), PROBAST (prediction model risk of bias assessment tool), and ROBINS-I (Risk of Bias in Nonrandomized Studies - of Interventions). Data on model types, applications, prompts, input/output formats, and outcomes were analyzed descriptively, with narrative synthesis per synthesis without meta-analysis (SWiM) guidelines.ResultsOf 8880 records, 37 studies met inclusion criteria (2023‐2026), mostly from China and the United States, with GPT series most frequently evaluated. Overall risk of bias was low in 10/37 studies (27.0%), moderate in 14/37 (37.8%), unclear in 7/37 (18.9%), and high or serious in 6/37 (16.2%). Problematic domains included outcome measurement, intervention classification, patient selection, and lack of blinded assessment. LLMs showed utility in automating data extraction from clinical texts, supporting patient education, aiding diagnosis, and assisting clinical decision-making, with emerging visual interpretation and multimodal capacities. Domain-specific and multimodal models showed advantages over general-purpose models in certain tasks. Performance was significantly influenced by prompt design, from zero-shot queries to fine-tuning. Despite efficiency and outcome benefits, challenges persist regarding methodological quality, data privacy, and generalizability.ConclusionsThis review provides an integrative framework synthesizing evidence across study designs and LLM categories in CRC care. Unlike prior reviews addressing gastroenterology broadly or limited to one design, it covers the full CRC continuum and, for the first time, comparatively evaluates general-purpose, domain-specific, and multimodal LLMs, clarifying how prompt engineering and heterogeneous metrics shape outcomes. Although findings support LLMs’ clinical potential, results must be interpreted cautiously, given low overall evidence quality. Most studies lacked safeguards against bias—blinded assessment, confounder adjustment, or prospective multicenter validation. Substantial heterogeneity across tasks, LLM types, prompts, reference standards, and outcomes means reported advantages cannot be generalized. Future work should prioritize real-world integration via prospective multicenter validation, robust privacy frameworks, and rigorous human oversight. Amid rising global CRC burden and health care disparities, this review informs clinical translation, equitable scaling, and policy on LLM deployment.

  • Research Article
  • Cite Count Icon 14
  • 10.2196/70535
Unveiling the Potential of Large Language Models in Transforming Chronic Disease Management: Mixed Methods Systematic Review.
  • Apr 16, 2025
  • Journal of medical Internet research
  • Caixia Li + 7 more

Chronic diseases are a major global health burden, accounting for nearly three-quarters of the deaths worldwide. Large language models (LLMs) are advanced artificial intelligence systems with transformative potential to optimize chronic disease management; however, robust evidence is lacking. This review aims to synthesize evidence on the feasibility, opportunities, and challenges of LLMs across the disease management spectrum, from prevention to screening, diagnosis, treatment, and long-term care. Following the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analysis) guidelines, 11 databases (Cochrane Central Register of Controlled Trials, CINAHL, Embase, IEEE Xplore, MEDLINE via Ovid, ProQuest Health & Medicine Collection, ScienceDirect, Scopus, Web of Science Core Collection, China National Knowledge Internet, and SinoMed) were searched on April 17, 2024. Intervention and simulation studies that examined LLMs in the management of chronic diseases were included. The methodological quality of the included studies was evaluated using a rating rubric designed for simulation-based research and the risk of bias in nonrandomized studies of interventions tool for quasi-experimental studies. Narrative analysis with descriptive figures was used to synthesize the study findings. Random-effects meta-analyses were conducted to assess the pooled effect estimates of the feasibility of LLMs in chronic disease management. A total of 20 studies examined general-purpose (n=17) and retrieval-augmented generation-enhanced LLMs (n=3) for the management of chronic diseases, including cancer, cardiovascular diseases, and metabolic disorders. LLMs demonstrated feasibility across the chronic disease management spectrum by generating relevant, comprehensible, and accurate health recommendations (pooled accurate rate 71%, 95% CI 0.59-0.83; I2=88.32%) with retrieval-augmented generation-enhanced LLMs having higher accuracy rates compared to general-purpose LLMs (odds ratio 2.89, 95% CI 1.83-4.58; I2=54.45%). LLMs facilitated equitable information access; increased patient awareness regarding ailments, preventive measures, and treatment options; and promoted self-management behaviors in lifestyle modification and symptom coping. Additionally, LLMs facilitate compassionate emotional support, social connections, and health care resources to improve the health outcomes of chronic diseases. However, LLMs face challenges in addressing privacy, language, and cultural issues; undertaking advanced tasks, including diagnosis, medication, and comorbidity management; and generating personalized regimens with real-time adjustments and multiple modalities. LLMs have demonstrated the potential to transform chronic disease management at the individual, social, and health care levels; however, their direct application in clinical settings is still in its infancy. A multifaceted approach that incorporates robust data security, domain-specific model fine-tuning, multimodal data integration, and wearables is crucial for the evolution of LLMs into invaluable adjuncts for health care professionals to transform chronic disease management. PROSPERO CRD42024545412; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024545412.

  • Research Article
  • Cite Count Icon 11
  • 10.1186/s40561-025-00406-0
How do generative artificial intelligence (AI) tools and large language models (LLMs) influence language learners’ critical thinking in EFL education? A systematic review
  • Aug 4, 2025
  • Smart Learning Environments
  • Jing Liu + 2 more

As generative artificial intelligence (AI) tools and large language models (LLMs)-powered applications develop rapidly in the era of algorithms, it should be integrated thoughtfully to enhance English as a Foreign Language (EFL) teaching and learning without replacing learners’ critical thinking (CT). This study systematically analyzes the impact of generative AI tools and LLMs on language learners’ CT in EFL education using the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) framework to identify, evaluate, and synthesize relevant studies from 2022 to 2025. A thorough review of 15 selected studies focuses on generative AI tools and LLMs’ dual nature, research methods, main focuses, theory and models, limitations and challenges, and future directions in the field based on Web of Science (WoS), SCOPUS, ERIC, ProQuest, and Google Scholar. The findings identified generative AI tools and LLMs possessed both the potential to nurture and the risk of hindering CT in EFL education. 66.67% of studies reported generative AI tools and LLMs’ positive role in CT, while 33.33% of studies reported its negative role in CT. Furthermore, 3 types of research methods, 3 key themes of research focus, and 4 groups of theoretical perspectives were examined. However, 4 kinds of limitations in this field remain, including research scope, user dependency, generative AI reliability, and pedagogical integration. Future research can focus on assessing long-term effects, broadening research scope, promoting responsible AI use, and refining pedagogical strategies. Finally, Limitations, implications and future direction of this study were discussed.

  • Research Article
  • Cite Count Icon 8
  • 10.1007/s00405-025-09504-8
Clinical decision support using large language models in otolaryngology: a systematic review.
  • Jun 6, 2025
  • European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery
  • Rania Filali Ansary + 1 more

This systematic review evaluated the diagnostic accuracy of large language models (LLMs) in otolaryngology-head and neck surgery clinical decision-making. PubMed/MEDLINE, Cochrane Library, and Embase databases were searched for studies investigating clinical decision support accuracy of LLMs in otolaryngology. Three investigators searched the literature for peer-reviewed studies investigating the application of LLMs as clinical decision support for real clinical cases according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. The following outcomes were considered: diagnostic accuracy, additional examination and treatment recommendations. Study quality was assessed using the modified Methodological Index for Non-Randomized Studies (MINORS). Of the 285 eligible publications, 17 met the inclusion criteria, accounting for 734 patients across various otolaryngology subspecialties. ChatGPT-4 was the most evaluated LLM (n = 14/17), followed by Claude-3/3.5 (n = 2/17), and Gemini (n = 2/17). Primary diagnostic accuracy ranged from 45.7 to 80.2% across different LLMs, with Claude often outperforming ChatGPT. LLMs demonstrated lower accuracy in recommending appropriate additional examinations (10-29%) and treatments (16.7-60%), with substantial subspecialty variability. Treatment recommendation accuracy was highest in head and neck oncology (55-60%) and lowest in rhinology (16.7%). There was substantial heterogeneity across studies for the inclusion criteria, information entered in the application programming interface, and the methods of accuracy assessment. LLMs demonstrate promising moderate diagnostic accuracy in otolaryngology clinical decision support, with higher performance in providing diagnoses than in suggesting appropriate additional examinations and treatments. Emerging findings support that Claude often outperforms ChatGPT. Methodological standardization is needed for future research. NA.

  • Research Article
  • Cite Count Icon 19
  • 10.1057/s41599-025-04471-1
LLM-based collaborative programming: impact on students’ computational thinking and self-efficacy
  • Feb 7, 2025
  • Humanities and Social Sciences Communications
  • Yi-Miao Yan + 3 more

At present, collaborative programming is a prevalent approach in programming education, yet its effectiveness often falls short due to the varying levels of coding skills among team members. To address these challenges, Large Language Models (LLMs) can be introduced as a supportive tool to enhance both the efficiency and outcomes of collaborative programming. In this shift, the structure of collaborative teams evolves from human-to-human to a new paradigm consisting of human, human, and AI. To investigate the effectiveness of integrating LLMs into collaborative programming, this study designed a quasi-experiment. To explore the effectiveness of integrating LLMs into collaborative programming, we conducted a quasi-experiment involving 82 sixth- and seventh-grade students, who were randomly assigned to either an experimental group or a control group. The results showed that incorporating LLMs into collaborative programming significantly reduced students’ cognitive load and improved their computational thinking skills. However, no significant difference in self-efficacy was observed between the two groups, likely due to the cognitive demand students faced when transitioning from graphical programming to text-based coding. Despite this, the study remains optimistic about the potential of LLM-enhanced collaborative programming, as students learning in this way exhibit lower cognitive load than those in conventional environments.

  • Research Article
  • 10.1093/sleep/zsaf090.1401
1401 Bringing Medicine Expertise to Your Screen: A New Frontier in Curbside Sleep Consultation Leveraging Large Language Models?
  • May 19, 2025
  • SLEEP
  • Nina Kuei + 4 more

Introduction Advances in large language models (LLMs) have opened new avenues for healthcare applications. Recently, ChatGPT-4 successfully achieved the pass mark >80% in 5 of 10 sleep medicine examination domains, indicating a strong foundational knowledge of sleep medicine. However, the ability to answer USMLE-type multiple-choice questions may not equate to the capacity to offer accurate and comprehensive answers to clinical queries or real-world case scenarios. Current literature exhibits a significant gap in validation research examining the clinical utility of LLMs in evidence-based medicine practice. This investigation aims to evaluate the potential usefulness and reliability of LLMs as adjunctive clinical decision-support tools (namely, “curbside consultants”) in sleep medicine. Methods Six clinical sleep queries and six case scenarios were presented to 6 LLMs including 3 general-purpose LLMs (ChatGPT-4o, Gemini-1.5-Pro, and Llama-3.1-405B) and 3 medical-specialized LLMs (OpenEvidence, Clara AI, and MediGPT). Performance assessment was conducted independently utilizing two independently developed 5-point Likert scales evaluating two primary domains: answer content (accuracy, relevance, comprehensiveness/depth, clarity/coherence, and unique insightfulness) and reference quality (accuracy, relevance, currentness, comprehensiveness/depth, and searchability). Benchmark answers were established through consensus among four sleep medicine specialists. Analysis of variance was used to compare the performance of the LLMs. Results The medical LLMs demonstrated superior overall performance compared to the general LLMs (P< 0.001). The primary distinction was observed in reference quality metrics, where medical LLMs significantly outperformed general LLMs across all parameters: accuracy, relevance, and searchability (p< 0.001), currentness (p=0.001), and comprehensiveness/depth (p=0.010). Notably, Open Evidence achieved the highest reference quality (p< 0.001). In contrast, the answer content analysis revealed no significant overall differences between medical and general LLMs (p=0.659). Most answer content-related metrics, including accuracy, relevance, or unique insightfulness, did not differ significantly (p>0.050). An exception was clarify/coherence, where medical LLMs were superior to general LLMs (p=0.030). Furthermore, MediGPT and ChatGPT-4o displayed better content comprehensiveness/depth relative to other LLMs (p< 0.001). These two LLMs exhibited comparable performance in overall metrics (p=0.615), answer contents (p=0.922), and reference quality (p=0.621). Conclusion Both medical-specialized and general-purpose LLMs show promise as adjunctive decision-support tools in clinical practice. However, substantial improvements in reference quality are critically needed across most LLM platforms. Support (if any)

  • Supplementary Content
  • Cite Count Icon 120
  • 10.2196/22769
Applications and Concerns of ChatGPT and Other Conversational Large Language Models in Health Care: Systematic Review
  • Nov 7, 2024
  • Journal of Medical Internet Research
  • Leyao Wang + 7 more

BackgroundThe launch of ChatGPT (OpenAI) in November 2022 attracted public attention and academic interest to large language models (LLMs), facilitating the emergence of many other innovative LLMs. These LLMs have been applied in various fields, including health care. Numerous studies have since been conducted regarding how to use state-of-the-art LLMs in health-related scenarios.ObjectiveThis review aims to summarize applications of and concerns regarding conversational LLMs in health care and provide an agenda for future research in this field.MethodsWe used PubMed, ACM, and the IEEE digital libraries as primary sources for this review. We followed the guidance of PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) to screen and select peer-reviewed research articles that (1) were related to health care applications and conversational LLMs and (2) were published before September 1, 2023, the date when we started paper collection. We investigated these papers and classified them according to their applications and concerns.ResultsOur search initially identified 820 papers according to targeted keywords, out of which 65 (7.9%) papers met our criteria and were included in the review. The most popular conversational LLM was ChatGPT (60/65, 92% of papers), followed by Bard (Google LLC; 1/65, 2% of papers), LLaMA (Meta; 1/65, 2% of papers), and other LLMs (6/65, 9% papers). These papers were classified into four categories of applications: (1) summarization, (2) medical knowledge inquiry, (3) prediction (eg, diagnosis, treatment recommendation, and drug synergy), and (4) administration (eg, documentation and information collection), and four categories of concerns: (1) reliability (eg, training data quality, accuracy, interpretability, and consistency in responses), (2) bias, (3) privacy, and (4) public acceptability. There were 49 (75%) papers using LLMs for either summarization or medical knowledge inquiry, or both, and there are 58 (89%) papers expressing concerns about either reliability or bias, or both. We found that conversational LLMs exhibited promising results in summarization and providing general medical knowledge to patients with a relatively high accuracy. However, conversational LLMs such as ChatGPT are not always able to provide reliable answers to complex health-related tasks (eg, diagnosis) that require specialized domain expertise. While bias or privacy issues are often noted as concerns, no experiments in our reviewed papers thoughtfully examined how conversational LLMs lead to these issues in health care research.ConclusionsFuture studies should focus on improving the reliability of LLM applications in complex health-related tasks, as well as investigating the mechanisms of how LLM applications bring bias and privacy issues. Considering the vast accessibility of LLMs, legal, social, and technical efforts are all needed to address concerns about LLMs to promote, improve, and regularize the application of LLMs in health care.

  • Research Article
  • Cite Count Icon 4
  • 10.34190/ejel.22.3.3992
Augmented and Virtual Reality in Computational Thinking: A Systematic Review of Their Individual Impacts, Advantages, Challenges, and Future Directions
  • May 19, 2025
  • Electronic Journal of e-Learning
  • Muhammad Aizri Fadillah + 3 more

Computational thinking (CT) skills are increasingly important in education to prepare students for the challenges of the digital age. Augmented Reality (AR) and Virtual Reality (VR) have been introduced as immersive technologies that have the potential to enhance CT skills through more interactive learning experiences. However, there is still a gap in understanding the effectiveness of these technologies in supporting the development of CT, particularly in different levels of education and disciplines. Although several studies have highlighted the benefits of AR and VR in education, no systematic review integrates these findings to identify advantages, challenges, and opportunities for further implementation. Therefore, this study conducted a systematic review based on the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines by analyzing 25 empirical studies (AR=17, VR=8) obtained from the Scopus database (2008-2024). The analysis addresses four key research questions: (1) the current state of AR/VR in CT development, (2) their advantages, (3) implementation challenges, and (4) future research directions. The results show that AR is more widespread than VR at various levels of education, with dominance in higher education followed by secondary and primary schools. Computer science is the main field of application of AR and VR, while AR is also widely applied in mathematics to increase interest and problem-solving. A total of 11 studies reported significant impacts of these technologies on CT, with AR being superior in increasing student motivation and engagement, as well as aiding in problem-solving and debugging. In contrast, VR provides a more immersive learning experience by strengthening concept understanding, especially in programming and recursion. However, several obstacles in the application of AR and VR, such as hardware limitations, costs, and user skills, affect the effectiveness of these technologies in the learning environment. This study also identified potential future research, including the exploration of VR in primary and kindergarten education, the application of VR in non-computer science fields, and the efficient use of these technologies in supporting the CT process. This study provides more precise insights into the optimal ways of utilizing AR and VR in developing CT skills. It is a reference for educators, policymakers, and researchers in supporting CT learning.

  • Research Article
  • Cite Count Icon 8
  • 10.2196/79202
Large Language Models for Health Care Text Classification: Systematic Review.
  • Jun 16, 2025
  • JMIR AI
  • Hajar Sakai + 1 more

Large language models (LLMs) have fundamentally transformed approaches to natural language processing tasks across diverse domains. In health care, accurate and cost-efficient text classification is crucial-whether for clinical note analysis, diagnosis coding, or other related tasks-and LLMs present promising potential. Text classification has long faced multiple challenges, including the need for manual annotation during training, the handling of imbalanced data, and the development of scalable approaches. In health care, additional challenges arise, particularly the critical need to preserve patient data privacy and the complexity of medical terminology. Numerous studies have leveraged LLMs for automated health care text classification and compared their performance with traditional machine learning-based methods, which typically require embedding, annotation, and training. However, existing systematic reviews of LLMs either do not specialize in text classification or do not focus specifically on the health care domain. This research synthesizes and critically evaluates the current evidence in the literature on the use of LLMs for text classification in health care settings. Major databases (eg, Google Scholar, Scopus, PubMed, ScienceDirect) and other resources were queried for papers published between 2018 and 2024, following the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, resulting in 65 eligible research articles. These studies were categorized by text classification type (eg, binary classification, multilabel classification), application (eg, clinical decision support, public health and opinion analysis), methodology, type of health care text, and the metrics used for evaluation and validation. The systematic review includes 65 research articles published between 2020 and Q3 2024, showing a significant increase in publications over time, with 28 papers published in Q1-Q3 2024 alone. Fine-tuning was the most common LLM-based approach (35 papers), followed by prompt engineering (17 papers). BERT (Bidirectional Encoder Representations from Transformers) variants were predominantly used for multilabel classification (50%), whereas closed-source LLMs were most commonly applied to binary (44.0%) and multiclass (30.6%) classification tasks. Clinical decision support was the most frequent application (29 papers). Over 80% of studies used English-language datasets, with clinical notes being the most common text type. All studies employed accuracy-related metrics for evaluation, and the findings consistently showed that LLMs outperformed traditional machine learning approaches in health care text classification tasks. This review identifies existing gaps in the literature and highlights future research directions for further investigation.

  • Supplementary Content
  • 10.3390/healthcare14010045
Large Language Models for Cardiovascular Disease, Cancer, and Mental Disorders: A Review of Systematic Reviews
  • Dec 24, 2025
  • Healthcare
  • Andreas Triantafyllidis + 8 more

Background/Objective: The use of Large Language Models (LLMs) has recently gained significant interest from the research community toward the development and adoption of Generative Artificial Intelligence (GenAI) solutions for healthcare. The present work introduces the first meta-review (i.e., review of systematic reviews) in the field of LLMs for chronic diseases, focusing particularly on cardiovascular, cancer, and mental diseases, to identify their value in patient care, and challenges for their implementation and clinical application. Methods: A literature search in the bibliographic databases of PubMed and Scopus was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, to identify systematic reviews incorporating LLMs. The original studies included in the reviews were synthesized according to their target disease, specific application, LLMs used, data sources, accuracy, and key outcomes. Results: The literature search identified 5 systematic reviews respecting our inclusion and exclusion criteria, which examined 81 unique LLM-based solutions. The highest percentage of the solutions targeted mental disease (86%), followed by cancer (7%) and cardiovascular disease (6%), implying a large research focus in mental health. Generative Pre-trained Transformer (GPT)-family models were used most frequently (~55%), followed by Bidirectional Encoder Representations from Transformers (BERT) variants (~40%). Key application areas included depression detection and classification (38%), suicidal ideation detection (7%), question answering based on treatment guidelines and recommendations (7%), and emotion classification (5%). Study aims and designs were highly heterogeneous, and methodological quality was generally moderate with frequent risk-of-bias concerns. Reported performance varied widely across domains and datasets, and many evaluations relied on fictional vignettes or non-representative data, limiting generalisability. The most significant found challenges in the development and evaluation of LLMs include inconsistent accuracy, bias detection and mitigation, model transparency, data privacy, need for continual human oversight, ethical concerns and guidelines, as well as the design and conduction of high-quality studies. Conclusions: While LLMs show promise for screening, triage, decision support, and patient education—particularly in mental health—the current literature is descriptive and constrained by data, transparency, and safety gaps. We recommend prioritizing rigorous real-world evaluations, diverse benchmark datasets, bias-auditing, and governance frameworks before LLM clinical deployment and large adoption.

  • Research Article
  • Cite Count Icon 2
  • 10.3352/jeehp.2025.22.36
Performance of large language models in medical licensing examinations: a systematic review and meta-analysis.
  • Nov 18, 2025
  • Journal of educational evaluation for health professions
  • Haniyeh Nouri + 5 more

This study systematically evaluates and compares the performance of large language models (LLMs) in answering medical licensing examination questions. By conducting subgroup analyses based on language, question format, and model type, this meta-analysis aims to provide a comprehensive overview of LLM capabilities in medical education and clinical decision-making. This systematic review, registered in PROSPERO and following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, searched MEDLINE (PubMed), Scopus, and Web of Science for relevant articles published up to February 1, 2025. The search strategy included Medical Subject Headings (MeSH) terms and keywords related to ("ChatGPT" OR "GPT" OR "LLM variants") AND ("medical licensing exam*" OR "medical exam*" OR "medical education" OR "radiology exam*"). Eligible studies evaluated LLM accuracy on medical licensing examination questions. Pooled accuracy was estimated using a random-effects model, with subgroup analyses by LLM type, language, and question format. Publication bias was assessed using Egger's regression test. This systematic review identified 2,404 studies. After removing duplicates and excluding irrelevant articles through title and abstract screening, 36 studies were included after full-text review. The pooled accuracy was 72% (95% confidence interval, 70.0% to 75.0%) with high heterogeneity (I2=99%, P<0.001). Among LLMs, GPT-4 achieved the highest accuracy (81%), followed by Bing (79%), Claude (74%), Gemini/Bard (70%), and GPT-3.5 (60%) (P=0.001). Performance differences across languages (range, 62% in Polish to 77% in German) were not statistically significant (P=0.170). LLMs, particularly GPT-4, can match or exceed medical students' examination performance and may serve as supportive educational tools. However, due to variability and the risk of errors, they should be used cautiously as complements rather than replacements for traditional learning methods.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 103
  • 10.1186/s12909-024-05239-y
Large language models for generating medical examinations: systematic review.
  • Mar 29, 2024
  • BMC medical education
  • Yaara Artsi + 5 more

Writing multiple choice questions (MCQs) for the purpose of medical exams is challenging. It requires extensive medical knowledge, time and effort from medical educators. This systematic review focuses on the application of large language models (LLMs) in generating medical MCQs. The authors searched for studies published up to November 2023. Search terms focused on LLMs generated MCQs for medical examinations. Non-English, out of year range and studies not focusing on AI generated multiple-choice questions were excluded. MEDLINE was used as a search database. Risk of bias was evaluated using a tailored QUADAS-2 tool. Overall, eight studies published between April 2023 and October 2023 were included. Six studies used Chat-GPT 3.5, while two employed GPT 4. Five studies showed that LLMs can produce competent questions valid for medical exams. Three studies used LLMs to write medical questions but did not evaluate the validity of the questions. One study conducted a comparative analysis of different models. One other study compared LLM-generated questions with those written by humans. All studies presented faulty questions that were deemed inappropriate for medical exams. Some questions required additional modifications in order to qualify. LLMs can be used to write MCQs for medical examinations. However, their limitations cannot be ignored. Further study in this field is essential and more conclusive evidence is needed. Until then, LLMs may serve as a supplementary tool for writing medical examinations. 2 studies were at high risk of bias. The study followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant