Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Reliable decision support with LLMs: a framework for evaluating consistency in binary text classification applications

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

ABSTRACT Introduction This study introduces a framework for evaluating consistency in large language model (LLM) binary text classification, addressing the lack of established reliability assessment methods. Methodology Adapting psychometric principles, we determine sample size requirements, develop metrics for invalid responses, and evaluate intra- and inter-rater reliability. Our case study examines financial news sentiment classification across 14 LLMs (including claude-3–7-sonnet, gpt-4o, deepseek-r1, gemma3, llama3.2, phi4, and command-r-plus), with five replicates per model on 1,350 articles. Results Models demonstrated high intra-rater consistency, achieving perfect agreement on 90–98% of examples, with minimal differences between expensive and economical models from the same families. When validated against StockNewsAPI labels, models achieved strong performance (accuracy 0.76–0.88), with smaller models like gemma3:1B, llama3.2:3B, and claude-3–5-haiku outperforming larger counterparts. All models performed at chance when predicting actual market movements, indicating task constraints rather than model limitations. Practical implications Our framework provides systematic guidance for LLM selection, sample size planning, and reliability assessment, enabling organisations to optimise resources for classification tasks.

Similar Papers
  • Research Article
  • Cite Count Icon 11
  • 10.1287/ijds.2023.0007
How Can IJDS Authors, Reviewers, and Editors Use (and Misuse) Generative AI?
  • Apr 1, 2023
  • INFORMS Journal on Data Science
  • Galit Shmueli + 7 more

How Can <i>IJDS</i> Authors, Reviewers, and Editors Use (and Misuse) Generative AI?

  • Research Article
  • 10.1200/jco.2025.43.16_suppl.e13705
Utilization of large language models to facilitate guideline-directed therapy of genitourinary cancer screening and primary workup.
  • Jun 1, 2025
  • Journal of Clinical Oncology
  • Matthew Driscoll + 5 more

e13705 Background: The rise of large language models (LLMs) offers a unique opportunity to streamline healthcare and simplify medical decision-making. These AI tools are increasingly used in clinical medicine, supporting physicians in tasks ranging from diagnosis to treatment planning. Here, the authors investigate the ability of four common LLMs to direct genitourinary cancer screening and primary workup in alignment with the American Urological Association’s (AUA) Clinical Practice Guidelines (CPGs). Methods: Four LLMs - ChatGPT 4o, Claude 3.5 Sonnet, Google Gemini 1.5 Flash, and OpenEvidence - were selected for evaluation on the basis of popularity, accessibility, and target audience. AUA guidelines for Microhematuria and Early Detection of Prostate Cancer were reworded into question prompts and submitted to each LLM. All queries were conducted within a 72-hour time frame to minimize variability due to LLM adaptation. Two board certified urologists blindly reviewed and assigned responses a rating of concordant, discordant, or indeterminate with AUA CPGs. The responses were evaluated using chi-square tests with the assumption that the LLMs were equally as likely to generate concordant, discordant, or indeterminate responses. Interrater agreement was assessed using a Cohen’s kappa. Results: Of the total 456 unique assessments of LLM performance made, 95% of the assessments were concordant with the AUA CPGs. Each of the four LLMs performed significantly better than chance (p &lt; 0.00001, α &lt; 0.05). Of the four LLMs, ChatGPT produced the highest rate of responses noted as concordant by both reviewers at 98.6%, compared to Gemini which produced the lowest at 84.2%. ChatGPT was the only of the four LLMs to not produce a response that was rated as discordant by both reviewers whereas the other three LLMs each produced just one response that was unanimously discordant. Interrater agreement between reviewers was strong, in line with the moderate range (Cohen’s kappa = 0.506). Lastly, there was a statistically significant difference in the concordance with AUA CPGs between the four LLMs (p=0.0001). However, there was no statistically significant difference with concordance between ChatGPT, OpenEvidence, and Claude responses themselves (p = 0.06). Conclusions: These results suggest that LLMs are significantly making recommendations in line with AUA CPGs. These findings indicate that LLMs could play a valuable role in assisting physicians with the screening and initial evaluation of common genitourinary cancers, particularly those that present in outpatient or primary care settings. It is important to recognize that LLMs are continuously evolving and improving with each query. While this ongoing learning enhances their potential to become more accurate and reliable over time, it also underscores the need for continuous oversight to ensure their responses remain aligned with CPGs.

  • Research Article
  • 10.1136/bmjhci-2025-101956
Open-source large language model-based on-premises pipeline for automated data extraction from unstructured electronic health records: a pilot study
  • Jun 1, 2026
  • BMJ Health & Care Informatics
  • Vasileios Ntinopoulos + 4 more

ObjectivesWe evaluated an on-premises, open-source large language model (LLM)-based data extraction pipeline for automated data extraction from unstructured electronic health records (EHRs).MethodsAutomated script-based EHR data preprocessing extracted 50 medical texts in German, which were entered into the LLM pipeline. 4-bit and 8-bit quantizations of 14 mid-sized LLMs (30B–90B parameters) were evaluated in 6 information extraction, 11 binary classification and 5 multilevel classification tasks comprising all variables of the European System for Cardiac Operative Risk Evaluation 2 (EuroSCORE 2) model for 1100 predictions each. LLM response consistency was assessed over three same-prompt iterations.ResultsIn overall accuracy, Qwen3-30b-a3b-q8 presented the highest value (0.954) and 13 LLMs had values over 0.90. In information extraction accuracy, 12 LLMs exhibited a value of 1.0 and all 14 LLMs had values over 0.96. In binary classification accuracy, Llama3.2-vision-90b-q4 exhibited the highest value (0.972), 5 LLMs had values of at least 0.95 and all 14 LLMs showed values over 0.93. In multilevel classification accuracy, Qwen3-30b-a3b-q8 exhibited the highest value (0.940), four LLMs had values over 0.90 and all LLMs presented values over 0.80. Nine LLMs exhibited perfect response consistency and the remaining five LLMs had a Krippendorff’s alpha value of 0.999.DiscussionMultiple LLMs exhibited high accuracy in information extraction, binary classification, multilevel classification and response consistency and seem able to reliably automate data extraction from EHRs.ConclusionThis pilot study demonstrates the feasibility of on-premises, privacy-preserving, LLM-based automated EHR data extraction pipelines. Larger-scope studies are warranted to validate their potential in healthcare.

  • Conference Article
  • Cite Count Icon 1
  • 10.54941/ahfe1006669
Enhancing Thematic Analysis with Local LLMs: A Scientific Evaluation of Prompt Engineering Techniques
  • Jan 1, 2025
  • AHFE international
  • Timothy Meyer + 2 more

Thematic Analysis (TA) is a powerful tool for human factors, HCI, and UX researchers to gather system usability insights from qualitative data like open-ended survey questions. However, TA is both time consuming and difficult, requiring researchers to review and compare hundreds, thousands, or even millions of pieces of text. Recently, this has driven many to explore using Large Language Models (LLMs) to support such an analysis. However, LLMs have their own processing limitations and usability challenges when implementing them reliably as part of a research process – especially when working with a large corpus of data that exceeds LLM context windows. These challenges are compounded when using locally hosted LLMs, which may be necessary to analyze sensitive and/or proprietary data. However, little human factors research has rigorously examined how various prompt engineering techniques can augment an LLM to overcome these limitations and improve usability. Accordingly, in the present paper, we investigate the impact of several prompt engineering techniques on the quality of LLM-mediated TA. Using a local LLM (Llama 3.1 8b) to ensure data privacy, we developed four LLM variants with progressively complex prompt engineering techniques and used them to extract themes from user feedback regarding the usability of a novel knowledge management system prototype. The LLM variants were as follows:1.A “baseline” variant with no prompt engineering or scalability2.A “naïve batch processing” variant that sequentially analyzed small batches of the user feedback to generate a single list of themes3.An “advanced batch processing” variant that built upon the naïve variant by adding prompt engineering techniques (e.g., chain-of-thought prompting)4.A “cognition-inspired” variant that incorporated advanced prompt engineering techniques and kept a working memory-like log of themes and their frequencyContrary to conventional approaches to studying LLMs, which largely rely upon descriptive statistics (e.g., % improvement), we systematically applied a set of evaluation methods from behavioral science and human factors. We performed three stages of evaluation of the outputs of each LLM variant: we compared the LLM outputs to our team’s original TA, we had human factors professionals (N = 4) rate the quality and usefulness of the outputs, and we compared the Inter-Rater Reliability (IRR) of other human factors professionals (N = 2) attempting to code the original data with the outputs generated by each variant. Results demonstrate that even small, locally deployed LLMs can produce high-quality TA when guided by appropriate prompts. While the “baseline” variant performed surprisingly well for small datasets, we found that the other, scalable methods were dependent upon advanced prompt engineering techniques to be successful. Only our novel "cognition-inspired" approach performed as well as the “baseline” variant in qualitative and quantitative comparisons of ratings and coding IRR. This research provides practical guidance for human factors researchers looking to integrate LLMs into their qualitative analysis workflows, disentangling and uncovering the importance of context window limitations, batch processing strategies, and advanced prompt engineering techniques. The findings suggest that local LLMs can serve as valuable and scalable tools in thematic analysis.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 108
  • 10.1038/s41746-024-01024-9
CancerGPT for few shot drug pair synergy prediction using large pretrained language models
  • Feb 19, 2024
  • NPJ Digital Medicine
  • Tianhao Li + 6 more

Large language models (LLMs) have been shown to have significant potential in few-shot learning across various fields, even with minimal training data. However, their ability to generalize to unseen tasks in more complex fields, such as biology and medicine has yet to be fully evaluated. LLMs can offer a promising alternative approach for biological inference, particularly in cases where structured data and sample size are limited, by extracting prior knowledge from text corpora. Here we report our proposed few-shot learning approach, which uses LLMs to predict the synergy of drug pairs in rare tissues that lack structured data and features. Our experiments, which involved seven rare tissues from different cancer types, demonstrate that the LLM-based prediction model achieves significant accuracy with very few or zero samples. Our proposed model, the CancerGPT (with ~ 124M parameters), is comparable to the larger fine-tuned GPT-3 model (with ~ 175B parameters). Our research contributes to tackling drug pair synergy prediction in rare tissues with limited data, and also advancing the use of LLMs for biological and medical inference tasks.

  • Research Article
  • Cite Count Icon 42
  • 10.1038/s41598-025-96508-3
Comparing large Language models and human annotators in latent content analysis of sentiment, political leaning, emotional intensity and sarcasm
  • Apr 3, 2025
  • Scientific Reports
  • Ljubiša Bojić + 6 more

In the era of rapid digital communication, vast amounts of textual data are generated daily, demanding efficient methods for latent content analysis to extract meaningful insights. Large Language Models (LLMs) offer potential for automating this process, yet comprehensive assessments comparing their performance to human annotators across multiple dimensions are lacking. This study evaluates the inter-rater reliability, consistency, and quality of seven state-of-the-art LLMs. These include variants of OpenAI’s GPT-4, Gemini, Llama-3.1-70B, and Mixtral 8 × 7B. Their performance is compared to human annotators in analyzing sentiment, political leaning, emotional intensity, and sarcasm detection. The study involved 33 human annotators and eight LLM variants assessing 100 curated textual items. This resulted in 3,300 human and 19,200 LLM annotations. LLM performance was also evaluated across three-time points to measure temporal consistency. The results reveal that both humans and most LLMs exhibit high inter-rater reliability in sentiment analysis and political leaning assessments, with LLMs demonstrating higher reliability than humans. In emotional intensity, LLMs displayed higher reliability compared to humans, though humans rated emotional intensity significantly higher. Both groups struggled with sarcasm detection, evidenced by low reliability. Most LLMs showed excellent temporal consistency across all dimensions, indicating stable performance over time. This research concludes that LLMs, especially GPT-4, can effectively replicate human analysis in sentiment and political leaning, although human expertise remains essential for emotional intensity interpretation. The findings demonstrate the potential of LLMs for consistent and high-quality performance in certain areas of latent content analysis.

  • Research Article
  • Cite Count Icon 17
  • 10.1098/rsos.240180
Personality testing of large language models: limited temporal stability, but highlighted prosociality
  • Oct 1, 2024
  • Royal Society Open Science
  • Bojana Bodroža + 2 more

As large language models (LLMs) continue to gain popularity due to their human-like traits and the intimacy they offer to users, their societal impact inevitably expands. This leads to the rising necessity for comprehensive studies to fully understand LLMs and reveal their potential opportunities, drawbacks and overall societal impact. With that in mind, this research conducted an extensive investigation into seven LLMs, aiming to assess the temporal stability and inter-rater agreement on their responses on personality instruments in two time points. In addition, LLMs’ personality profile was analysed and compared with human normative data. The findings revealed varying levels of inter-rater agreement in the LLMs’ responses over a short time, with some LLMs showing higher agreement (e.g. Llama3 and GPT-4o) compared with others (e.g. GPT-4 and Gemini). Furthermore, agreement depended on used instruments as well as on domain or trait. This implies the variable robustness in LLMs’ ability to reliably simulate stable personality characteristics. In the case of scales which showed at least fair agreement, LLMs displayed mostly a socially desirable profile in both agentic and communal domains, as well as a prosocial personality profile reflected in higher agreeableness and conscientiousness and lower Machiavellianism. Exhibiting temporal stability and coherent responses on personality traits is crucial for AI systems due to their societal impact and AI safety concerns.

  • Research Article
  • 10.1200/op.2025.21.10_suppl.624
Leveraging large language models on personal computers for automated identification of lung cancer progression in radiology reports.
  • Oct 1, 2025
  • JCO Oncology Practice
  • Alexander Kim + 8 more

624 Background: Radiology reports are used to identify lung cancer treatment response for clinical care and research. Disease progression is typically documented as free text. Extracting instances of progression from numerous reports during chart review for clinical research requires time-intensive manual labor. Conversely, large language model (LLM) tools can quickly extract this information, but they demand significant computational resources such as graphics processing units (GPUs). We aimed to develop an LLM tool using structured “few-shot” prompting, in which examples are provided to optimize LLM output, to accurately identify lung cancer progression from radiology reports on a personal computer without a GPU. Methods: To generate a gold standard for review, we manually annotated 400 chest CT reports using RECIST 1.1 criteria. Each report listed measurements of all thoracic (i.e., lung, mediastinal, and chest wall) masses and indicated whether these lesions were smaller or larger than in the prior study. We tested three LLMs (Llama3-8B, Mistral-7B, and Phi3-4B) for mass size extraction, report summarization, and binary classification (stable vs. progression) using chain-of-thought reasoning. The model first extracted size changes for each lesion. The model was then queried to classify the report as indicating progression or stability, along with confidence scores. Predictions with low confidence scores were defaulted to stable. Models were compared using the same prompts, differing only in model architecture and parameter count. Results: 39% of reports indicated progression per gold standard labels. Among the models tested, Llama-3 achieved the best performance with an F1 score of 0.866 and overall accuracy of 82.5%. Of all the cases Llama-3 predicted as stable, 95.8% were actually stable. Mistral followed with F1 score of 0.807, an overall accuracy of 76.4%, and precision of stable predictions of 97.5%. Smaller models such as Phi-3 did not consistently provide the requested labels. Upon manual review of the cases incorrectly labeled by Llama-3, 60% of incorrect labels were applied to borderline cases (e.g., slight disease growth that did not meet RECIST criteria, or growth that the radiologist considered to be attributable to factors such as inflammation). Conclusions: LLMs utilizing few-shot prompting can reliably and efficiently identify lung cancer progression from free-text radiology reports on a personal computer. Llama-3, operating without GPU support, achieved high sensitivity and F1 scores, enabling accurate extraction of clinical outcomes from unstructured text. Locally deployed LLMs can enable scalable, cost-effective data abstraction for retrospective research and may offer practical tools for researchers without high-end computing infrastructure.

  • Research Article
  • 10.1161/circ.152.suppl_3.4369198
Abstract 4369198: Performance of Large Language Models in Analyzing Common Hypertension Scenarios in Clinical Practice
  • Nov 4, 2025
  • Circulation
  • Jaleh Zand + 7 more

Background/Objective: Hypertension is the most prevalent chronic disease in primary care and a leading cause of cardiovascular morbidity and mortality. Despite existing guidelines, therapeutic inertia and suboptimal control persist. Large language models (LLMs) offer a potential valuable addition to augment clinical decision-making, yet their reliability for guideline-driven tasks remains unverified. This study evaluated the accuracy and safety of hypertension management recommendations generated by three LLMs compared to expert responses. Methods: Fifty-one clinical vignettes representing 17 core hypertension management concepts were constructed by hypertension experts. Each case was submitted to three LLMs (GPT-4, Gemini, MedLM) and a hypertension expert also wrote the “gold standard” answers. Three blinded expert reviewers rated each response on a 4-point accuracy scale, a binary safety (safe/unsafe) scale, and attempted to identify the source (LLM vs. expert) providing the response. Ratings were analyzed using mean scores, percentages of accurate and safe responses, and inter-rater agreement. Results: GPT-4 had the highest accuracy (83%) and safety (86%) scores among LLMs but remained inferior to expert responses (92% accuracy, 93% safety). Gemini and MedLM performed significantly worse (accuracy: 64% and 35%; safety: 73% and 39%, respectively). GPT-4 generated the most guideline-concordant responses (46%) among the three LLMs (Gemini 35%, MedLM 14%), but remains lower than experts’ responses (68%). Evaluators misidentified LLM responses as expert-written in 10 to 25% of cases, particularly with GPT-4. Inter-rater reliability for accuracy ratings was highest for expert-generated responses (ICC 0.81), with progressively lower agreement for GPT-4 (0.76), Gemini (0.70), and MedLM (0.68). A similar pattern was observed for safety and source discrimination ratings. The agreement was strongest for safety assessments and weakest for source discrimination. Conclusion: Among three tested LLMs, GPT-4 demonstrated closer agreement to expert decisions thereby showing greater potential for supporting hypertension management. However, current LLMs’ versions frequently produce inaccurate or unsafe recommendations and remain inferior to expert judgment. Human-in-the-loop supervision remains essential when deploying LLMs for clinical decision-making.

  • Research Article
  • 10.2196/86630
Evaluation of GPT-5 for Esophageal Cancer Staging Using Fluorodeoxyglucose Positron Emission Tomography Maximum-Intensity Projection Images: Comparative Pilot Study.
  • Feb 23, 2026
  • JMIR cancer
  • Hiroki Maruyama + 7 more

Accurate esophageal cancer staging relies on 18F fluorodeoxyglucose positron emission tomography (18F FDG-PET), but its interpretation is complex and time-intensive. This diagnostic burden is exacerbated by significant workforce shortages in both radiology and surgery, thus necessitating automated support systems. The emergence of advanced large language models (LLMs) has raised expectations for their potential to fulfill this role in complex medical tasks. We evaluated the diagnostic accuracy of LLMs for staging esophageal cancer using 18F FDG-PET images, with a focus on their ability to assess lymph nodes (LNs; clinical N [cN]) and distant metastases (clinical M [cM]) for automated radiology reporting. This retrospective study included 120 consecutive adult patients who were diagnosed with esophageal squamous cell carcinoma and underwent 18F FDG-PET/computed tomography at Tohoku University Hospital between January 2019 and December 2021. Patients with prior treatment, nonsquamous cell carcinoma histology, or blood glucose levels ≥200 mg/dL were excluded. Frontal maximum-intensity projection positron emission tomography images were extracted, standardized, and analyzed along with information regarding the tumor location. Six LLMs (GPT-5, GPT-4.5, GPT-4.1, OpenAI-o3, -o1, and GPT-4 Turbo) and 4 blinded human evaluators (a nuclear medicine specialist, a gastrointestinal surgeon, and 2 radiology residents) assessed the presence of thoracic and abdominal LN metastases on a region-level basis and determined cN and cM staging on a patient-level basis. The model analyses were performed using the application programming interface in a zero-shot setting. Radiology reports served as the reference standard. Diagnostic agreement and accuracy were evaluated using Cohen κ and the Cochran Q test. Additionally, to account for the class imbalance in the dataset, the Matthews Correlation Coefficient was calculated as a robust metric for binary classification performance. Post hoc McNemar tests were performed with Bonferroni correction; statistical significance for pairwise comparisons was set at P<.0083 (adjusted from P<.05) using JMP Pro (version 18.0; SAS Institute Inc). The average accuracy was 41/120 (34%) to 94/120 (78%) for LLMs and 72/120 (60%) to 102/120 (85%) for physicians, with significantly higher accuracy for physicians (P<.05) in the thoracic LN, abdominal LN, and cN stages. Interrater reliability was slight to fair for LLMs (κ: -0.07 to 0.25) and fair to substantial for physicians (κ: 0.27 to 0.74). Matthews Correlation Coefficient scores were consistently higher for physicians (0.28 to 0.75) than for LLMs (-0.07 to 0.32). Among the LLMs, GPT-5 demonstrated the highest overall accuracy, with newer LLMs showing improved diagnostic accuracy when compared with previous models in identifying abdominal LN metastases and cM staging, though they showed weaker consistency for cN staging. For example, in thoracic LN detection, GPT-5 achieved 76/120 (63%) accuracy, whereas other LLMs achieved 72/120 (60%) or lower accuracy. Although current LLMs have not yet reached physician-level accuracy in comprehensive staging, recent models show promise in assisting with specific diagnostic tasks.

  • Research Article
  • Cite Count Icon 2
  • 10.1371/journal.pdig.0000980
Development and evaluation of large-language models (LLMs) for oncology: A scoping review
  • Aug 7, 2025
  • PLOS Digital Health
  • Namya Mehan + 2 more

Large language models (LLMs), a significant development in artificial intelligence (AI), are continuing to demonstrate seminal improvement in performance for various text analysis and generation tasks. There are limited systematic studies on LLM applications that were developed/evaluated in relevance to oncology. Our scoping review explores applications of LLMs in oncology to determine (1) the nature of LLM applications relevant to a cancer/tumor type, (2) the phases of cancer care addressed by the LLMs, (3) which LLMs were used in these applications, (4) the sources and pre-processing of datasets used, (5) the techniques used to optimize the performance of LLMs, (6) the methods of evaluation, and (7) the common limitations noted by the authors of these LLM applications and to study their implications in research and practice. A librarian-assisted search was performed across the following databases: Association for Computing Machinery (ACM), Embase, Engineering Village, IEEE Xplore, Medline, Scopus, SPIE and Web of Science till Jan 12, 2024. Pre-prints from this search were considered if they were published/accepted by Feb 29, 2024. From the initial search of 14863 articles, 60 were finally included. Our results demonstrated that LLMs were mostly evaluated across a diverse set of oncology-related applications. Generative pre-trained transformer (GPT)-based LLMs were mostly used. In the subset of studies where the phase(s) of cancer care was/were provided or implied, treatment and diagnosis were the most included phases. Data for development and evaluation extended from patient health records, synthetic patient records, research and professional society publications to social media. Prompt-designing and engineering were performed as data pre-processing steps in several studies. Clinicians, trainees, researchers, and patients were among the variety of users targeted by the applications. In the17% studies that developed LLMs for oncological aspects, domain adaptation through pre-training and fine-tuning were often performed and resulted in performance improvement. The evaluation of an LLM’s performance involved usage of both standard, validated, non-standardized, and/or customized performance measures considering a variety of constructs, other than accuracy. Six primary themes emerged as limitations including limitation of generalizability/applicability, sample size, bias and subjectivity, and evaluation metrics. This review highlights that LLMs, specific to oncological aspects, are less common than general-purpose LLMs. The application areas were heterogeneous, used diverse data sources, were directed towards a variety of users, and resulted in variety of evaluation methods. Despite the diversity of LLM applications in oncology, future research needs to address the limited generalizability of these applications, mitigation of bias and subjectivity, and standardization of evaluation methodologies. Future applications of LLMs in oncology should include developing oncology-specific LLMs that can mitigate knowledge gaps and extend to diverse areas of oncology training and practice not considered so far.

  • Research Article
  • Cite Count Icon 8
  • 10.2196/79202
Large Language Models for Health Care Text Classification: Systematic Review.
  • Jun 16, 2025
  • JMIR AI
  • Hajar Sakai + 1 more

Large language models (LLMs) have fundamentally transformed approaches to natural language processing tasks across diverse domains. In health care, accurate and cost-efficient text classification is crucial-whether for clinical note analysis, diagnosis coding, or other related tasks-and LLMs present promising potential. Text classification has long faced multiple challenges, including the need for manual annotation during training, the handling of imbalanced data, and the development of scalable approaches. In health care, additional challenges arise, particularly the critical need to preserve patient data privacy and the complexity of medical terminology. Numerous studies have leveraged LLMs for automated health care text classification and compared their performance with traditional machine learning-based methods, which typically require embedding, annotation, and training. However, existing systematic reviews of LLMs either do not specialize in text classification or do not focus specifically on the health care domain. This research synthesizes and critically evaluates the current evidence in the literature on the use of LLMs for text classification in health care settings. Major databases (eg, Google Scholar, Scopus, PubMed, ScienceDirect) and other resources were queried for papers published between 2018 and 2024, following the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, resulting in 65 eligible research articles. These studies were categorized by text classification type (eg, binary classification, multilabel classification), application (eg, clinical decision support, public health and opinion analysis), methodology, type of health care text, and the metrics used for evaluation and validation. The systematic review includes 65 research articles published between 2020 and Q3 2024, showing a significant increase in publications over time, with 28 papers published in Q1-Q3 2024 alone. Fine-tuning was the most common LLM-based approach (35 papers), followed by prompt engineering (17 papers). BERT (Bidirectional Encoder Representations from Transformers) variants were predominantly used for multilabel classification (50%), whereas closed-source LLMs were most commonly applied to binary (44.0%) and multiclass (30.6%) classification tasks. Clinical decision support was the most frequent application (29 papers). Over 80% of studies used English-language datasets, with clinical notes being the most common text type. All studies employed accuracy-related metrics for evaluation, and the findings consistently showed that LLMs outperformed traditional machine learning approaches in health care text classification tasks. This review identifies existing gaps in the literature and highlights future research directions for further investigation.

  • Research Article
  • 10.1037/tra0002087
Can large language models enhance assessment of Criterion A for PTSD from self-report?
  • Dec 1, 2025
  • Psychological trauma : theory, research, practice and policy
  • Mikael Rubin + 4 more

Based on the Diagnostic and Statistical Manual of Mental Disorders, fifth edition, posttraumatic stress disorder (PTSD) involves assessing whether a traumatic event meets Criterion A, which is necessary to establish symptom severity and a potential PTSD diagnosis. As research moves online, methods to establish Criterion A have varied widely, influencing the accuracy and consistency of PTSD diagnoses. Literature suggests relying solely on self-assessment of trauma experiences may be problematic. This study evaluated whether integration of large language models (LLMs) directly into online self-report data collection could enhance assessment of Criterion A for PTSD. The present study leveraged LLMs to test a new method for enhancing the online assessment of Criterion A from self-report. Adults (N = 110) completed the extended Life Events Checklist for the Diagnostic and Statistical Manual of Mental Disorders, fifth edition. An LLM was integrated directly into the online survey tool Qualtrics and utilized via Application Programming Interface to code text descriptions to actively follow up with participants by providing additional questions/prompts. Four clinician raters independently evaluated the text descriptions after data collection was complete to determine the proportion of individuals meeting Criterion A and to establish interrater reliability with LLMs. The percentage of participants who met Criterion A based on clinician ratings was increased from an average of 65% (range: 59%-71%) at the first description to an average of 86% across all follow-up clarifications. However, interrater reliability of LLMs with clinician raters was only fair, original LLM mean κ = 0.26 (κ range: 0.18-0.46), newer LLM mean κ = 0.35 (κ range: 0.23-0.47). Findings suggest that use of LLMs for enhancing Criterion A assessment led to increased information from participants, leading to greater reporting of events meeting Criterion A. However, LLMs did not provide determination of Criterion A on par with clinicians. Findings highlight the need for further assessment of integrating LLMs into online research or treatment. (PsycInfo Database Record (c) 2025 APA, all rights reserved).

  • Research Article
  • 10.2196/84668
Disclaimers and Referral Patterns for Medical Advice Across Urgency Levels: Large Language Model Evaluation Study
  • Mar 16, 2026
  • Journal of Medical Internet Research
  • Florian Reis + 6 more

Background“I’m not a doctor, but...” is a typical response when asking considerate laypeople for health advice. However, seeking medical advice has also shifted to digital settings, where the expertise of the other party is less transparent than in face-to-face interactions. Recently, large language models (LLMs) have emerged as easily accessible tools, offering a novel way to formulate medical questions and receive seemingly qualified advice. Given the sensitive nature of health-related queries and the lack of professional supervision, incorrect advice can pose serious health risks. Therefore, including explicit disclaimers and precise referrals in LLM responses to medical queries is crucial. However, little is known about how LLMs adapt their safety implementations in response to different urgency levels.ObjectiveThis study evaluates disclaimer and referral patterns in responses from LLMs to authentic medical queries of different urgency levels using a systematic evaluation framework.MethodsThis prospective, multimodel evaluation study generated and analyzed 908 responses from 4 popular LLMs (GPT-4o, Claude Sonnet-4, Grok-3, and DeepSeek-V3) to 227 authentic patient queries from a public dataset. Two human raters classified all 227 patient queries using a 3-level urgency scale. LLM responses were evaluated using a 5-point ordinal classification system for disclaimer and referral advice, ranging from “no disclaimer” to “urgent advice to consult a medical professional.” GPT-4o served as the primary rater model for this task after conducting a subset validation against human expert annotations. Statistical analyses included Jonckheere-Terpstra tests to examine the relationship between case urgency and disclaimer ratings and Kruskal-Wallis tests for intermodel comparisons.ResultsThe 227 patient queries were distributed as 77 (34%) low-urgency, 110 (48%) intermediate-urgency, and 40 (18%) high-urgency cases. All 4 LLMs demonstrated statistically significant ordered trends (all P<.001), with higher-urgency queries receiving more explicit referral advice. Disclaimer and referral advice clustered toward higher categories across all models, with 97% (881/908) of responses indicating that a medical professional should be consulted. Sonnet-4, Grok-3, and GPT-4o demonstrated a conservative approach, with 89%, 89%, and 88%, respectively, of their responses being either explicit or urgent referrals. In contrast, DeepSeek-V3 showed a broader distribution, with 74% of responses falling into these categories. Interrater reliability between GPT-4o and human raters achieved moderate to substantial agreement, with weighted Cohen κ values between 0.415 and 0.707.ConclusionsCurrent LLMs exhibit urgency-responsive safety mechanisms when providing medical advice. All evaluated models adaptively incorporate more explicit disclaimers and urgent referrals for higher-urgency queries. However, variability between LLMs highlights the need for standardized safety measures and appropriate regulatory frameworks. Although these findings indicate progress regarding safety concerns, the public availability of LLMs requires careful consideration to ensure consistent protection against patient harm while preserving the benefits of low-threshold access to health information.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 9
  • 10.1001/jamanetworkopen.2025.6359
Semantic Clinical Artificial Intelligence vs Native Large Language Model Performance on the USMLE
  • Apr 22, 2025
  • JAMA Network Open
  • Peter L Elkin + 9 more

Large language models (LLMs) are being implemented in health care. Enhanced accuracy and methods to maintain accuracy over time are needed to maximize LLM benefits. To evaluate whether LLM performance on the US Medical Licensing Examination (USMLE) can be improved by including formally represented semantic clinical knowledge. This comparative effectiveness research study was conducted between June 2024 and February 2025 at the Department of Biomedical Informatics, Jacobs School of Medicine and Biomedical Sciences, University at Buffalo, Buffalo, New York, using sample questions from the USMLE Steps 1, 2, and 3. Semantic clinical artificial intelligence (SCAI) was developed to insert formally represented semantic clinical knowledge into LLMs using retrieval augmented generation (RAG). The SCAI method was evaluated by comparing the performance of 3 Llama LLMs (13B, 70B, and 405B; Meta) with and without SCAI RAG on text-based questions from the USMLE Steps 1, 2, and 3. LLM accuracy for answering questions was determined by comparing the LLM output with the USMLE answer key. The LLMs were tested on 87 questions in the USMLE Step 1, 103 in Step 2, and 123 in Step 3. The 13B LLM enhanced by SCAI RAG was associated with significantly improved performance on Steps 1 and 3 but only met the 60% passing threshold on Step 3 (74 questions correct [60.2%]). The 70B and 405B LLMs passed all the USMLE steps with and without SCAI RAG. The SCAI RAG 70B model scored 80 questions (92.0%) correctly on Step 1, 82 (79.6%) on Step 2, and 112 (91.1%) on Step 3. The SCAI RAG 405B model scored 79 (90.8%) correctly on Step 1, 87 (84.5%) on Step 2, and 117 (95.1%) on Step 3. Significant improvements associated with SCAI RAG were found for the 13B model on Steps 1 and 3, the 70B model on Step 2, and the 405B parameter model on Step 3. The 70B model was significantly better than the 13B model, and the 405B model was not significantly better than the 70B model. In this comparative effectiveness research study, SCAI RAG was associated with significantly improved scores on the USMLE Steps 1, 2, and 3. The 13B model passed Step 3 with RAG, and the 70B and 405B models passed and scored well on Steps 1, 2, and 3 with or without augmentation. New forms of reasoning by LLMs, like semantic reasoning, have potential to improve the accuracy of LLM performance on important medical questions. Improving LLM performance in health care with targeted, up-to-date clinical knowledge is an important step in LLM implementation and acceptance.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant