Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Large language model-based screening of substances and their composition from safety data sheets for high-resolution chemical exposure assessment.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

SDSs provide information on chemical substances in professional and consumer products. Large language models (LLMs) offer a rapid approach to screening chemical information from SDSs. This study aimed to validate the performance of LLMs in accurately extracting substance information from SDSs of products. Chemical information was extracted from the SDSs of cleaning products using the LLMs ChatGPT-4o and Gemini 2.5 Pro. The performance of the LLMs was evaluated against manually extracted data using precision, sensitivity, and F1 score. A total of 301 substance-composition combinations across 59 products were included in the validation. The Gemini 2.5 Pro model showed a higher F1 score (1.00) than ChatGPT-4o (0.94). LLMs enable high-throughput chemical information extraction from SDSs, reducing the burden of manual screening and supporting large-scale combined-exposure assessments. The accurate identification of chemical substances in professional and consumer products is a major challenge in assessing complex chemical exposures in exposure science and epidemiology. This study validated the use of Large Language Models (LLMs) as a novel, rapid, and highly accurate method of extracting high-throughput chemical data from multilingual Safety Data Sheets. Our findings illustrated that LLMs could be effectively used to overcome the labor-intensive and time-consuming limitations of manual data screening. This approach allows high-throughput analysis, enabling comprehensive combined-exposure assessments in large-scale epidemiological and exposure assessment studies. Additionally, as LLMs can extract data in different languages, they can facilitate international research collaboration.

Similar Papers
  • Research Article
  • 10.1016/j.jbi.2026.105034
A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning
  • Mar 27, 2026
  • Journal of biomedical informatics
  • Cheng Peng + 5 more

A Study of Large Language Models for Patient Information Extraction: Model Architecture, Fine-Tuning Strategy, and Multi-task Instruction Tuning

  • Research Article
  • 10.65521/ijacect.v14i3s.1636
An Iterative Comprehensive Evaluation of Large Language and Vision Models in Medical AI: Benchmarks, Adaptability, and Deployment Challenges
  • Dec 22, 2025
  • International Journal on Advanced Computer Engineering and Communication Technology
  • Wani H Bisen + 1 more

PRISMA principles provide a thorough analysis of current advances in large language models (LLMs) and multimodal transformers for medical applications. As LLMs like GPT-4, BioGPT, Med-PaLM, and hybrid frameworks like COMCARE enter clinical processes, thorough synthesis is essential to increase performance, methodological adaptability, and implementation practicality in many healthcare situations. Their creativity in medical report writing, decision support, and diagnosis is notable, but the literature has not established a cohesive taxonomy that evaluates these models by uniform metrics, domain-specific generalizability, and ethical acceptability. Over 40 studies examined radiology report production, clinical question responding, cognitive assessment, and causal reasoning. After testing vision-language transformer architectures like PEGASUS and ETB MII for automated imaging-based reporting, graph-based reasoning was used to evaluate drug safety and interpretability of knowledge- integrated models like KELLM. As needed, BLEU, ROUGE, F1 score, CIDEr, and qualitative evaluations were used. Domain- adapted and hybrid models improve diagnostic accuracy, task- specific explainability, and clinician workload differently. Model illusion, biases, hostile manipulation, and resource-intensive fine- tuning persist. The report recommends strong benchmarking, public evaluation standards, and ethical frameworks for LLMs in high-stakes medical applications. This study defines LLMs' therapeutic utility and recommends infrastructure, ethics, and technology for safe and successful integration. This effort prepares scalable, interpretable, and equitable medical AI systems.

  • Research Article
  • Cite Count Icon 84
  • 10.1001/jamanetworkopen.2024.12687
Assessing the Risk of Bias in Randomized Clinical Trials With Large Language Models
  • May 22, 2024
  • JAMA Network Open
  • Honghao Lai + 17 more

Large language models (LLMs) may facilitate the labor-intensive process of systematic reviews. However, the exact methods and reliability remain uncertain. To explore the feasibility and reliability of using LLMs to assess risk of bias (ROB) in randomized clinical trials (RCTs). A survey study was conducted between August 10, 2023, and October 30, 2023. Thirty RCTs were selected from published systematic reviews. A structured prompt was developed to guide ChatGPT (LLM 1) and Claude (LLM 2) in assessing the ROB in these RCTs using a modified version of the Cochrane ROB tool developed by the CLARITY group at McMaster University. Each RCT was assessed twice by both models, and the results were documented. The results were compared with an assessment by 3 experts, which was considered a criterion standard. Correct assessment rates, sensitivity, specificity, and F1 scores were calculated to reflect accuracy, both overall and for each domain of the Cochrane ROB tool; consistent assessment rates and Cohen κ were calculated to gauge consistency; and assessment time was calculated to measure efficiency. Performance between the 2 models was compared using risk differences. Both models demonstrated high correct assessment rates. LLM 1 reached a mean correct assessment rate of 84.5% (95% CI, 81.5%-87.3%), and LLM 2 reached a significantly higher rate of 89.5% (95% CI, 87.0%-91.8%). The risk difference between the 2 models was 0.05 (95% CI, 0.01-0.09). In most domains, domain-specific correct rates were around 80% to 90%; however, sensitivity below 0.80 was observed in domains 1 (random sequence generation), 2 (allocation concealment), and 6 (other concerns). Domains 4 (missing outcome data), 5 (selective outcome reporting), and 6 had F1 scores below 0.50. The consistent rates between the 2 assessments were 84.0% for LLM 1 and 87.3% for LLM 2. LLM 1's κ exceeded 0.80 in 7 and LLM 2's in 8 domains. The mean (SD) time needed for assessment was 77 (16) seconds for LLM 1 and 53 (12) seconds for LLM 2. In this survey study of applying LLMs for ROB assessment, LLM 1 and LLM 2 demonstrated substantial accuracy and consistency in evaluating RCTs, suggesting their potential as supportive tools in systematic review processes.

  • Research Article
  • 10.1136/bmjhci-2025-101956
Open-source large language model-based on-premises pipeline for automated data extraction from unstructured electronic health records: a pilot study
  • Jun 1, 2026
  • BMJ Health & Care Informatics
  • Vasileios Ntinopoulos + 4 more

ObjectivesWe evaluated an on-premises, open-source large language model (LLM)-based data extraction pipeline for automated data extraction from unstructured electronic health records (EHRs).MethodsAutomated script-based EHR data preprocessing extracted 50 medical texts in German, which were entered into the LLM pipeline. 4-bit and 8-bit quantizations of 14 mid-sized LLMs (30B–90B parameters) were evaluated in 6 information extraction, 11 binary classification and 5 multilevel classification tasks comprising all variables of the European System for Cardiac Operative Risk Evaluation 2 (EuroSCORE 2) model for 1100 predictions each. LLM response consistency was assessed over three same-prompt iterations.ResultsIn overall accuracy, Qwen3-30b-a3b-q8 presented the highest value (0.954) and 13 LLMs had values over 0.90. In information extraction accuracy, 12 LLMs exhibited a value of 1.0 and all 14 LLMs had values over 0.96. In binary classification accuracy, Llama3.2-vision-90b-q4 exhibited the highest value (0.972), 5 LLMs had values of at least 0.95 and all 14 LLMs showed values over 0.93. In multilevel classification accuracy, Qwen3-30b-a3b-q8 exhibited the highest value (0.940), four LLMs had values over 0.90 and all LLMs presented values over 0.80. Nine LLMs exhibited perfect response consistency and the remaining five LLMs had a Krippendorff’s alpha value of 0.999.DiscussionMultiple LLMs exhibited high accuracy in information extraction, binary classification, multilevel classification and response consistency and seem able to reliably automate data extraction from EHRs.ConclusionThis pilot study demonstrates the feasibility of on-premises, privacy-preserving, LLM-based automated EHR data extraction pipelines. Larger-scope studies are warranted to validate their potential in healthcare.

  • Research Article
  • Cite Count Icon 79
  • 10.1016/j.jclinepi.2025.111746
Large language models for conducting systematic reviews: on the rise, but not yet ready for use-a scoping review.
  • May 1, 2025
  • Journal of clinical epidemiology
  • Judith-Lisa Lieberum + 8 more

Machine learning promises versatile help in the creation of systematic reviews (SRs). Recently, further developments in the form of large language models (LLMs) and their application in SR conduct attracted attention. We aimed at providing an overview of LLM applications in SR conduct in health research. We systematically searched MEDLINE, Web of Science, IEEEXplore, ACM Digital Library, Europe PMC (preprints), Google Scholar, and conducted an additional hand search (last search: February 26, 2024). We included scientific articles in English or German, published from April 2021 onwards, building upon the results of a mapping review that has not yet identified LLM applications to support SRs. Two reviewers independently screened studies for eligibility; after piloting, 1 reviewer extracted data, checked by another. Our database search yielded 8054 hits, and we identified 33 articles from our hand search. We finally included 37 articles on LLM support. LLM approaches covered 10 of 13 defined SR steps, most frequently literature search (n = 15, 41%), study selection (n = 14, 38%), and data extraction (n = 11, 30%). The mostly recurring LLM was Generative Pretrained Transformer (GPT) (n = 33, 89%). Validation studies were predominant (n = 21, 57%). In half of the studies, authors evaluated LLM use as promising (n = 20, 54%), one-quarter as neutral (n = 9, 24%) and one-fifth as nonpromising (n = 8, 22%). Although LLMs show promise in supporting SR creation, fully established or validated applications are often lacking. The rapid increase in research on LLMs for evidence synthesis production highlights their growing relevance. Systematic reviews are a crucial tool in health research where experts carefully collect and analyze all available evidence on a specific research question. Creating these reviews is typically time- and resource-intensive, often taking months or even years to complete, as researchers must thoroughly search, evaluate, and synthesize an immense number of scientific studies. For the present article, we conducted a review to understand how new artificial intelligence (AI) tools, specifically large language models (LLMs) like Generative Pretrained Transformer (GPT), can be used to help create systematic reviews in health research. We searched multiple scientific databases and finally found 37 relevant articles. We found that LLMs have been tested to help with various parts of the systematic review process, particularly in 3 main areas: searching scientific literature (41% of studies), selecting relevant studies (38%), and extracting important information from these studies (30%). GPT was the most commonly used LLM, appearing in 89% of the studies. Most of the research (57%) focused on testing whether these AI tools actually work as intended in this context of systematic review production. The results were mixed: about half of the studies found LLMs promising, a quarter were neutral, and one-fifth found them not promising. While LLMs show potential for making the systematic review process more efficient, there is still a lack of fully tested and validated applications. However, the increasing number of studies in this field suggests that these AI tools are becoming increasingly important in creating systematic reviews.

  • Research Article
  • Cite Count Icon 2
  • 10.1016/j.jclinepi.2026.112221
Large language models show promising performance for some systematic review tasks but call for cautious implementation: a systematic review.
  • Jun 1, 2026
  • Journal of clinical epidemiology
  • Florian Laignelot + 9 more

With the exponential growth of biomedical literature, the challenge of conducting systematic reviews is becoming increasingly burdensome. We aimed to evaluate the performance of large language models (LLMs) in the automation of some or all steps of systematic reviews and meta-analyses. In this systematic review, we searched PubMed, Embase, the Cochrane Library and preprint platforms up to January 14, 2025. We included any studies assessing the performance of LLMs (eg, generative pre-trained transformer [GPT], Claude, Mistral) in any step of the systematic review process. Pairs of reviewers independently extracted data and assessed risk of bias. We conducted analyses using median (interquartile range [IQR]) for positive (PPA) and negative percent agreements (NPA), respectively, analogous to sensitivity and specificity, between LLMs and human reviewers. From 3889 unique references, we included 63 studies of which 52 reporting performance metrics for a total of 148 LLM performance assessments. Most assessments concerned GPT models (n = 114, 77%). The most frequently evaluated tasks were title and abstract screening (n = 78, 53%), data extraction (n = 23, 16%), and full-text screening (n = 20, 14%). For title and abstract screening, overall median PPA was 0.92 (IQR 0.69-0.98) and median NPA was 0.89 (0.72-0.95). For full-text screening, the overall median PPA was 0.93 (0.87-1.00) and median NPA was 0.92 (0.78-0.97). Late-generation LLMs released after GPT-4 seemed to provide higher performance than earlier models. For other tasks, authors reported overall good performances, but variability of performance metrics precluded complete quantitative synthesis. Global accuracy for data extraction tasks ranged from 0.36 to 1.00, with a median accuracy of 0.95 (IQR 0.91-0.97, n = 11). For the "risk of bias assessment" task, accuracy ranged from 0.44 to 0.90 (median = 0.62, IQR 0.53-0.76, n = 6). The performance of LLMs, particularly newer generations, shows promise in automating some repetitive steps of systematic reviews such as screening. However, their successful integration will require appropriate safeguards and careful implementation. Systematic reviews are one of the most reliable ways to answer medical and public health questions. They bring together all available studies on a topic and help clinicians and policymakers make informed decisions. However, producing a high-quality systematic review takes a lot of time and effort. Whole teams of researchers spend months screening thousands of articles, extracting data, and double-checking results. With little more than a million of new publications every year, keeping reviews up to date is becoming increasingly difficult. LLMs, such as ChatGPT, may help reduce this workload. These tools can read and summarize text and might assist with repetitive tasks like selecting relevant studies or extracting information from articles. But it is still unclear how reliable these tools are for research purposes. This is the first systematic review to assess LLMs' performance to facilitate systematic reviews. We sought to review all studies that tested LLMs in the different steps of systematic reviews and found 63 studies evaluating how well these tools performed compared with human reviewers. Overall, LLMs showed good agreement with humans for tasks such as screening titles and abstracts, and full-text articles. Newer models seemed to perform better than older ones. However, performance was more variable for complex tasks that require interpretation, such as extracting detailed data or assessing methodological quality. Our findings suggest that LLMs could help researchers work faster and make systematic reviews more efficient. However, they are not ready to replace human judgment. These tools can make mistakes, produce inconsistent results, or generate inaccurate information if not carefully supervised. In practice, LLMs should be used as assistants rather than substitutes. With proper safeguards, transparent reporting, and human oversight, they may become valuable tools to support evidence-based healthcare and help keep research up to date.

  • Research Article
  • Cite Count Icon 16
  • 10.2196/67488
Accuracy of Large Language Models for Literature Screening in Thoracic Surgery: Diagnostic Study.
  • Mar 11, 2025
  • Journal of medical Internet research
  • Zhang-Yi Dai + 6 more

Systematic reviews and meta-analyses rely on labor-intensive literature screening. While machine learning offers potential automation, its accuracy remains suboptimal. This raises the question of whether emerging large language models (LLMs) can provide a more accurate and efficient approach. This paper evaluates the sensitivity, specificity, and summary receiver operating characteristic (SROC) curve of LLM-assisted literature screening. We conducted a diagnostic study comparing the accuracy of LLM-assisted screening versus manual literature screening across 6 thoracic surgery meta-analyses. Manual screening by 2 investigators served as the reference standard. LLM-assisted screening was performed using ChatGPT-4o (OpenAI) and Claude-3.5 (Anthropic) sonnet, with discrepancies resolved by Gemini-1.5 pro (Google). In addition, 2 open-source, machine learning-based screening tools, ASReview (Utrecht University) and Abstrackr (Center for Evidence Synthesis in Health, Brown University School of Public Health), were also evaluated. We calculated sensitivity, specificity, and 95% CIs for the title and abstract, as well as full-text screening, generating pooled estimates and SROC curves. LLM prompts were revised based on a post hoc error analysis. LLM-assisted full-text screening demonstrated high pooled sensitivity (0.87, 95% CI 0.77-0.99) and specificity (0.96, 95% CI 0.91-0.98), with the area under the curve (AUC) of 0.96 (95% CI 0.94-0.97). Title and abstract screening achieved a pooled sensitivity of 0.73 (95% CI 0.57-0.85) and specificity of 0.99 (95% CI 0.97-0.99), with an AUC of 0.97 (95% CI 0.96-0.99). Post hoc revisions improved sensitivity to 0.98 (95% CI 0.74-1.00) while maintaining high specificity (0.98, 95% CI 0.94-0.99). In comparison, the pooled sensitivity and specificity of ASReview tool-assisted screening were 0.58 (95% CI 0.53-0.64) and 0.97 (95% CI 0.91-0.99), respectively, with an AUC of 0.66 (95% CI 0.62-0.70). The pooled sensitivity and specificity of Abstrackr tool-assisted screening were 0.48 (95% CI 0.35-0.62) and 0.96 (95% CI 0.88-0.99), respectively, with an AUC of 0.78 (95% CI 0.74-0.82). A post hoc meta-analysis revealed comparable effect sizes between LLM-assisted and conventional screening. LLMs hold significant potential for streamlining literature screening in systematic reviews, reducing workload without sacrificing quality. Importantly, LLMs outperformed traditional machine learning-based tools (ASReview and Abstrackr) in both sensitivity and AUC values, suggesting that LLMs offer a more accurate and efficient approach to literature screening.

  • Research Article
  • 10.1186/s12911-026-03366-8
Evaluating large language models for clinical note processing: local fine-tuning and internal-external validation using electronic health records from South Asia.
  • Feb 25, 2026
  • BMC medical informatics and decision making
  • Seyed Alireza Hasheminasab + 18 more

Large Language Models (LLMs) hold the potential for clinical task-shifting by processing unstructured clinical text, enabling tasks such as clinical concept extraction and medical question answering from electronic health records. If implemented reliably, such approaches may benefit over-burdened healthcare systems, particularly in resource-limited settings and for traditionally overlooked populations, provided that local fine-tuning is supported by appropriate clinical and technical expertise. However, this powerful technology remains largely understudied in real-world contexts, particularly in the Global South. This study aims to assess whether openly available LLMs can be used reliably for processing medical notes in real-world settings in South Asia. We used publicly available LLMs to parse de-identified clinical notes from a large electronic health records (EHR) database in Pakistan, containing hospital records for 8.2 million patients. ChatGPT (GPT-3.5) as a general-purpose LLM, and GatorTron (base), BioMegatron, BioBert and ClinicalBERT as medical LLMs were evaluated when applied to these data, after fine-tuning them with (a) publicly available clinical datasets namely Informatics for Integrating Biology & the Bedside (I2B2) and National NLP Clinical Challenges (N2C2) for medical concept extraction (MCE) and emrQA for medical question answering (MQA), and (b) the local Pakistani de-identified EHR dataset, which includes inpatient Discharge Summaries (DS) and Subjective, Objective, Assessment, and Plan (SOAP) notes, as detailed in this paper. MCE models were applied to these clinical notes using both 3-label and 9-label formats, while MQA models were applied to medical questions. Internal and external validation performance was measured for (a) and (b) using F1 score, precision, recall, and accuracy for MCE and BLEU and ROUGE-L, which measure lexical and sequence similarity, for MQA. When clinical LLMs were not fine-tuned on the local EHR dataset, their performance during external validation on local data was notably poorer compared to internal validation on the dataset used for fine-tuning, with reductions of at least 15% in F1 scores for MCE and 35% in ROUGE-L and BLEU scores for MQA tasks. This suggests potential bias and highlights the inability of the medical LLMs to reliably handle the data distribution of the local population without further fine-tuning and adaptation. This trend persisted across two distinct natural language processing tasks: concept extraction and question answering, spanning a spectrum of task complexities. However, fine-tuning the LLMs with local EHR data significantly improved model performance across both tasks, yielding a 7.5% to 15% increase in the F1 score for MCE and a 27% to 53% increase in ROUGE-L and BLEU scores for MQA. Notably, ChatGPT, as a general-purpose LLM, stood out as an exception, demonstrating superior performance across all measured metrics on the local dataset compared to the publicly available dataset, with improvements ranging from 3% to 17% on the local EHR dataset, even without fine-tuning on the local data. Publicly available LLMs, predominantly trained on data from high-income regions, were found to be unreliable when applied in a real-world clinical setting in Pakistan. Fine-tuning them with local EHR data and regional clinical contexts improved their reliability, demonstrating a feasible adaptation strategy that is substantially less resource-intensive than training large language models from scratch. Close collaboration between local clinical and technical experts to curate and leverage more representative, inclusive, and unbiased medical datasets, can play a crucial role in further ensuring reliability of LLMs for resource-limited, overburdened settings, to be used in ways that are safe, fair, and beneficial for all.

  • Research Article
  • Cite Count Icon 2
  • 10.2196/73605
Large Language Model Versus Manual Review for Clinical Data Curation in Breast Cancer: Retrospective Comparative Study
  • Nov 6, 2025
  • JMIR Medical Informatics
  • Young-Joon Kang + 11 more

BackgroundManual review of electronic health records for clinical research is labor-intensive and prone to reviewer-dependent variations. Large language models (LLMs) offer potential for automated clinical data extraction; however, their feasibility in surgical oncology remains underexplored.ObjectiveThis study aimed to evaluate the feasibility and accuracy of LLM-based processing compared with manual physician review for extracting clinical data from breast cancer records.MethodsWe conducted a retrospective comparative study analyzing breast cancer records from 5 academic hospitals (January 2019-December 2019). Two data extraction pathways were compared: (1) manual physician review with direct electronic health record access (group 1: 1366/3100, 44.06%) and (2) LLM-based processing using Claude 3.5 Sonnet (Anthropic) on deidentified data automatically extracted through a clinical data warehouse platform (group 2: 1734/3100, 55.94%). The automated extraction system provided prestructured, deidentified data sheets organized by clinical domains, which were then processed by the LLM. The LLM prompt was developed through a 3-phase iterative process over 2 days. Primary outcomes included missing value rates, extraction accuracy, and concordance between groups. Secondary outcomes included comparison with the Korean Breast Cancer Society national registry data, processing time, and resource use. Validation involved 50 stratified random samples per group (900 data points each), assessed by 4 breast surgical oncologists. Statistical analysis included chi-square tests, 2-tailed t tests, Cohen κ, and intraclass correlation coefficients. The accuracy threshold was set at 90%.ResultsThe LLM achieved 90.8% (817) accuracy in validation analysis. Missing data patterns differed between groups: group 2 showed better lymph node documentation (missing: 152/1734, 8.76% vs 294/1366, 21.52%) but higher missing rates for cancer staging (211/1734, 12.17% vs 43/1366, 3.15%). Both groups demonstrated similar breast-conserving surgery rates (1107/1734, 63.84% vs 868/1366, 63.54%). Processing efficiency differed substantially: LLM processing required 12 days with 2 physicians versus 7 months with 5 physicians for manual review, representing a 91% reduction in physician hours (96 h vs 1025 h). The LLM group captured significantly more survival events (41 vs 11; P=.002). Stage distribution in the LLM group aligned better with national registry data (Cramér V=0.03 vs 0.07). Application programming interface costs totaled US $260 for 1734 cases (US $0.15 per case).ConclusionsLLM-based curation of automatically extracted, deidentified clinical data demonstrated comparable effectiveness to manual physician review while reducing processing time by 95% and physician hours by 91%. This 2-step approach—automated data extraction followed by LLM curation—addresses both privacy concerns and efficiency needs. Despite limitations in integrating multiple clinical events, this methodology offers a scalable solution for clinical data extraction in oncology research. The 90.8% accuracy rate and superior capture of survival events suggest that combining automated data extraction systems with LLM processing can accelerate retrospective clinical research while maintaining data quality and patient privacy.

  • Research Article
  • 10.1016/j.esmorw.2026.100718
Validating large language model\u2013assisted data extraction from clinical notes
  • May 12, 2026
  • ESMO Real World Data and Digital Oncology
  • J.W Van Koevorden + 12 more

Validating large language model\u2013assisted data extraction from clinical notes

  • Research Article
  • Cite Count Icon 1
  • 10.1158/1557-3265.aimachine-b006
Abstract B006: Using large language models for scalable extraction of real-world progression events across multiple cancer types
  • Jul 10, 2025
  • Clinical Cancer Research
  • Aaron B Cohen + 9 more

Background: Accurate identification of cancer progression events from electronic health records (EHRs) can help enable promising oncology applications such as predicting disease trajectory, assessing treatment efficacy, and generating real-world evidence. These use cases require both large-scale and high-quality data but manual abstraction of real-world progression (rwP) is time-intensive, difficult to scale, and inherently challenging given the varied and unstructured ways it can be documented across cancer types. Large language models (LLMs) offer a scalable alternative, but their accuracy relative to expert human abstractors is unclear. We evaluated the ability of LLMs to extract rwP events and dates across 7 cancer types and assessed how using LLM-extracted data impacted real-world progression-free survival (rwPFS) estimates compared to using human-abstracted data. Methods: We applied LLM-based extraction techniques to unstructured EHR text for 7 cancer types from the Flatiron Health Research Database: bladder (N=377), breast (N=1000), colorectal (N=564), hepatocellular (N=217), renal cell (N= 229), non–small cell lung (N=1000), and small cell lung (N=955). Prompt engineering strategies including zero-shot, few-shot, and chain-of-thought were tested to optimize performance. We measured agreement between the LLM and abstractor on the presence of rwP (Y/N) and first rwP date (± 30 days) in the first-line (1L) setting. To contextualize the LLM’s ability to extract rwP relative to an expert human abstractor, we evaluated the difference in F1 scores between both curation approaches using a duplicate human-abstracted reference dataset. We also compared rwPFS calculated from LLM-curated data versus human abstractor-curated data for 1000 patients in each cancer type, indexed to 1L start date. Results: Across all cancer types, agreement between the LLM and abstractor on the presence of at least 1 rwP event ranged from 86%-90% while first rwP date agreement ranged from 80%-92%. The difference in F1 score between the LLM and human abstraction was within 3-8 points across cancer types. A comparison of rwPFS between the LLM and human abstractors showed <1 month difference in median rwPFS and overlapping 95% confidence intervals across all cancer types. Discussion: LLMs extracted rwP with high performance, achieving F1 scores similar to expert human abstraction. Across 7 distinct cancers, agreement with human-abstracted data aligned with published inter-abstractor reliability benchmarks, and rwPFS estimates were nearly identical across curation approaches, demonstrating both the generalizability and validity of the approach. These results highlight the potential of LLMs to extract high-quality clinical endpoints at scale, helping to advance research, enhance applications such as predictive algorithms, and ultimately supporting more personalized and effective cancer care. Citation Format: Aaron B. Cohen, Konstantin Krismer, Kelly Magee, James Gippetti, Aaron Dolor, Tori Williams, Erin Fidyk, Hank Kim, Qianyu Yuan, Melissa Estevez. Using large language models for scalable extraction of real-world progression events across multiple cancer types [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B006.

  • Research Article
  • Cite Count Icon 5
  • 10.1371/journal.pdig.0000943
Leveraging large language models for automated depression screening
  • Jul 28, 2025
  • PLOS Digital Health
  • Bazen Gashaw Teferra + 14 more

Mental health diagnoses possess unique challenges that often lead to nuanced difficulties in managing an individual’s well-being and daily functioning. Self-report questionnaires are a common practice in clinical settings to help mitigate the challenges involved in mental health disorder screening. However, these questionnaires rely on an individual’s subjective response which can be influenced by various factors. Despite the advancements of Large Language Models (LLMs), quantifying self-reported experiences with natural language processing has resulted in imperfect accuracy. This project aims to demonstrate the effectiveness of zero-shot learning LLMs for screening and assessing item scales for depression using LLMs. The DAIC-WOZ is a publicly available mental health dataset that contains textual data from clinical interviews and self-report questionnaires with relevant mental health disorder labels. The RISEN prompt engineering framework was utilized to evaluate LLMs’ effectiveness in predicting depression symptoms based on individual PHQ-8 items. Various LLMs, including GPT models, Llama3_8B, Cohere, and Gemini were assessed based on performance. The GPT models, especially GPT-4o, were consistently better than other LLMs (Llama3_8B, Cohere, Gemini) across all eight items of the PHQ-8 scale in accuracy (M = 75.9%), and F1 score (0.74). GPT models were able to predict PHQ-8 items related to emotional and cognitive states. Llama 3_8B demonstrated superior detection of anhedonia-related symptoms and the Cohere LLM’s strength was identifying and predicting psychomotor activity symptoms. This study provides a novel outlook on the potential of LLMs for predicting self-reported questionnaire scores from textual interview data. The promising preliminary performance of the various models indicates there is potential that these models could effectively assist in the screening of depression. Further research is needed to establish a framework for which LLM can be used for specific mental health symptoms and other disorders. As well, analysis of additional datasets while fine-tuning models should be explored.

  • PDF Download Icon
  • Preprint Article
  • 10.2196/preprints.68320
Knowledge Enhancement of Small-Scale Models in Medical Question Answering (Preprint)
  • Nov 3, 2024
  • Xinbai Li + 3 more

BACKGROUND Medical question answering (QA) is essential for various medical applications. While small-scale pre-training language models (PLMs) are widely adopted in open-domain QA tasks through fine-tuning with related datasets, applying this approach in the medical domain requires significant and rigorous integration of external knowledge. Knowledge-enhanced small-scale PLMs have been proposed to incorporate knowledge bases (KBs) to improve performance, as KBs contain vast amounts of factual knowledge. Large language models (LLMs) contain a vast amount of knowledge and have attracted significant research interest due to their outstanding natural language processing (NLP) capabilities. KBs and LLMs can provide external knowledge to enhance small-scale models in medical QA. OBJECTIVE KBs consist of structured factual knowledge that must be converted into sentences to align with the input format of PLMs. However, these converted sentences often lack semantic coherence, potentially causing them to deviate from the intrinsic knowledge of KBs. LLMs, on the other hand, can generate natural, semantically rich sentences, but they may also produce irrelevant or inaccurate statements. Retrieval-augmented generation (RAG) paradigm enhances LLMs by retrieving relevant information from an external database before responding. By integrating LLMs and KBs using the RAG paradigm, it is possible to generate statements that combine the factual knowledge of KBs with the semantic richness of LLMs, thereby enhancing the performance of small-scale models. In this paper, we explore a RAG fine-tuning method, RAG-mQA, that combines KBs and LLMs to improve small-scale models in medical QA. METHODS In the RAG fine-tuning scenario, we adopt medical KBs as an external database to augment the text generation of LLMs, producing statements that integrate medical domain knowledge with semantic knowledge. Specifically, KBs are used to extract medical concepts from the input text, while LLMs are tasked with generating statements based on these extracted concepts. In addition, we introduce two strategies for constructing knowledge: KB-based and LLM-based construction. In the KB-based scenario, we extract medical concepts from the input text using KBs and convert them into sentences by connecting the concepts sequentially. In the LLM-based scenario, we provide the input text to an LLM, which generates relevant statements to answer the question. For downstream QA tasks, the knowledge produced by these three strategies is inserted into the input text to fine-tune a small-scale PLM. F1 and exact match (EM) scores are employed as evaluation metrics for performance comparison. Fine-tuned PLMs without knowledge insertion serve as baselines. Experiments are conducted on two medical QA datasets: emrQA (English) and MedicalQA (Chinese). RESULTS RAG-mQA achieved the best results on both datasets. On the MedicalQA dataset, compared to the KB-based and LLM-based enhancement methods, RAG-mQA improved the F1 score by 0.59% and 2.36%, and the EM score by 2.96% and 11.18%, respectively. On the emrQA dataset, the EM score of RAG-mQA exceeded those of the KB-based and LLM-based methods by 4.65% and 7.01%, respectively. CONCLUSIONS Experimental results demonstrate that RAG fine-tuning method can improve the model performance in medical QA. RAG-mQA achieves greater improvements compared to other knowledge-enhanced methods. CLINICALTRIAL This study does not involve trial registration.

  • Research Article
  • Cite Count Icon 5
  • 10.1111/epi.18475
Extracting epilepsy-related information from unstructured clinic letters using large language models.
  • Jul 10, 2025
  • Epilepsia
  • Shichao Fang + 7 more

The emergence of large language models (LLMs) and the increasing prevalence of electronic health records (EHRs) present significant opportunities for advancing health care research and practice. However, research that compares and applies LLMs to extract key epilepsy-related information from unstructured medical free text is under-explored. This study fills this gap by comparing and applying different open-source LLMs and methods to extract epilepsy information from unstructured clinic letters, thereby optimizing EHRs as a resource for the benefit of epilepsy research. We also highlight some limitations of LLMs. Employing a dataset of 280 annotated clinic letters from King's College Hospital, we explored the efficacy of open-source LLMs (Llama and Mistral series) for extracting key epilepsy-related information, including epilepsy type, seizure type, current anti-seizure medications (ASMs), and associated symptoms. The study used various extraction methods, including direct extraction, summarized extraction, and contextualized extraction, complemented by role-prompting and few-shot prompting techniques. Performance was evaluated against a gold standard dataset, and was also compared to advanced fine-tuned models and human annotations. Llama 2 13b (a 13-billion-parameter LLM developed by Meta) demonstrated superior extraction capabilities across tasks by consistently outperforming other LLMs (F1 = .80 in epilepsy-type extraction, F1 = .76 in seizure-type extraction, and F1 = .90 in current ASMs extraction). Here, F1 score is a balanced metric indicating the model's accuracy in correctly identifying relevant information without excessive false positives. The study highlights the direct extraction showing consistent high performance. Comparative analysis showed that LLMs outperformed current approaches like MedCAT (Medical Concept Annotation Tool) in extracting epilepsy-related information (.2 higher in F1). The results affirm the potential of LLMs in medical information extraction relating to epilepsy, offering insights into leveraging these models for detailed and accurate data extraction from unstructured texts. The study underscores the importance of method selection in optimizing extraction performance and suggests a promising avenue for enhancing medical research and patient care through advanced natural language processing technologies.

  • Research Article
  • Cite Count Icon 1
  • 10.1145/3725411
Logical and Physical Optimizations for SQL Query Execution over Large Language Models
  • Jun 17, 2025
  • Proceedings of the ACM on Management of Data
  • Dario Satriani + 5 more

Interacting with Large Language Models (LLMs) via declarative queries is increasingly popular for tasks like question answering and data extraction, thanks to their ability to process vast unstructured data. However, LLMs often struggle with answering complex factual questions, exhibiting low precision and recall in the returned data. This challenge highlights that executing queries on LLMs remains a largely unexplored domain, where traditional data processing assumptions often fall short. Conventional query optimization, typically cost-driven, overlooks LLM-specific quality challenges such as contextual understanding. Just as new physical operators are designed to address the unique characteristics of LLMs, optimization must consider these quality challenges. Our results highlight that adhering strictly to conventional query optimization principles fails to generate the best plans in terms of result quality. To tackle this challenge, we present a novel approach to enhance SQL results by applying query optimization techniques specifically adapted for LLMs. We introduce a database system, GALOIS, that sits between the query and the LLM, effectively using the latter as a storage layer. We design alternative physical operators tailored for LLM-based query execution and adapt traditional optimization strategies to this novel context. For example, while pushing down operators in the query plan reduces execution cost (fewer calls to the model), it might complicate the call to the LLM and deteriorate result quality. Additionally, these models lack a traditional catalog for optimization, leading us to develop methods to dynamically gather such metadata during query execution. Our solution is compatible with any LLM and balances the trade-off between query result quality and execution cost. Experiments show up to 144% quality improvement over questions in Natural Language and 29% over direct SQL execution, highlighting the advantages of integrating database solutions with LLMs.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant