Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Performance and reliability of state-of-the-art LLMs in complex hand surgery scenarios: A prospective cross-sectional, double-blinded study.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

BackgroundIntegrating large language models (LLMs) into decision-making and education has shown promise across various healthcare disciplines. The study aimed to evaluate the performance of leading LLMs-ChatGPT-5, Gemini 2, Grok 3, and DeepSeek R1-in accurately responding to structured multiple-choice and open-ended queries about complex case scenarios in hand surgery.MethodsA prospective cross-sectional analysis used 50 clinically relevant, guideline-based case scenarios developed for hand surgery. Each scenario consisted of four open-ended and two multiple-choice questions, totaling 300 points per LLM. Responses were independently assessed by blinded expert reviewers using a standardized six-point Likert scale evaluating accuracy, completeness, and adherence to international surgical guidelines.ResultsIn multiple-choice queries, Gemini (5.9 ± 0.2) and Grok (5.9 ± 0.1) outperformed ChatGPT (5.7 ± 0.3; p = 0.031 and p = 0.009, respectively) and DeepSeek (5.6 ± 0.4; p = 0.004 and p = 0.001, respectively). In open-ended queries, Gemini (5.6 ± 0.3 accuracy) and Grok (5.5 ± 0.4 accuracy) demonstrated superior results across all measured dimensions-accuracy, completeness, and guideline adherence-markedly surpassing ChatGPT (5.1 ± 0.5 accuracy, p < 0.001) and DeepSeek (4.9 ± 0.6 accuracy; p < 0.001). Notably, Gemini and Grok demonstrated consistently high performance with minimal variability, while ChatGPT, particularly DeepSeek, exhibited considerable inconsistency in complex clinical judgments.ConclusionGemini 2 and Grok 3 showed reliable and clinically relevant performance, positioning them as promising adjunctive tools for decision-making and education in hand surgery. The limitations in ChatGPT-5 and the significant shortcomings of DeepSeek underscore the necessity for cautious deployment and continued refinement.

Similar Papers
  • Research Article
  • Cite Count Icon 6
  • 10.1002/bcp.70137
Evaluating and leveraging large language models in clinical pharmacology and therapeutics assessment: From exam takers to exam shapers.
  • Jun 10, 2025
  • British journal of clinical pharmacology
  • Alexandre O Gérard + 11 more

In medical education, the ability of large language models (LLMs) to match human performance raises questions about their potential as educational tools. This study evaluates LLMs' performance on Clinical Pharmacology and Therapeutics (CPT) exams, comparing their results to medical students and exploring their ability to identify poorly formulated multiple-choice questions (MCQs). ChatGPT-4 Omni, Gemini Advanced, Le Chat and DeepSeek R1 were tested on local CPT exams (third year of bachelor's degree, first/second year of master's degree) and the European Prescribing Exam (EuroPE+). The exams included MCQs and open-ended questions assessing knowledge and prescribing skills. LLM results were analysed using the same scoring system as students. A confusion matrix was used to evaluate the ability of ChatGPT and Gemini to identify ambiguous/erroneous MCQs. LLMs achieved comparable or superior results to medical students across all levels. For local exams, LLMs outperformed M1 students and matched L3 and M2 students. In EuroPE+, LLMs significantly outperformed students in both the knowledge and prescribing skills sections. All LLM errors in EuroPE+ were genuine (100%), whereas local exam errors were frequently due to ambiguities or correction flaws (24.3%). When both ChatGPT and Gemini provided the same incorrect answer to an MCQ, the specificity for detecting ambiguous questions was 92.9%, with a negative predictive value of 85.5%. LLMs demonstrate capabilities comparable to or exceeding medical students in CPT exams. Their ability to flag potentially flawed MCQs highlights their value not only as educational tools but also as quality control instruments in exam preparation.

  • Research Article
  • 10.1093/sleep/zsaf090.1401
1401 Bringing Medicine Expertise to Your Screen: A New Frontier in Curbside Sleep Consultation Leveraging Large Language Models?
  • May 19, 2025
  • SLEEP
  • Nina Kuei + 4 more

Introduction Advances in large language models (LLMs) have opened new avenues for healthcare applications. Recently, ChatGPT-4 successfully achieved the pass mark &amp;gt;80% in 5 of 10 sleep medicine examination domains, indicating a strong foundational knowledge of sleep medicine. However, the ability to answer USMLE-type multiple-choice questions may not equate to the capacity to offer accurate and comprehensive answers to clinical queries or real-world case scenarios. Current literature exhibits a significant gap in validation research examining the clinical utility of LLMs in evidence-based medicine practice. This investigation aims to evaluate the potential usefulness and reliability of LLMs as adjunctive clinical decision-support tools (namely, “curbside consultants”) in sleep medicine. Methods Six clinical sleep queries and six case scenarios were presented to 6 LLMs including 3 general-purpose LLMs (ChatGPT-4o, Gemini-1.5-Pro, and Llama-3.1-405B) and 3 medical-specialized LLMs (OpenEvidence, Clara AI, and MediGPT). Performance assessment was conducted independently utilizing two independently developed 5-point Likert scales evaluating two primary domains: answer content (accuracy, relevance, comprehensiveness/depth, clarity/coherence, and unique insightfulness) and reference quality (accuracy, relevance, currentness, comprehensiveness/depth, and searchability). Benchmark answers were established through consensus among four sleep medicine specialists. Analysis of variance was used to compare the performance of the LLMs. Results The medical LLMs demonstrated superior overall performance compared to the general LLMs (P&amp;lt; 0.001). The primary distinction was observed in reference quality metrics, where medical LLMs significantly outperformed general LLMs across all parameters: accuracy, relevance, and searchability (p&amp;lt; 0.001), currentness (p=0.001), and comprehensiveness/depth (p=0.010). Notably, Open Evidence achieved the highest reference quality (p&amp;lt; 0.001). In contrast, the answer content analysis revealed no significant overall differences between medical and general LLMs (p=0.659). Most answer content-related metrics, including accuracy, relevance, or unique insightfulness, did not differ significantly (p&amp;gt;0.050). An exception was clarify/coherence, where medical LLMs were superior to general LLMs (p=0.030). Furthermore, MediGPT and ChatGPT-4o displayed better content comprehensiveness/depth relative to other LLMs (p&amp;lt; 0.001). These two LLMs exhibited comparable performance in overall metrics (p=0.615), answer contents (p=0.922), and reference quality (p=0.621). Conclusion Both medical-specialized and general-purpose LLMs show promise as adjunctive decision-support tools in clinical practice. However, substantial improvements in reference quality are critically needed across most LLM platforms. Support (if any)

  • Research Article
  • 10.1016/j.chest.2025.11.034
Evaluating the Accuracy of Large Language Models in Answering Asthma Multiple Choice and Objective Structured Clinical Examination Questions.
  • Dec 1, 2025
  • Chest
  • Pei Ye Li + 2 more

Evaluating the Accuracy of Large Language Models in Answering Asthma Multiple Choice and Objective Structured Clinical Examination Questions.

  • Research Article
  • Cite Count Icon 17
  • 10.7759/cureus.81871
Evaluating the Accuracy and Reliability of Large Language Models (ChatGPT, Claude, DeepSeek, Gemini, Grok, and Le Chat) in Answering Item-Analyzed Multiple-Choice Questions on Blood Physiology.
  • Apr 8, 2025
  • Cureus
  • Mayank Agarwal + 2 more

Background Previous research has highlighted the potential of large language models (LLMs) in answering multiple-choice questions (MCQs) in medical physiology. However, their accuracy and reliability in specialized fields, such as blood physiology, remain underexplored. This study evaluates the performance of six free-to-use LLMs (ChatGPT, Claude, DeepSeek, Gemini, Grok, and Le Chat) in solving item-analyzed MCQs on blood physiology. The findings aim to assess their suitability as educational aids. Methods This cross-sectional study at the All India Institute of Medical Sciences, Raebareli, India, involved administering a 40-item MCQ test on blood physiology to 75 first-year medical students. Item analysis utilized the Difficulty Index (DIF I), Discrimination Index (DI), and Distractor Effectiveness (DE). Internal consistency was assessed with the Kuder-Richardson 20 (KR-20) coefficient. These 40 item-analyzed MCQs were presented to six selected LLMs (ChatGPT, Claude, DeepSeek, Gemini, Grok, Le Chat) available as standalone Android applications on March 19, 2025. Three independent users accessed each LLM simultaneously, uploading the compiled MCQs in a Portable Document Format (PDF) file. Accuracy was determined as the percentage of correct responses averaged across all three users. Reliability was measured as the percentage of MCQs consistently answered correctly by LLM to all three users. Descriptive statistics were presented as mean ± standard deviation and percentages. Pearson's correlation coefficient or Spearman's rho was used to evaluate the associations between variables, with p < 0.05 considered significant. Results Item analysis confirmed the validity and reliability of the assessment tool, with a DIF I of 63.2 ± 20.4, a DI of 0.38 ± 0.20, a DE of 66.7 ± 33.3, and a KR-20 of 0.804. Among LLMs, Claude 3.7demonstrated the highest reliable accuracy (95%), followed by DeepSeek (93%), Grok 3 beta (93%), ChatGPT (90%), Gemini 2.0 (88%), and Mistral Le Chat (70%). No significant correlations were found between LLM performance and MCQ difficulty, discrimination power, or distractor effectiveness. Conclusions The MCQ assessment tool exhibited an appropriate difficulty level, strong discriminatory power, and adequately constructed distractors. LLMs, particularly Claude, DeepSeek, and Grok, demonstrated high accuracy and reliability in solving blood physiology MCQs, supporting their role as supplementary educational tools. LLMs handled questions of varying difficulty, discrimination power, and distractor effectiveness with similar competence. However, given occasional errors, they should be used alongside traditional teaching methods and expert supervision.

  • Research Article
  • Cite Count Icon 3
  • 10.1080/10872981.2025.2592430
When AI models take the exam: large language models vs medical students on multiple-choice course exams
  • Nov 29, 2025
  • Medical Education Online
  • Pablo Ros-Arlanzón + 9 more

Large language models (LLMs) are increasingly used in healthcare and medical education, but their performance on institution-authored multiple-choice questions (MCQs), particularly with negative marking, remains unclear. To compare the examination performance of five contemporary LLMs with enrolled medical students on final multiple-choice (MCQ-style) course exams across four clinical courses. We conducted a comparative cross-sectional study at Miguel Hernández University (Spain) in 2025. Final exams in Infectious Diseases, Neurology, Respiratory Medicine, and Cardiovascular Medicine were administered under routine conditions in Spanish. Five LLMs (OpenAI o1, GPT-4o, DeepSeek R1, Microsoft Copilot, and Google Gemini 1.5 Flash) completed all MCQs in two independent runs. Scores were averaged and test–retest was estimated with Gwet’s AC1. Student scores (n = 442) were summarized as mean ± SD or median (IQR). Pairwise differences between models were explored with McNemar’s test; student–LLM contrasts were descriptive. Across courses, LLMs consistently exceeded the student median and, in several instances, the highest student score. Mean LLM courses scores ranged 7.46–9.88, versus student means 4.28–7.32. OpenAI o1 achieved the highest mean in three courses; Copilot led in Cardiovascular Medicine (text-only subset due to image limitations). All LLMs answered every MCQ and short term test–retest agreement was high (AC1 0.79–1.00). Aggregated across courses, LLMs averaged 8.75 compared with 5.76 for students. On department-set Spanish MCQ exams with negative marking, LLMs outperformed enrolled medical students, answered every item, and showed high short-term reproducibility. These findings support cautious, faculty-supervised use of LLMs as adjuncts to MCQ assessment (e.g. automated pretesting, feedback). Confirmation across institutions, languages, and image-rich formats, and evaluation of educational impact beyond accuracy are needed.

  • Abstract
  • Cite Count Icon 3
  • 10.1182/blood-2023-185854
Evaluating the Performance of Large Language Models in Hematopoietic Stem Cell Transplantation Decision Making
  • Nov 2, 2023
  • Blood
  • Ivan Civettini + 14 more

Evaluating the Performance of Large Language Models in Hematopoietic Stem Cell Transplantation Decision Making

  • Research Article
  • 10.1016/j.jdent.2026.106724
Temporal trends in large language model (LLM) accuracy: A meta-analysis of multiple-choice question performance in dentistry and dental education.
  • Aug 1, 2026
  • Journal of dentistry
  • Alison F Doubleday + 26 more

Temporal trends in large language model (LLM) accuracy: A meta-analysis of multiple-choice question performance in dentistry and dental education.

  • Research Article
  • 10.1177/15589447251371089
Do ChatGPT and Gemini's Recommendations Align With Established Guidelines for Hand and Upper Extremity Surgery?
  • Sep 18, 2025
  • Hand (New York, N.Y.)
  • Yibin B Zhang + 5 more

The use of large language models (LLMs) such as ChatGPT and Gemini in clinical settings has surged, presenting potential benefits in reducing administrative workload and enhancing patient communication. However, concerns about the clinical accuracy of these tools persist. This study evaluated the concordance of ChatGPT and Gemini's recommendations with American Academy of Orthopedic Surgeons (AAOS) clinical practice guidelines (CPGs) for carpal tunnel syndrome, distal radius fractures, and glenohumeral joint osteoarthritis. ChatGPT (version 4o) and Gemini (version 1.5 Flash) were queried using structured text-based prompts aligned with AAOS CPGs. The LLMs' outputs were analyzed by blinded reviewers to determine concordance with the guidelines. Concordance rates were compared across models, topics, and guideline strength using descriptive statistics and McNemar's test. The transparency of responses, including source citation, was also assessed. A total of 174 recommendations were generated, with an overall concordance rate of 62.1%. When comparing concordance rates between LLMs, there was no statistically significant difference between ChatGPT and Gemini (66.7% vs 57.5%, P = .131). Concordance varied by topic and guideline strength, with ChatGPT performing best for moderately supported guidelines. Both models demonstrated low citation transparency. Gemini provided sources for 39.1% of recommendations, significantly more than ChatGPT's 3.5% (P < .0001). Despite modest concordance rates, both models exhibited significant limitations, including variability across topics and guideline strengths, as well as insufficient citation transparency. These findings highlight the challenges in integrating LLMs into clinical practice and emphasize the need for further refinement and evaluation before adoption in hand surgery.

  • Research Article
  • 10.1016/j.prosdent.2026.03.043
Evaluating the performance of large language models versus prosthodontic residents on the 2024 and 2025 National Prosthodontic Resident Examination.
  • Apr 15, 2026
  • The Journal of prosthetic dentistry
  • Soni Prasad + 5 more

Evaluating the performance of large language models versus prosthodontic residents on the 2024 and 2025 National Prosthodontic Resident Examination.

  • Research Article
  • Cite Count Icon 1
  • 10.1080/0142159x.2025.2497891
The answer may vary: large language model response patterns challenge their use in test item analysis
  • May 2, 2025
  • Medical Teacher
  • Lauren K Buhl

Introduction The validation of multiple-choice question (MCQ)-based assessments typically requires administration to a test population, which is resource-intensive and practically demanding. Large language models (LLMs) are a promising tool to aid in many aspects of assessment development, including the challenge of determining the psychometric properties of test items. This study investigated whether LLMs could predict the difficulty and point biserial indices of MCQs, potentially alleviating the need for preliminary analysis in a test population. Methods Sixty MCQs developed by subject matter experts in anesthesiology were presented one hundred times each to five different LLMs (ChatGPT-4o, o1-preview, Claude 3.5 Sonnet, Grok-2, and Llama 3.2) and to clinical fellows. Response patterns were analyzed, and difficulty indices (proportion of correct responses) and point biserial indices (item-test score correlation) were calculated. Spearman correlation coefficients were used to compare difficulty and point biserial indices between the LLMs and fellows. Results Marked differences in response patterns were observed among LLMs: ChatGPT-4o, o1-preview, and Grok-2 showed variable responses across trials, while Claude 3.5 Sonnet and Llama 3.2 gave consistent responses. The LLMs outperformed fellows with mean scores of 58% to 85% compared to 57% for the fellows. Three LLMs showed a weak correlation with fellow difficulty indices (r = 0.28–0.29), while the two highest scoring models showed no correlation. No LLM predicted the point biserial indices. Discussion These findings suggest LLMs have limited utility in predicting MCQ performance metrics. Notably, higher-scoring models showed less correlation with human performance, suggesting that as models become more powerful, their ability to predict human performance may decrease. Understanding the consistency of an LLM’s response pattern is critical for both research methodology and practical applications in test development. Future work should focus on leveraging the language-processing capabilities of LLMs for overall assessment optimization (e.g., inter-item correlation) rather than predicting item characteristics.

  • Research Article
  • 10.1007/s11695-025-08418-y
Performance of Large Language Models in Metabolic Bariatric Surgery: a Comparative Study.
  • Dec 11, 2025
  • Obesity surgery
  • Hassan El-Masry + 10 more

The rapid integration of Large Language Models (LLMs) into healthcare necessitates a rigorous evaluation of their performance in specialized medical fields. In metabolic bariatric surgery (MBS), LLMs have the potential to revolutionize education and clinical support, yet their accuracy and reliability are not well-established. This study provides a critical assessment of the capabilities of current LLMs in the context of MBS. This cross-sectional validation study assessed the performance of six LLMs (ChatGPT-3.5, ChatGPT-4o, Gemini, Copilot, GROK, and DeepSeek) in answering 100 evidence-based binary and multiple-choice questions related to MBS. Questions were constructed from international guidelines and categorized into six thematic domains. Expert consensus answers served as the reference standard, with inter-rater reliability measured using Fleiss’ κ. Model outputs were scored for accuracy. Comparisons across LLMs were first assessed using an overall test for differences between multiple related groups. Pairwise comparisons were then conducted between LLMs to identify specific differences in performance. Across the dataset, the mean number of correct LLM responses per question was 3.9 (SD = 1.8). ChatGPT-4o achieved the highest accuracy (66.0%), while DeepSeek recorded the lowest (60.0%). Accuracy varied across domains, highest for indications/contraindications (78.7%) and complications/management (68.0%), and lowest for preoperative preparation (52.0%) and postoperative care (58.4%). Binary questions yielded higher accuracy (69.1%) than multiple-choice questions (62.0%). Inter-expert reliability was substantial (κ = 0.742, 95% CI: 0.71–0.77). Agreement between LLMs and experts ranged from fair (DeepSeek κ = 0.349) to moderate (ChatGPT-4o κ = 0.446). No significant accuracy differences were detected across models (Friedman test, p = 0.662). LLMs represent a promising, yet imperfect, adjunct in MBS education. Their utility is currently limited by inconsistencies in accuracy, particularly in areas requiring nuanced clinical judgment. While these models can supplement traditional learning resources, they are not yet a substitute for expert clinical guidance. This study underscores the need for continued refinement and validation of LLMs to ensure their safe and effective integration into clinical practice. LLMs show moderate accuracy in bariatric surgery education, strongest in guideline-based domains. Newer models (ChatGPT-4o, Gemini, Copilot) performed slightly better, but gains were modest. Accuracy was higher for binary than multiple-choice questions.

  • Research Article
  • 10.3390/dj14020072
Assessing the Efficacy of Artificial Intelligence Platforms in Answering Dental Caries Multiple-Choice Questions: A Comparative Study of ChatGPT and Google Gemini Language Models.
  • Jan 27, 2026
  • Dentistry journal
  • Amr Ahmed Azhari + 5 more

Objective: This study aimed to compare the accuracy of two large language models (LLMs)-ChatGPT (version 3.5) and Google Gemini (formerly Bard)-in answering dental caries-related multiple-choice questions (MCQs) using a simulated student examination framework across seven examination lengths. Materials and Methods: A total of 125 validated dental caries MCQs were extracted from Dental Decks and Oxford University Press question banks. Seven examination groups were constructed with varying question counts (25, 35, 45, 55, 65, 75, and 85 questions). For each group, 100 simulations were generated per LLM (ChatGPT and Gemini), resulting in 1400 simulated examinations. Each simulated student received a unique randomized subset of questions. MCQs were answered by each LLM using a standardized prompt to minimize ambiguity. Outcomes included mean score, passing rate (≥60%), and performance differences between LLMs. Statistical analyses included independent t-tests, one-way ANOVA within each LLM, and two-way ANOVA examining interactions between LLM type and question count. Results: Across all seven examination formats, Gemini significantly outperformed ChatGPT (p < 0.001). Gemini achieved higher passing rates and higher mean scores in every examination length. One-way ANOVA revealed significant score variation with increasing exam length for both LLMs (p < 0.05). Two-way ANOVA demonstrated significant main effects of LLM type and question count, with no significant interaction. Randomization had no measurable effect on Gemini performance but influenced ChatGPT scores. Conclusions: Gemini demonstrated superior accuracy and higher passing rates compared to ChatGPT in all simulated examination formats. While both LLMs struggled with complex caries-related content, Gemini provided more reliable performance across question quantities. Educators should exercise caution in relying on LLMs for automated assessment or self-study, and future research should evaluate human-AI hybrid models and LLM performance across broader dental domains.

  • Research Article
  • 10.1161/circ.152.suppl_3.4369198
Abstract 4369198: Performance of Large Language Models in Analyzing Common Hypertension Scenarios in Clinical Practice
  • Nov 4, 2025
  • Circulation
  • Jaleh Zand + 7 more

Background/Objective: Hypertension is the most prevalent chronic disease in primary care and a leading cause of cardiovascular morbidity and mortality. Despite existing guidelines, therapeutic inertia and suboptimal control persist. Large language models (LLMs) offer a potential valuable addition to augment clinical decision-making, yet their reliability for guideline-driven tasks remains unverified. This study evaluated the accuracy and safety of hypertension management recommendations generated by three LLMs compared to expert responses. Methods: Fifty-one clinical vignettes representing 17 core hypertension management concepts were constructed by hypertension experts. Each case was submitted to three LLMs (GPT-4, Gemini, MedLM) and a hypertension expert also wrote the “gold standard” answers. Three blinded expert reviewers rated each response on a 4-point accuracy scale, a binary safety (safe/unsafe) scale, and attempted to identify the source (LLM vs. expert) providing the response. Ratings were analyzed using mean scores, percentages of accurate and safe responses, and inter-rater agreement. Results: GPT-4 had the highest accuracy (83%) and safety (86%) scores among LLMs but remained inferior to expert responses (92% accuracy, 93% safety). Gemini and MedLM performed significantly worse (accuracy: 64% and 35%; safety: 73% and 39%, respectively). GPT-4 generated the most guideline-concordant responses (46%) among the three LLMs (Gemini 35%, MedLM 14%), but remains lower than experts’ responses (68%). Evaluators misidentified LLM responses as expert-written in 10 to 25% of cases, particularly with GPT-4. Inter-rater reliability for accuracy ratings was highest for expert-generated responses (ICC 0.81), with progressively lower agreement for GPT-4 (0.76), Gemini (0.70), and MedLM (0.68). A similar pattern was observed for safety and source discrimination ratings. The agreement was strongest for safety assessments and weakest for source discrimination. Conclusion: Among three tested LLMs, GPT-4 demonstrated closer agreement to expert decisions thereby showing greater potential for supporting hypertension management. However, current LLMs’ versions frequently produce inaccurate or unsafe recommendations and remain inferior to expert judgment. Human-in-the-loop supervision remains essential when deploying LLMs for clinical decision-making.

  • Research Article
  • 10.2196/82842
Performance of Large Language Models in the Japanese Public Health Nurse National Examination: Comparative Cross-Sectional Study.
  • Feb 20, 2026
  • JMIR nursing
  • Yutaro Takahashi + 3 more

Large language models (LLMs) have shown promising results on Japanese national medical and nursing examinations. However, no study has evaluated LLM performance on the Japanese Public Health Nurse National Examination, which requires specialized knowledge in community health and public health nursing practice. This study aimed to compare the performance of multiple LLMs on the Japanese Public Health Nurse National Examination and evaluate their potential utility in public health nursing education. Three LLMs were evaluated: GPT-4o, Claude Opus 4, and Gemini 2.5 Pro. All 110 questions from the 111th Public Health Nurse National Examination were administered using standardized prompts. Questions were classified by format (text vs figure or calculation), content (general vs situational), and selection type (single vs multiple choice). Accuracy rates and 95% CIs were calculated, with statistical comparisons performed using chi-square tests. All LLMs exceeded the passing criterion (60%). The accuracy rates were as follows: 85.5% (94/110) for GPT-4o (95% CI 77.5%-91.5%), 91.8% (101/110) for Claude Opus 4 (95% CI 85.0%-96.2%), and 92.7% (102/110) for Gemini 2.5 Pro (95% CI 86.2%-96.8%). No significant differences were found among the LLMs (P>.99). However, all models showed lower accuracy on multiple-choice questions than on single-choice questions, with significant intramodel differences observed for GPT-4o (10/16, 62.5% vs 82/92, 89.1%; P=.01) and Claude Opus 4 (12/16, 75% vs 87/92, 94.6%; P=.03). LLMs demonstrated high performance on a public health nursing examination but showed limitations in complex reasoning requiring multiple-choice selection. These findings suggest the potential for LLM use as educational support tools while highlighting the need for cautious implementation in specialized nursing education.

  • Research Article
  • Cite Count Icon 13
  • 10.2196/67677
Improving Dietary Supplement Information Retrieval: Development of a Retrieval-Augmented Generation System With Large Language Models.
  • Mar 19, 2025
  • Journal of medical Internet research
  • Yu Hou + 3 more

Dietary supplements (DSs) are widely used to improve health and nutrition, but challenges related to misinformation, safety, and efficacy persist due to less stringent regulations compared with pharmaceuticals. Accurate and reliable DS information is critical for both consumers and health care providers to make informed decisions. This study aimed to enhance DS-related question answering by integrating an advanced retrieval-augmented generation (RAG) system with the integrated Dietary Supplement Knowledgebase 2.0 (iDISK2.0), a dietary supplement knowledge base, to improve accuracy and reliability. We developed iDISK2.0 by integrating updated data from authoritative sources, including the Natural Medicines Comprehensive Database, the Memorial Sloan Kettering Cancer Center database, Dietary Supplement Label Database, and Licensed Natural Health Products Database, and applied advanced data cleaning and standardization techniques to reduce noise. The RAG system combined the retrieval power of a biomedical knowledge graph with the generative capabilities of large language models (LLMs) to address limitations of stand-alone LLMs, such as hallucination. The system retrieves contextually relevant subgraphs from iDISK2.0 based on user queries, enabling accurate and evidence-based responses through a user-friendly interface. We evaluated the system using true-or-false and multiple-choice questions derived from the Memorial Sloan Kettering Cancer Center database and compared its performance with stand-alone LLMs. iDISK2.0 integrates 174,317 entities across 7 categories, including 8091 dietary supplement ingredients; 163,806 dietary supplement products; 786 diseases; and 625 drugs, along with 6 types of relationships. The RAG system achieved an accuracy of 99% (990/1000) for true-or-false questions on DS effectiveness and 95% (948/100) for multiple-choice questions on DS-drug interactions, substantially outperforming stand-alone LLMs like GPT-4o (OpenAI), which scored 62% (618/1000) and 52% (517/1000) on these respective tasks. The user interface enabled efficient interaction, supporting free-form text input and providing accurate responses. Integration strategies minimized data noise, ensuring access to up-to-date, DS-related information. By integrating a robust knowledge graph with RAG and LLM technologies, iDISK2.0 addresses the critical limitations of stand-alone LLMs in DS information retrieval. This study highlights the importance of combining structured data with advanced artificial intelligence methods to improve accuracy and reduce misinformation in health care applications. Future work includes extending the framework to broader biomedical domains and improving evaluation with real-world, open-ended queries.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant