Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Human-AI collaboration in clinical reasoning: a UK replication and interaction analysis.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

A paper from Goh etal. found that a large language model (LLM) working alone outperformed American clinicians assisted by the same LLM in diagnostic reasoning tests (Goh E, Gallo R, Hom J, Strong E, Weng J, Kerman H, etal. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Netw Open 2024;7:e2440969). We aimed to replicate this experiment in a UK setting and explore how interactions with the LLM might explain the observed gaps in performance. This was a within-subjects study of UK physicians. 22 participants answered structured questions on four clinical vignettes. For 2 cases physicians had access to an LLM via a custom-built web-application. Results were analysed using a mixed-effects model accounting for case difficulty and the variability of clinicians at baseline. Qualitative analysis involved coding of participant-LLM interaction logs and evaluating the rates of LLM use per question. Physicians with LLM assistance scored significantly lower than the LLM alone (mean difference 21.3 percentage points, p<0.001). Access to the LLM was associated with improved physician performance compared to using conventional resources (74.3 vs. 65.7 %, p=0.001). There was significant heterogeneity in the degree of LLM-assisted improvement (SD 12.8 %). Qualitative analysis revealed that only 30 % of case questions were directly posed to the LLM, which suggests that under-utilisation of the LLM contributed to the observed performancegap. While access to an LLM can improve diagnostic accuracy, realising the full potential of human-AI collaboration may require a focus on training clinicians to integrate these tools into their cognitive workflows and on designing systems that make these integrations the default rather than an optionalextra.

Similar Papers
  • Research Article
  • Cite Count Icon 504
  • 10.1001/jamanetworkopen.2024.40969
Large Language Model Influence on Diagnostic Reasoning
  • Oct 28, 2024
  • JAMA Network Open
  • Ethan Goh + 15 more

Large language models (LLMs) have shown promise in their performance on both multiple-choice and open-ended medical reasoning examinations, but it remains unknown whether the use of such tools improves physician diagnostic reasoning. To assess the effect of an LLM on physicians' diagnostic reasoning compared with conventional resources. A single-blind randomized clinical trial was conducted from November 29 to December 29, 2023. Using remote video conferencing and in-person participation across multiple academic medical institutions, physicians with training in family medicine, internal medicine, or emergency medicine were recruited. Participants were randomized to either access the LLM in addition to conventional diagnostic resources or conventional resources only, stratified by career stage. Participants were allocated 60 minutes to review up to 6 clinical vignettes. The primary outcome was performance on a standardized rubric of diagnostic performance based on differential diagnosis accuracy, appropriateness of supporting and opposing factors, and next diagnostic evaluation steps, validated and graded via blinded expert consensus. Secondary outcomes included time spent per case (in seconds) and final diagnosis accuracy. All analyses followed the intention-to-treat principle. A secondary exploratory analysis evaluated the standalone performance of the LLM by comparing the primary outcomes between the LLM alone group and the conventional resource group. Fifty physicians (26 attendings, 24 residents; median years in practice, 3 [IQR, 2-8]) participated virtually as well as at 1 in-person site. The median diagnostic reasoning score per case was 76% (IQR, 66%-87%) for the LLM group and 74% (IQR, 63%-84%) for the conventional resources-only group, with an adjusted difference of 2 percentage points (95% CI, -4 to 8 percentage points; P = .60). The median time spent per case for the LLM group was 519 (IQR, 371-668) seconds, compared with 565 (IQR, 456-788) seconds for the conventional resources group, with a time difference of -82 (95% CI, -195 to 31; P = .20) seconds. The LLM alone scored 16 percentage points (95% CI, 2-30 percentage points; P = .03) higher than the conventional resources group. In this trial, the availability of an LLM to physicians as a diagnostic aid did not significantly improve clinical reasoning compared with conventional resources. The LLM alone demonstrated higher performance than both physician groups, indicating the need for technology and workforce development to realize the potential of physician-artificial intelligence collaboration in clinical practice. ClinicalTrials.gov Identifier: NCT06157944.

  • Research Article
  • Cite Count Icon 5
  • 10.1038/s44360-025-00007-8
Large language model diagnostic assistance for physicians in a lower-middle-income country: a randomized controlled trial
  • Feb 6, 2026
  • Nature Health
  • Ihsan Ayyub Qazi + 5 more

Diagnostic errors remain a major source of preventable patient harm, particularly in low- and middle-income countries. Large language models (LLMs) have the potential to bridge these gaps in diagnostic capacity, but can generate inaccurate information, necessitating artificial intelligence (AI)-literacy training. However, whether AI-trained physicians can leverage LLMs to improve diagnostic reasoning compared with conventional resources is unknown. Here we conducted a single-blind randomized controlled trial involving 60 licensed physicians from multiple medical institutions in Pakistan between January and May 2025. Participants completed a 20-hour AI-literacy curriculum covering LLM capabilities, appropriate use and limitations. Post-training, physicians were randomized to either LLM access plus conventional resources or conventional resources only, with 75 minutes to review up to 6 clinical vignettes. The primary outcome was diagnostic reasoning score using an expert-validated grading rubric; the secondary outcome was time per vignette. Of 58 physicians completing the study, those with LLM access achieved a mean diagnostic reasoning score of 71.4% versus 42.6% with conventional resources alone, an adjusted difference of 27.5 percentage points (95% confidence interval (CI), 22.8 to 32.2; P < 0.001). Time per case was similar (mean difference −6.4 s; 95% CI, −68.2 to 55.3; P = 0.84). While a secondary exploratory analysis showed that LLM alone outperformed LLM-assisted physicians (11.5 percentage points; 95% CI, 5.5 to 17.5; P < 0.001), in 31.4% of cases, the physician-plus-LLM group exceeded the median LLM-alone performance, indicating complementarities. Among AI-trained physicians, access to an LLM substantially improved diagnostic reasoning without slowing case review, suggesting that effective LLM utilization could help address diagnostic gaps in resource-limited settings. ClinicalTrials.gov registration: NCT06774612 . In a randomized controlled study involving 58 physicians in Pakistan, assistance by a large language model in diagnostic reasoning resulted in a 27.5% increase in performance on 6 clinical vignettes.

  • Research Article
  • 10.64898/2026.02.03.26345402
Evaluating Diagnostic Accuracy and Clinical Reasoning of Multiple Large Language Models in Psychiatry.
  • Feb 11, 2026
  • medRxiv : the preprint server for health sciences
  • Kevin W Jin + 20 more

Existing large language model (LLM) evaluations rely on accuracy benchmarks that fail to capture whether models reason well while making diagnoses. Studies that do analyse reasoning focus on post hoc explanations accompanying model outputs rather than distinct, clinician-visible artifacts such as detailed reasoning traces. This creates a translational gap in domains such as psychiatry, where diagnosis relies on narrative interpretation, diagnostic reasoning, and clinical judgment under uncertainty. We conducted a mixed-methods evaluation of four state-of-the-art LLMs using a clinician-curated dataset of 196 psychiatric case vignettes, including 135 published cases and 61 novel clinician-authored vignettes. Diagnostic accuracy was assessed using multiple metrics (top-1 accuracy, top-5 accuracy, recall@5, and mean reciprocal rank) based on ranked lists of five differential diagnoses per vignette. Clinical reasoning quality was evaluated on a randomly selected subset of 30 vignettes through clinician assessment of model-generated diagnostic reasoning traces alongside qualitative commentary from board-certified psychiatrists. We examined the association between clinician-rated reasoning quality and diagnostic correctness and included an illustrative comparison with psychiatry residents. Clinician-rated diagnostic reasoning quality was strongly associated with diagnostic correctness in mixed-effects logistic regression analyses (β = 1·80; p < 0·001), whereas data extraction quality alone was not. Across the full vignette set, models demonstrated moderate to high diagnostic accuracy. The highest-performing model achieved a top-5 accuracy of 0·801 and also received the highest clinician-rated reasoning scores. In an illustrative comparison, model diagnostic accuracy fell within the range observed for psychiatry residents. Diagnostic reasoning quality captures clinically meaningful variation in LLM performance beyond accuracy metrics. Psychiatry may represent a stringent testbed for evaluating reasoning in narrative-driven clinical domains. Evaluation frameworks for LLM-based clinical decision support should incorporate structured assessment of reasoning processes, not accuracy alone. Evidence before this study: We searched PubMed and Scopus for studies evaluating large language models for psychiatric diagnosis and/or differential diagnosis from text vignettes. Searches were run from database inception to February 6, 2026, using terms including ("large language model" OR LLM "artificial intelligence" OR "generative AI" OR AI OR ChatGPT OR GPT OR Claude OR Gemini OR DeepSeek OR Llama) AND (psychiatr* OR mental OR DSM OR "differential diagnosis" OR diagnos*) AND (vignette OR case OR "case report"). We included empirical studies that evaluated model diagnostic outputs using psychiatric cases/vignettes and excluded editorials/commentaries and studies that did not report case-level diagnostic performance. We did not formally assess risk of bias/quality; studies were heterogeneous in vignette sources, model access, and outcome definitions, precluding quantitative pooling. Prior studies have typically examined small vignette sets, focused on narrow diagnostic domains, evaluated single models, or relied primarily on outcome-based accuracy metrics. When diagnostic reasoning has been assessed, it has usually been inferred from post hoc explanations accompanying model outputs rather than evaluated as a distinct, clinician-visible artifact. Clinician-grounded evaluations of diagnostic reasoning across multiple contemporary models remain limited.Added value of this study: This study provides a large-scale, clinician-grounded evaluation of diagnostic accuracy and diagnostic reasoning quality across four contemporary large language models using a diverse dataset of psychiatric case vignettes. Rather than relying solely on outcome-based explanations, we directly evaluated model-generated diagnostic reasoning traces as clinician-visible artifacts using structured clinician ratings and qualitative analysis. By integrating multiple accuracy metrics with clinician assessment of reasoning coherence, flexibility, and plausibility, we demonstrate that clinician-rated reasoning quality is strongly associated with diagnostic correctness, whereas data extraction quality alone is not. Our analysis also identifies recurrent reasoning failure modes not captured by accuracy metrics, highlighting psychiatry as a stringent testbed for evaluating reasoning in narrative-driven clinical domains.Implications of all the available evidence: Evaluations of large language models for clinical decision support should extend beyond accuracy to include systematic assessment of clinician-visible diagnostic reasoning. Mixed-methods, clinician-grounded evaluation frameworks that examine both diagnostic outcomes and reasoning artifacts may be critical for responsible assessment of LLMs in psychiatry and other areas of medicine where diagnosis depends on interpretation, judgment, and tolerance of uncertainty.

  • Research Article
  • Cite Count Icon 122
  • 10.1038/s41591-024-03456-y
GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial.
  • Feb 5, 2025
  • Nature medicine
  • Ethan Goh + 17 more

While large language models (LLMs) have shown promise in diagnostic reasoning, their impact on management reasoning, which involves balancing treatment decisions and testing strategies while managing risk, is unknown. This prospective, randomized, controlled trial assessed whether LLM assistance improves physician performance on open-ended management reasoning tasks compared to conventional resources. From November 2023 to April 2024, 92 practicing physicians were randomized to use either GPT-4 plus conventional resources or conventional resources alone to answer five expert-developed clinical vignettes in a simulated setting. All cases were based on real, de-identified patient encounters, with information revealed sequentially to mirror the nature of clinical environments. The primary outcome was the difference in total score between groups on expert-developed scoring rubrics. Secondary outcomes included domain-specific scores and time spent per case. Physicians using the LLM scored significantly higher compared to those using conventional resources (mean difference = 6.5%, 95% confidence interval (CI) = 2.7 to 10.2, P < 0.001). LLM users spent more time per case (mean difference = 119.3 s, 95% CI = 17.4 to 221.2, P = 0.02). There was no significant difference between LLM-augmented physicians and LLM alone (-0.9%, 95% CI = -9.0 to 7.2, P = 0.8). LLM assistance can improve physician management reasoning in complex clinical vignettes compared to conventional resources and should be validated in real clinical practice. ClinicalTrials.gov registration: NCT06208423 .

  • Research Article
  • 10.1016/j.ijmedinf.2026.106342
Success and failure of human-AI collaboration in clinical reasoning: An experimental study on challenging real-world cases.
  • May 1, 2026
  • International journal of medical informatics
  • Kai Tzu-Iunn Ong + 7 more

Success and failure of human-AI collaboration in clinical reasoning: An experimental study on challenging real-world cases.

  • Research Article
  • Cite Count Icon 2
  • 10.1001/jamanetworkopen.2026.4003
Large Language Model Performance and Clinical Reasoning Tasks
  • Apr 13, 2026
  • JAMA Network Open
  • Arya S Rao + 16 more

Large language models (LLMs) are increasingly marketed for clinical use, yet their ability to replicate full-spectrum clinical reasoning remains uncertain. Existing evaluations often rely on multiple-choice examinations that do not reflect the complexity of patient care. To evaluate the longitudinal clinical reasoning ability of state-of-the-art LLMs and to introduce a multidimensional, clinically meaningful benchmark for clinical-grade artificial intelligence (AI). In this cross-sectional study, performance was evaluated using standardized clinical vignettes from the January 2025 update of MSD Manual vignettes. A total of 21 off-the-shelf LLMs, including recently released GPT-5, Claude 4.5 Opus, Gemini 3.0 Flash and Pro, and Grok 4, were evaluated. Models were assessed by medical student scorers in triplicate across sequential stages of the standard clinical workflow. Analyses were performed from January to December 2025. The primary outcome was the Proportional Index of Medical Evaluation for LLMs (PrIME-LLM) score, defined as the normalized polygonal area representing balanced accuracy across 5 domains of clinical reasoning as follows: differential diagnosis, diagnostic testing, final diagnosis, management, and miscellaneous clinical reasoning questions. Analyses including analyses of variance, t tests, and regression models were used to compare AI model performance and demographic associations. LLMs were tested across 29 clinical vignettes (representing 16 254 responses in total). PrIME-LLM scores ranged from 0.64 (range, 0.63-0.65) (Gemini 1.5 Flash) to 0.78 (range, 0.77-0.79) (Grok 4), with reasoning-optimized models outperforming nonreasoning models and GPT models scoring highest overall. Differential diagnosis was less accurate than diagnostic testing, while final diagnosis, management, and miscellaneous reasoning were more accurate. Failure rates exceeded 0.80 (range, 0.90-1.00) for differential diagnosis in all models but were less than 0.40 (range, 0.09-0.39) for final diagnosis. Multimodal performance was robust; most LLM models showed improved accuracy with image inputs. In this cross-sectional study of 21 LLMs, frontier LLMs achieved high accuracy on final diagnoses but performed poorly in generating differential diagnoses and navigating uncertainty relative to other reasoning stages. The PrIME-LLM framework provided greater separation than raw accuracy, revealing critical reasoning gaps obscured by traditional benchmarks. Thus, despite version-based improvements and advantages in reasoning-optimized models, off-the-shelf LLMs have not yet achieved the intelligence required for safe deployment and remain limited in demonstrating advanced clinical reasoning.

  • Research Article
  • Cite Count Icon 10
  • 10.1101/2024.08.05.24311485
Large Language Model Influence on Management Reasoning: A Randomized Controlled Trial.
  • Aug 7, 2024
  • medRxiv : the preprint server for health sciences
  • Ethan Goh + 17 more

Large language model (LLM) artificial intelligence (AI) systems have shown promise in diagnostic reasoning, but their utility in management reasoning with no clear right answers is unknown. To determine whether LLM assistance improves physician performance on open-ended management reasoning tasks compared to conventional resources. Prospective, randomized controlled trial conducted from 30 November 2023 to 21 April 2024. Multi-institutional study from Stanford University, Beth Israel Deaconess Medical Center, and the University of Virginia involving physicians from across the United States. 92 practicing attending physicians and residents with training in internal medicine, family medicine, or emergency medicine. Five expert-developed clinical case vignettes were presented with multiple open-ended management questions and scoring rubrics created through a Delphi process. Physicians were randomized to use either GPT-4 via ChatGPT Plus in addition to conventional resources (e.g., UpToDate, Google), or conventional resources alone. The primary outcome was difference in total score between groups on expert-developed scoring rubrics. Secondary outcomes included domain-specific scores and time spent per case. Physicians using the LLM scored higher compared to those using conventional resources (mean difference 6.5 %, 95% CI 2.7-10.2, p<0.001). Significant improvements were seen in management decisions (6.1%, 95% CI 2.5-9.7, p=0.001), diagnostic decisions (12.1%, 95% CI 3.1-21.0, p=0.009), and case-specific (6.2%, 95% CI 2.4-9.9, p=0.002) domains. GPT-4 users spent more time per case (mean difference 119.3 seconds, 95% CI 17.4-221.2, p=0.02). There was no significant difference between GPT-4-augmented physicians and GPT-4 alone (-0.9%, 95% CI -9.0 to 7.2, p=0.8). LLM assistance improved physician management reasoning compared to conventional resources, with particular gains in contextual and patient-specific decision-making. These findings indicate that LLMs can augment management decision-making in complex cases. ClinicalTrials.gov Identifier: NCT06208423; https://classic.clinicaltrials.gov/ct2/show/NCT06208423.

  • Research Article
  • Cite Count Icon 1
  • 10.2196/80167
Digitally Assisted Clinical Decision-Making in Traditional Chinese Medicine: Comparative Study of 5 Large Language Models.
  • Mar 2, 2026
  • JMIR formative research
  • Weiwei Liu + 13 more

Traditional Chinese medicine (TCM) clinical decision-making involves complex integration of syndrome differentiation, constitutional assessment, and individualized treatment selection, creating challenges for standardization and quality assurance. While large language models (LLMs) demonstrate capabilities in medical knowledge integration and clinical reasoning, their application to TCM remains largely unexplored, particularly regarding syndrome differentiation principles and prescription formulation. This study evaluated 5 contemporary LLMs in TCM clinical decision-making and assessed human-artificial intelligence (AI) collaboration compared with independent approaches. Specific objectives were to benchmark LLM performance in TCM knowledge assessment, evaluate clinical case analysis capabilities, identify the optimal model, and assess the quality, efficiency, and acceptability of human-AI collaboration. In total, 5 mainstream LLMs were evaluated-Claude 3.7 Sonnet-Extended (Anthropic), ChatGPT 4.5 (OpenAI), Grok3-DeepSearch (xAI), Gemini 2.0 Flash Thinking Experimental (Google), and DeepSeek-R1 (DeepSeek). The evaluation consisted of four phases, (1) TCM knowledge assessment using 160 standardized questions, (2) clinical case analysis of 30 cases representing different disease systems and complexity levels, (3) optimal model selection using weighted scoring (40% knowledge and 60% clinical analysis), and (4) clinical application assessment involving 10 TCM practitioners and 2 experts comparing physician-only, AI-only, and human-AI collaboration across 5 clinical cases. Statistical analysis included descriptive statistics, reliability analysis, comparative testing, and regression analysis. DeepSeek-R1 demonstrated superior performance across both evaluation domains, achieving 96.7% accuracy in knowledge assessment and 17.31/20 (SD 2.65) in clinical case analysis, significantly outperforming other models (P<.001). Human-AI collaboration achieved significant improvements compared with physician-only decision-making, with 16.1% quality enhancement (33.62 vs 28.97; P<.001) and 66.1% time reduction (162.6 s vs 479.2 s; P<.001). System usability was rated favorably (System Usability Scale score=76.8; P=.002), with high acceptance rates (74.25% adoption, 24% modification, and 1.75% rejection). AI assistance provided the greatest benefits in prescription formulation and medication selection (P<.001). LLMs, particularly DeepSeek-R1, demonstrate substantial capabilities in TCM knowledge assessment and clinical case analysis. Human-AI collaboration significantly enhanced clinical decision-making quality and efficiency while maintaining high physician acceptance. These findings provide compelling evidence for the clinical value of AI-assisted decision-making in TCM, suggesting potential solutions to current challenges in knowledge standardization, clinical training, and health care delivery efficiency. Strategic implementation of AI assistance could significantly enhance the quality, efficiency, and accessibility of TCM care while preserving fundamental principles of individualized treatment.

  • Research Article
  • Cite Count Icon 5
  • 10.1016/j.sleep.2025.106677
Diagnostic performance of Large Language Models (LLMs) compared with physicians in sleep medicine.
  • Oct 1, 2025
  • Sleep medicine
  • Anshum Patel + 5 more

Diagnostic performance of Large Language Models (LLMs) compared with physicians in sleep medicine.

  • Research Article
  • Cite Count Icon 13
  • 10.1152/advan.00209.2024
Transforming medical education: leveraging large language models to enhance PBL-a proof-of-concept study.
  • Jun 1, 2025
  • Advances in physiology education
  • Shoukat Ali Arain + 3 more

The alignment of learning materials with learning objectives (LOs) is critical for successfully implementing the problem-based learning (PBL) curriculum. This study investigated the capabilities of Gemini Advanced, a large language model (LLM), in creating clinical vignettes that align with LOs and comprehensive tutor guides. This study used a faculty-written clinical vignette about diabetes mellitus for third-year medical students. We submitted the LOs and the associated clinical vignette and tutor guide to the LLM to evaluate their alignment and generate new versions. Four faculty members compared both versions, using a structured questionnaire. The mean evaluation scores for original and LLM-generated versions are reported. The LLM identified new triggers for the clinical vignette to align it better with the LOs. Moreover, it restructured the tutor guide for better organization and flow and included thought-provoking questions. The medical information provided by the LLM was scientifically appropriate and accurate. The LLM-generated clinical vignette scored higher (3.0 vs. 1.25) for alignment with the LOs. However, the original version scored better for being educational level-appropriate (2.25 vs. 1.25) and adhering to PBL design (2.50 vs. 1.25). The LLM-generated tutor guide scored higher for better flow (3.0 vs. 1.25), comprehensive and relevant content (2.75 vs. 1.50), and thought-provoking questions (2.25 vs. 1.75). However, LLM-generated learning material lacked visual elements. In conclusion, this study demonstrated that Gemini could align and improve PBL learning materials. By leveraging the potential of LLMs while acknowledging their limitations, medical educators can create innovative and effective learning experiences for future physicians.NEW & NOTEWORTHY This study evaluated a large language model (LLM) (Gemini Advanced) for creating aligned problem-based learning (PBL) materials. The LLM improved the alignment of the clinical vignette with learning goals. The LLM also restructured the tutor guide and added thought-provoking questions. The LLM guide was well organized and informative, but the original vignette was considered more educational level-appropriate. Although the LLM could not generate visuals, AI can improve PBL materials, especially when combined with human expertise.

  • Research Article
  • 10.1093/geroni/igaf122.1180
Application of Large Language Models (LLMs) to Geriatric Practice and Its Evaluation at 4 VA GRECCs
  • Dec 1, 2025
  • Innovation in Aging
  • Huai Cheng + 2 more

LLMs application to clinical practice is growing fast. However, LLMs are less studied in geriatrics practice but are urgently needed. This symposium will address whether LLMs allocation to geriatric practice can be trusted via five approaches. 1) LLMs generated gender and race-biased outputs. We will demonstrate whether LLMs generated age-biased output by assessing their geriatric attitude evaluated by social workers. 2). LLMs passed USMLE and other examinations. We will demonstrate whether LLMs can pass geriatrics knowledge competence tests evaluated by geriatricians 3). LLMs performed well on clinical vignettes from different clinical disciplines. We will demonstrate whether LLMs can perform well on geriatrics 5M-based vignettes of older adults evaluated by clinical providers and trainees 4) LLMs reviewed and summarized clinical charts. We will demonstrate whether LLMs can review geriatrics and general medicine notes to extract Mobility (one of Geriatrics 5Ms) documentation evaluated by geriatricians 5). LLMs can generate deprescribing recommendations, tapering schedules, and patient education materials. We will demonstrate their accuracy, safety, and appropriateness compared to recommendations from a multidisciplinary team of pharmacists, geriatricians, and nurses. Specifically, this symposium will address the following topics: 1) Geriatric Attitude of ChatGPT4.o and Its Evaluation by Social Workers. 2) ChatGPT4.o Geriatrics Knowledge Competency and Its Evaluation by Geriatricians. 3) LLMs application to geriatrics 5Ms evaluated by clinical providers and trainees. 4) Using LLMs to Extract and Assess Mobility Documentation for Age-Friendly Health System evaluated by geriatricians. 5) Using LLMs to generate medication deprescribing recommendations compared to clinician-led deprescribing recommendations.

  • Research Article
  • Cite Count Icon 84
  • 10.1001/jamanetworkopen.2024.12687
Assessing the Risk of Bias in Randomized Clinical Trials With Large Language Models
  • May 22, 2024
  • JAMA Network Open
  • Honghao Lai + 17 more

Large language models (LLMs) may facilitate the labor-intensive process of systematic reviews. However, the exact methods and reliability remain uncertain. To explore the feasibility and reliability of using LLMs to assess risk of bias (ROB) in randomized clinical trials (RCTs). A survey study was conducted between August 10, 2023, and October 30, 2023. Thirty RCTs were selected from published systematic reviews. A structured prompt was developed to guide ChatGPT (LLM 1) and Claude (LLM 2) in assessing the ROB in these RCTs using a modified version of the Cochrane ROB tool developed by the CLARITY group at McMaster University. Each RCT was assessed twice by both models, and the results were documented. The results were compared with an assessment by 3 experts, which was considered a criterion standard. Correct assessment rates, sensitivity, specificity, and F1 scores were calculated to reflect accuracy, both overall and for each domain of the Cochrane ROB tool; consistent assessment rates and Cohen κ were calculated to gauge consistency; and assessment time was calculated to measure efficiency. Performance between the 2 models was compared using risk differences. Both models demonstrated high correct assessment rates. LLM 1 reached a mean correct assessment rate of 84.5% (95% CI, 81.5%-87.3%), and LLM 2 reached a significantly higher rate of 89.5% (95% CI, 87.0%-91.8%). The risk difference between the 2 models was 0.05 (95% CI, 0.01-0.09). In most domains, domain-specific correct rates were around 80% to 90%; however, sensitivity below 0.80 was observed in domains 1 (random sequence generation), 2 (allocation concealment), and 6 (other concerns). Domains 4 (missing outcome data), 5 (selective outcome reporting), and 6 had F1 scores below 0.50. The consistent rates between the 2 assessments were 84.0% for LLM 1 and 87.3% for LLM 2. LLM 1's κ exceeded 0.80 in 7 and LLM 2's in 8 domains. The mean (SD) time needed for assessment was 77 (16) seconds for LLM 1 and 53 (12) seconds for LLM 2. In this survey study of applying LLMs for ROB assessment, LLM 1 and LLM 2 demonstrated substantial accuracy and consistency in evaluating RCTs, suggesting their potential as supportive tools in systematic review processes.

  • Conference Article
  • Cite Count Icon 9
  • 10.1145/3708359.3712094
One Does Not Simply Meme Alone: Evaluating Co-Creativity Between LLMs and Humans in the Generation of Humor
  • Mar 24, 2025
  • Zhikun Wu + 2 more

Collaboration has been shown to enhance creativity, leading to more innovative and effective outcomes. While previous research has explored the abilities of Large Language Models (LLMs) to serve as co-creative partners in tasks like writing poetry or creating narratives, the collaborative potential of LLMs in humor-rich and culturally nuanced domains remains an open question. To address this gap, we conducted a user study to explore the potential of LLMs in co-creating memes---a humor-driven and culturally specific form of creative expression. We conducted a user study with three groups of 50 participants each: a human-only group creating memes without AI assistance, a human-AI collaboration group interacting with a state-of-the-art LLM model, and an AI-only group where the LLM autonomously generated memes. We assessed the quality of the generated memes through crowdsourcing, with each meme rated on creativity, humor, and shareability. Our results showed that LLM assistance increased the number of ideas generated and reduced the effort participants felt. However, it did not improve the quality of the memes when humans were collaborated with LLM. Interestingly, memes created entirely by AI performed better than both human-only and human-AI collaborative memes in all areas on average. However, when looking at the top-performing memes, human-created ones were better in humor, while human-AI collaborations stood out in creativity and shareability. These findings highlight the complexities of human-AI collaboration in creative tasks. While AI can boost productivity and create content that appeals to a broad audience, human creativity remains crucial for content that connects on a deeper level.

  • Research Article
  • 10.3389/fdgth.2025.1719340
Tests of large language models' medical competence and application for clinical decision support of musculoskeletal rehabilitation.
  • Jan 1, 2025
  • Frontiers in digital health
  • Ruikang Liu + 11 more

Large language models (LLMs) are currently abundant and diverse, yet clinicians lack clarity on top performers, with uncertainty about general LLMs' expertise in musculoskeletal rehabilitation. This study aims to investigate the potential and correctness of LLMs in clinical application, and to evaluate whether LLMs could assist primary rehabilitation therapists to prepare for rehabilitation examination. 8 primary doctors and therapists tested 10 LLMs in the first test, 5 senior doctors and therapists assessed answers in the second test, and 5 primary therapists acted as examinees in the third test. We assessed the quality of case analysis based on six different dimensions, including Case Understanding, Clinical Reasoning, Primary Diagnosis, Differential Diagnosis, Treatment Plan Accuracy and Safety, and Guidelines & Consensus. In the first test, only ERNIE Bot X1 Turbo and Doubao 1.5 pro had accuracy rates of over 90%, and Chinese LLMs had significantly fewer incorrect questions than English LLMs (9.6% vs. 14.8%, P < 0.001). In the second test, Doubao 1.5 pro achieved relatively high scores in both cases, and LLMs gained high scores in "Case understanding", "Clinical Reasoning" and "Diagnosis". In the third test, primary therapists achieving a mean accuracy rate of 76.9%, and Doubao 1.5 pro improved its accuracy rates to 85.8%. Doubao 1.5 pro possessed competent ability and application prospects, and was assessed as the best LLM for answering musculoskeletal rehabilitation questions. We also demonstrated that the response quality of local-language LLMs was significantly better than that of English LLMs in answering localized language questions.

  • Conference Article
  • 10.54941/ahfe1006042
Leveraging LLMs to emulate the design processes of different cognitive styles
  • Jan 1, 2025
  • AHFE international
  • Xiyuan Zhang + 5 more

Cognitive styles, which shape designers’ thinking, problem-solving, and decision-making, influence strategies and preferences in design tasks. In team collaboration, diversity cognitive styles enhance problem-solving efficiency, foster creativity, and improve team performance.The ‘Co-evolution of problem–solution’ model serves as a key theoretical framework for understanding differences in designers’ cognitive styles. Based on this model, designers can be categorized into two cognitive styles: problem-driven and solution-driven. Problem-driven designers prioritize structuring the problem before developing solutions, while solution-driven designers generate solutions when design problems still ill-defined, and then work backward to define the problem. Designers with different expertise and disciplinary backgrounds exhibit distinct cognitive style tendencies. Different cognitive styles also adapt differently to design tasks, excelling in some more than others.As a rapidly advancing technology, large language models (LLMs) have shown considerable potential in the field of design. Their powerful generative capabilities position them as potential collaborators in design teams, emulating different cognitive styles. These emulations aim to bridge cognitive differences among team members, enable designers to leverage their individual strengths, and ultimately produce more feasible and high-quality design solutions.However, previous studies have been limited in leveraging LLMs to directly generate design outcomes based on different cognitive styles, neglecting the emulation of the design process itself. In fact, the evolutionary development between problem and solution spaces better reflects the core differences in cognitive styles. Moreover, communication and collaboration within design teams extend beyond simply exchanging solutions, but span multiple stages of the design process—from problem analysis, idea generation, to evaluation. To better integrate LLMs into design teams, it is necessary to consider the emulation of the design cognition process.To this end, our study, based on the cognitive style taxonomy proposed by Dorst and Cross (2001), explores how LLMs can be used to emulate the design processes of problem-driven and solution-driven designers. We develop a zero-shot chain-of-thought (CoT)-based prompting strategy that enables LLMs to emulate the step-by-step cognitive flow of both design styles. The prompt design is inspired by Jiang et al. (2014) and Chen et al. (2023), who analyzed cognitive differences in conceptual design process using the FBS ontology model. Furthermore, to evaluate the effectiveness of LLMs in emulating cognitive styles, this study establishes a three-dimentional evaluation metrics: static distribution (the proportion and preference of cognitive issues), dynamic transformation (behavioral transition patterns), and the creativity of the design outcomes. Using previous studies identified human design behaviours as a benchmark, we compare the cognitive styles emulated by LLMs under different design constraints against human performance to assess their alignment and differences.The results show that LLM-generated design processes align well with human cognitive styles, effectively emulate static cognitive characteristics. Moreover, enhancing novelty and integrity in solutions and demonstrating superior creativity compared to baseline methods. However, LLMs lack the fully complex nonlinear transitions between problem and solution spaces observed in human designers.This process-based emulation has the potential to enhance the application of LLMs in design teams, enabling them to not only serve as tools for generating solutions but also provide support for collaboration during key stages of the design process. Future research should enhance LLMs' reasoning flexibility through fine-tuning or the GoT approach and explore their impact on human-AI collaboration across diverse design tasks to refine their role in design teams.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant