Influence of Model Evolution and System Roles on ChatGPT's Performance in Chinese Medical Licensing Exams: Comparative Study.

Shuai Ming,Qingge Guo,Wenjun Cheng,Bo Lei

doi:10.2196/52784

Abstract

With the increasing application of large language models like ChatGPT in various industries, its potential in the medical domain, especially in standardized examinations, has become a focal point of research. The aim of this study is to assess the clinical performance of ChatGPT, focusing on its accuracy and reliability in the Chinese National Medical Licensing Examination (CNMLE). The CNMLE 2022 question set, consisting of 500 single-answer multiple choices questions, were reclassified into 15 medical subspecialties. Each question was tested 8 to 12 times in Chinese on the OpenAI platform from April 24 to May 15, 2023. Three key factors were considered: the version of GPT-3.5 and 4.0, the prompt's designation of system roles tailored to medical subspecialties, and repetition for coherence. A passing accuracy threshold was established as 60%. The χ2 tests and κ values were employed to evaluate the model's accuracy and consistency. GPT-4.0 achieved a passing accuracy of 72.7%, which was significantly higher than that of GPT-3.5 (54%; P<.001). The variability rate of repeated responses from GPT-4.0 was lower than that of GPT-3.5 (9% vs 19.5%; P<.001). However, both models showed relatively good response coherence, with κ values of 0.778 and 0.610, respectively. System roles numerically increased accuracy for both GPT-4.0 (0.3%-3.7%) and GPT-3.5 (1.3%-4.5%), and reduced variability by 1.7% and 1.8%, respectively (P>.05). In subgroup analysis, ChatGPT achieved comparable accuracy among different question types (P>.05). GPT-4.0 surpassed the accuracy threshold in 14 of 15 subspecialties, while GPT-3.5 did so in 7 of 15 on the first response. GPT-4.0 passed the CNMLE and outperformed GPT-3.5 in key areas such as accuracy, consistency, and medical subspecialty expertise. Adding a system role insignificantly enhanced the model's reliability and answer coherence. GPT-4.0 showed promising potential in medical education and clinical practice, meriting further study.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: JMIR medical education	Publication Date: Aug 13, 2024
Citations: 1	License type: cc-by

R Discovery Prime

R Discovery Prime

Influence of Model Evolution and System Roles on ChatGPT's Performance in Chinese Medical Licensing Exams: Comparative Study.

Abstract

Talk to us

Similar Papers

More From: JMIR medical education

Lead the way for us

Similar Papers

Finding Historical Context in Medical Regulation: A Bibliographical Guide
David A Johnson
-
David A JohnsonDavid A Johnson
01 Dec 2019
01 Dec 2019

ChatGPT's performance in German OB/GYN exams - paving the way for AI-enhanced medical education and clinical practice.
Maximilian Riedel ... Bastian Meyer
Frontiers in medicine | VOL. 10
Maximilian Riedel, et. al.Maximilian Riedel ... Bastian Meyer
13 Dec 2023
Frontiers in medicine | VOL. 10

Improving ventilation classification in under-actuated zones: a k-nearest neighbor and data preprocessing approach
Yaddarabullah Yaddarabullah ... Amna Saad
Indonesian Journal of Electrical Engineering and Computer Science | VOL. 34
Yaddarabullah Yaddarabullah, et. al.Yaddarabullah Yaddarabullah ... Amna Saad
01 Apr 2024
Indonesian Journal of Electrical Engineering and Computer Science | VOL. 34

University of Miami Leonard M. Miller School of Medicine
...
Academic Medicine | VOL. 85
, et. al. ...
01 Sep 2010
Academic Medicine | VOL. 85

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Influence of Model Evolution and System Roles on ChatGPT's Performance in Chinese Medical Licensing Exams: Comparative Study.

Abstract

Talk to us

Similar Papers

More From: JMIR medical education