Assessing GPT-4's Performance in Delivering Medical Advice: Comparative Analysis With Human Experts.

Eunbeen Jo,Sanghoun Song,Jong-Ho Kim,Jong-Ho Kim,Subin Lim,Ju Hyeon Kim,Jung-Joon Cha,Young-Min Kim,Hyung Joon Joo,Hyung Joon Joo,Hyung Joon Joo

doi:10.2196/51282

Abstract

Accurate medical advice is paramount in ensuring optimal patient care, and misinformation can lead to misguided decisions with potentially detrimental health outcomes. The emergence of large language models (LLMs) such as OpenAI's GPT-4 has spurred interest in their potential health care applications, particularly in automated medical consultation. Yet, rigorous investigations comparing their performance to human experts remain sparse. This study aims to compare the medical accuracy of GPT-4 with human experts in providing medical advice using real-world user-generated queries, with a specific focus on cardiology. It also sought to analyze the performance of GPT-4 and human experts in specific question categories, including drug or medication information and preliminary diagnoses. We collected 251 pairs of cardiology-specific questions from general users and answers from human experts via an internet portal. GPT-4 was tasked with generating responses to the same questions. Three independent cardiologists (SL, JHK, and JJC) evaluated the answers provided by both human experts and GPT-4. Using a computer interface, each evaluator compared the pairs and determined which answer was superior, and they quantitatively measured the clarity and complexity of the questions as well as the accuracy and appropriateness of the responses, applying a 3-tiered grading scale (low, medium, and high). Furthermore, a linguistic analysis was conducted to compare the length and vocabulary diversity of the responses using word count and type-token ratio. GPT-4 and human experts displayed comparable efficacy in medical accuracy ("GPT-4 is better" at 132/251, 52.6% vs "Human expert is better" at 119/251, 47.4%). In accuracy level categorization, humans had more high-accuracy responses than GPT-4 (50/237, 21.1% vs 30/238, 12.6%) but also a greater proportion of low-accuracy responses (11/237, 4.6% vs 1/238, 0.4%; P=.001). GPT-4 responses were generally longer and used a less diverse vocabulary than those of human experts, potentially enhancing their comprehensibility for general users (sentence count: mean 10.9, SD 4.2 vs mean 5.9, SD 3.7; P<.001; type-token ratio: mean 0.69, SD 0.07 vs mean 0.79, SD 0.09; P<.001). Nevertheless, human experts outperformed GPT-4 in specific question categories, notably those related to drug or medication information and preliminary diagnoses. These findings highlight the limitations of GPT-4 in providing advice based on clinical experience. GPT-4 has shown promising potential in automated medical consultation, with comparable medical accuracy to human experts. However, challenges remain particularly in the realm of nuanced clinical judgment. Future improvements in LLMs may require the integration of specific clinical reasoning pathways and regulatory oversight for safe use. Further research is needed to understand the full potential of LLMs across various medical specialties and conditions.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: JMIR medical education	Publication Date: Jul 8, 2024
Citations: 1	License type: cc-by

R Discovery Prime

R Discovery Prime

Assessing GPT-4's Performance in Delivering Medical Advice: Comparative Analysis With Human Experts.

Abstract

Talk to us

Similar Papers

More From: JMIR medical education

Lead the way for us

Similar Papers

Large Language Models Can Enable Inductive Thematic Analysis of a Social Media Corpus in a Single Prompt: Human Validation Study.
Michael S Deiner ... Urmimala Sarkar
JMIR infodemiology | VOL. 4
Michael S Deiner, et. al.Michael S Deiner ... Urmimala Sarkar
29 Aug 2024
JMIR infodemiology | VOL. 4

Leveraging Large Language Models for Improved Patient Access and Self-Management: Assessor-Blinded Comparison Between Expert- and AI-Generated Content.
Xiaolei Lv ... Xiaomeng Zhang
Journal of Medical Internet Research | VOL. 26
Xiaolei Lv, et. al.Xiaolei Lv ... Xiaomeng Zhang
25 Apr 2024
Journal of Medical Internet Research | VOL. 26

Evaluating large language models for selection of statistical test for research: A pilot study
Himel Mondal ... Shaikat Mondal
Perspectives in Clinical Research | VOL. 15
Himel Mondal, et. al.Himel Mondal ... Shaikat Mondal
08 Apr 2024
Perspectives in Clinical Research | VOL. 15

Leveraging Large Language Models for Decision Support in Personalized Oncology
Manuela Benary ... Damian T Rieke
JAMA network open | VOL. 6
Manuela Benary, et. al.Manuela Benary ... Damian T Rieke
17 Nov 2023
JAMA network open | VOL. 6

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Assessing GPT-4's Performance in Delivering Medical Advice: Comparative Analysis With Human Experts.

Abstract

Talk to us

Similar Papers

More From: JMIR medical education