Performance of Chatgpt in ophthalmology exam; human versus AI.

Ali Safa Balci,Zeliha Yazar,Banu Turgut Ozturk,Cigdem Altan

doi:10.1007/s10792-024-03353-w

Abstract

This cross-sectional study focuses on evaluating the success rate of ChatGPT in answering questions from the 'Resident Training Development Exam' and comparing these results with the performance of the ophthalmology residents. The 75 exam questions, across nine sections and three difficulty levels, were presented to ChatGPT. The responses and explanations were recorded. The readability and complexity of the explanations were analyzed and The Flesch Reading Ease (FRE) score (0-100) was recorded using the program named Readable. Residents were categorized into four groups based on their seniority. The overall and seniority-specific success rates of the residents were compared separately with ChatGPT. Out of 69 questions, ChatGPT answered 37 correctly (53.62%). The highest success was in Lens and Cataract (77.77%), and the lowest in Pediatric Ophthalmology and Strabismus (0.00%). Of 789 residents, overall accuracy was 50.37%. Seniority-specific accuracy rates were 43.49%, 51.30%, 54.91%, and 60.05% for 1st to 4th-year residents. ChatGPT ranked 292nd among residents. Difficulty-wise, 11 questions were easy, 44 moderate, and 14 difficult. ChatGPT's accuracy for each level was 63.63%, 54.54%, and 42.85%, respectively. The average FRE score of responses generated by ChatGPT was found to be 27.56 ± 12.40. ChatGPT correctly answered 53.6% of questions in an exam for residents. ChatGPT has a lower success rate on average than a 3rd year resident. The readability of responses provided by ChatGPT is low, and they are difficult to understand. As difficulty increases, ChatGPT's success decreases. Predictably, these results will change with more information loaded into ChatGPT.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Performance of Chatgpt in ophthalmology exam; human versus AI.

Abstract

Talk to us

Similar Papers

More From: International ophthalmology

Lead the way for us

Similar Papers

Readability Analysis of Patient-Accessible Information Regarding Ambulatory Surgical Center Procedures
Conor P Lynch ... Nathaniel W Jenkins
International Journal of Spine Surgery | VOL. 15
Conor P Lynch, et. al.Conor P Lynch ... Nathaniel W Jenkins
01 Oct 2021
International Journal of Spine Surgery | VOL. 15

Preparation, Validation and User-testing of Patient Information Leaflets on Diabetes and Hypertension
Santosha Vooradi ... G Thunga
Indian Journal of Pharmaceutical Sciences | VOL. 80
Santosha Vooradi, et. al.Santosha Vooradi ... G Thunga
01 Jan 2018
Indian Journal of Pharmaceutical Sciences | VOL. 80

Facial Paralysis Online Educational Resources: Readability and Benefit to Patient Education.
Mithila Somasundaram ... Christine B Novak
Plastic and reconstructive surgery | VOL. 148
Mithila Somasundaram, et. al.Mithila Somasundaram ... Christine B Novak
09 Jul 2021
Plastic and reconstructive surgery | VOL. 148

Quality of Online Health Information Regarding Sickle Cell Disease Transition
Steffi Shilly ... Sophia Jan
Blood | VOL. 132
Steffi Shilly, et. al.Steffi Shilly ... Sophia Jan
29 Nov 2018
Blood | VOL. 132

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Performance of Chatgpt in ophthalmology exam; human versus AI.

Abstract

Talk to us

Similar Papers

More From: International ophthalmology