GPT-4 performance on querying scientific publications: reproducibility, accuracy, and impact of an instruction sheet

Kaiming Tao,Zachary A Osman,Philip L Tzou,Soo-Yon Rhee,Vineet Ahluwalia,Robert W Shafer

doi:10.1186/s12874-024-02253-y

Abstract

BackgroundLarge language models (LLMs) that can efficiently screen and identify studies meeting specific criteria would streamline literature reviews. Additionally, those capable of extracting data from publications would enhance knowledge discovery by reducing the burden on human reviewers.MethodsWe created an automated pipeline utilizing OpenAI GPT-4 32 K API version “2023–05-15” to evaluate the accuracy of the LLM GPT-4 responses to queries about published papers on HIV drug resistance (HIVDR) with and without an instruction sheet. The instruction sheet contained specialized knowledge designed to assist a person trying to answer questions about an HIVDR paper. We designed 60 questions pertaining to HIVDR and created markdown versions of 60 published HIVDR papers in PubMed. We presented the 60 papers to GPT-4 in four configurations: (1) all 60 questions simultaneously; (2) all 60 questions simultaneously with the instruction sheet; (3) each of the 60 questions individually; and (4) each of the 60 questions individually with the instruction sheet.ResultsGPT-4 achieved a mean accuracy of 86.9% – 24.0% higher than when the answers to papers were permuted. The overall recall and precision were 72.5% and 87.4%, respectively. The standard deviation of three replicates for the 60 questions ranged from 0 to 5.3% with a median of 1.2%. The instruction sheet did not significantly increase GPT-4’s accuracy, recall, or precision. GPT-4 was more likely to provide false positive answers when the 60 questions were submitted individually compared to when they were submitted together.ConclusionsGPT-4 reproducibly answered 3600 questions about 60 papers on HIVDR with moderately high accuracy, recall, and precision. The instruction sheet's failure to improve these metrics suggests that more sophisticated approaches are necessary. Either enhanced prompt engineering or finetuning an open-source model could further improve an LLM's ability to answer questions about highly specialized HIVDR papers.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

GPT-4 performance on querying scientific publications: reproducibility, accuracy, and impact of an instruction sheet

Abstract

Talk to us

Similar Papers

More From: BMC Medical Research Methodology

Lead the way for us

Journal: BMC Medical Research Methodology	Publication Date: Jun 25, 2024
License type: CC BY 4.0

Similar Papers

Human Immunodeficiency Virus (HIV) Drug Resistance: A Global Narrative Review
Maureen Nkandu Phiri ... Steward Mudenda
Journal of Biomedical Research & Environmental Sciences | VOL. 2
Maureen Nkandu Phiri, et. al.Maureen Nkandu Phiri ... Steward Mudenda
01 Sep 2021
Journal of Biomedical Research & Environmental Sciences | VOL. 2

Recommendations on data sharing in HIV drug resistance research.
...
PLOS Medicine | VOL. 20
, et. al. ...
22 Sep 2023
PLOS Medicine | VOL. 20

Maternal Human Immunodeficiency Virus (HIV) Drug Resistance Is Associated With Vertical Transmission and Is Prevalent in Infected Infants.
Ceejay L Boyce ... Patricia Demarrais
Clinical Infectious Diseases | VOL. 74
Ceejay L Boyce, et. al.Ceejay L Boyce ... Patricia Demarrais
01 Sep 2021
Clinical Infectious Diseases | VOL. 74

Pretreatment Human Immunodeficiency Virus (HIV) Drug Resistance Among Treatment-Naive Infants Newly Diagnosed With HIV in 2016 in Namibia: Results of a Nationally Representative Study.
Michael R Jordan ... Eric J Dziuban
Open Forum Infectious Diseases | VOL. 9
Michael R Jordan, et. al.Michael R Jordan ... Eric J Dziuban
24 Mar 2022
Open Forum Infectious Diseases | VOL. 9

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

GPT-4 performance on querying scientific publications: reproducibility, accuracy, and impact of an instruction sheet

Abstract

Talk to us

Similar Papers

More From: BMC Medical Research Methodology