Improving the Reliability of Peer Review Without a Gold Standard

Tarmo Äijö,Daniel Elgort,Murray Becker,Richard Herzog,Richard K. J. Brown,Benjamin L. Odry,Ron Vianu

doi:10.1007/s10278-024-00971-9

Abstract

Peer review plays a crucial role in accreditation and credentialing processes as it can identify outliers and foster a peer learning approach, facilitating error analysis and knowledge sharing. However, traditional peer review methods may fall short in effectively addressing the interpretive variability among reviewing and primary reading radiologists, hindering scalability and effectiveness. Reducing this variability is key to enhancing the reliability of results and instilling confidence in the review process. In this paper, we propose a novel statistical approach called “Bayesian Inter-Reviewer Agreement Rate” (BIRAR) that integrates radiologist variability. By doing so, BIRAR aims to enhance the accuracy and consistency of peer review assessments, providing physicians involved in quality improvement and peer learning programs with valuable and reliable insights. A computer simulation was designed to assign predefined interpretive error rates to hypothetical interpreting and peer-reviewing radiologists. The Monte Carlo simulation then sampled (100 samples per experiment) the data that would be generated by peer reviews. The performances of BIRAR and four other peer review methods for measuring interpretive error rates were then evaluated, including a method that uses a gold standard diagnosis. Application of the BIRAR method resulted in 93% and 79% higher relative accuracy and 43% and 66% lower relative variability, compared to “Single/Standard” and “Majority Panel” peer review methods, respectively. Accuracy was defined by the median difference of Monte Carlo simulations between measured and pre-defined “actual” interpretive error rates. Variability was defined by the 95% CI around the median difference of Monte Carlo simulations between measured and pre-defined “actual” interpretive error rates. BIRAR is a practical and scalable peer review method that produces more accurate and less variable assessments of interpretive quality by accounting for variability within the group’s radiologists, implicitly applying a standard derived from the level of consensus within the group across various types of interpretive findings.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Improving the Reliability of Peer Review Without a Gold Standard

Abstract

Talk to us

Similar Papers

More From: Journal of imaging informatics in medicine

Lead the way for us

Journal: Journal of imaging informatics in medicine	Publication Date: Feb 5, 2024
License type: CC BY 4.0

Similar Papers

Give until it hurts.
Gautam Allahbadia
The Journal of Obstetrics and Gynecology of India | VOL. 64
Gautam AllahbadiaGautam Allahbadia
12 Mar 2014
The Journal of Obstetrics and Gynecology of India | VOL. 64

Peer review on the Internet: A better class of conversation
Craig Bingham
The Lancet | VOL. 351
Craig BinghamCraig Bingham
01 Mar 1998
The Lancet | VOL. 351

Peer Review: The Best of the Blemished?
Joseph S Alpert
The American Journal of Medicine | VOL. 120
Joseph S AlpertJoseph S Alpert
29 Mar 2007
The American Journal of Medicine | VOL. 120

JID Innovations and Peer Review
Russell P Hall
Journal of Investigative Dermatology Innovations | VOL. 1
Russell P HallRussell P Hall
01 Sep 2021
Journal of Investigative Dermatology Innovations | VOL. 1

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Improving the Reliability of Peer Review Without a Gold Standard

Abstract

Talk to us

Similar Papers

More From: Journal of imaging informatics in medicine