Automated Extraction of Tumor Staging and Diagnosis Information From Surgical Pathology Reports.

Sajjad Abedian,Marika M Cusick,Stephanie E Weiner,Jim C Hu,Thomas R Campion,Evan T Sholle,Prakash M Adekkanattu,Jonathan E Shoag

doi:10.1200/cci.21.00065

Sajjad Abedian, Marika M Cusick + Show 6 more

Open Access

https://doi.org/10.1200/cci.21.00065

Copy DOI

Abstract

Typically stored as unstructured notes, surgical pathology reports contain data elements valuable to cancer research that require labor-intensive manual extraction. Although studies have described natural language processing (NLP) of surgical pathology reports to automate information extraction, efforts have focused on specific cancer subtypes rather than across multiple oncologic domains. To address this gap, we developed and evaluated an NLP method to extract tumor staging and diagnosis information across multiple cancer subtypes. The NLP pipeline was implemented on an open-source framework called Leo. We used a total of 555,681 surgical pathology reports of 329,076 patients to develop the pipeline and evaluated our approach on subsets of reports from patients with breast, prostate, colorectal, and randomly selected cancer subtypes. Averaged across all four cancer subtypes, the NLP pipeline achieved an accuracy of 1.00 for International Classification of Diseases, Tenth Revision codes, 0.89 for T staging, 0.90 for N staging, and 0.97 for M staging. It achieved an F1 score of 1.00 for International Classification of Diseases, Tenth Revision codes, 0.88 for T staging, 0.90 for N staging, and 0.24 for M staging. The NLP pipeline was developed to extract tumor staging and diagnosis information across multiple cancer subtypes to support the research enterprise in our institution. Although it was not possible to demonstrate generalizability of our NLP pipeline to other institutions, other institutions may find value in adopting a similar NLP approach-and reusing code available at GitHub-to support the oncology research enterprise with elements extracted from surgical pathology reports.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Automated Extraction of Tumor Staging and Diagnosis Information From Surgical Pathology Reports.

Abstract

Talk to us

Similar Papers

More From: JCO Clinical Cancer Informatics

Lead the way for us

Journal: JCO Clinical Cancer Informatics	Publication Date: Dec 1, 2021
Citations: 12

Similar Papers

Cost-efficient quality assurance of natural language processing tools through continuous monitoring with continuous integration
Marc Schreiber ... Bodo Kraft
-
Marc Schreiber, et. al.Marc Schreiber ... Bodo Kraft
14 May 2016
14 May 2016

Identification of Preanesthetic History Elements by a Natural Language Processing Engine.
Harrison S Suh ... Rodney A Gabriel
Anesthesia & Analgesia | VOL. 135
Harrison S Suh, et. al.Harrison S Suh ... Rodney A Gabriel
15 Jul 2022
Anesthesia & Analgesia | VOL. 135

A Methodological Approach to Validate Pneumonia Encounters from Radiology Reports Using Natural Language Processing.
Aloksagar Panny ... Harshad Hegde
Methods of Information in Medicine | VOL. 61
Aloksagar Panny, et. al.Aloksagar Panny ... Harshad Hegde
01 May 2022
Methods of Information in Medicine | VOL. 61

Information extraction from free text for aiding transdiagnostic psychiatry: constructing NLP pipelines tailored to clinicians’ needs
Rosanne J Turner ... Karin Hagoort
BMC Psychiatry | VOL. 22
Rosanne J Turner, et. al.Rosanne J Turner ... Karin Hagoort
17 Jun 2022
BMC Psychiatry | VOL. 22

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Automated Extraction of Tumor Staging and Diagnosis Information From Surgical Pathology Reports.

Abstract

Talk to us

Similar Papers

More From: JCO Clinical Cancer Informatics