Diversity in Renal Mass Data Cohorts: Implications for Urology AI Researchers

Harmony Selena Cen,Siddhartha Dandamudi,Xiaomeng Lei,Chris Weight,Mihir Desai,Inderbir Gill,Vinay Duddalwar

doi:10.1159/000535841

Abstract

Introduction: We examine the heterogeneity and distribution of the cohort populations in two publicly used radiological image cohorts, the Cancer Genome Atlas Kidney Renal Clear Cell Carcinoma (TCIA TCGA KIRC) collection and 2019 MICCAI Kidney Tumor Segmentation Challenge (KiTS19), and deviations in real-world population renal cancer data from the National Cancer Database (NCDB) Participant User Data File (PUF) and tertiary center data. PUF data are used as an anchor for prevalence rate bias assessment. Specific gene expression and, therefore, biology of RCC differ by self-reported race, especially between the African American and Caucasian populations. AI algorithms learn from datasets, but if the dataset misrepresents the population, reinforcing bias may occur. Ignoring these demographic features may lead to inaccurate downstream effects, thereby limiting the translation of these analyses to clinical practice. Consciousness of model training biases is vital to patient care decisions when using models in clinical settings. Methods: Data elements evaluated included gender, demographics, reported pathologic grading, and cancer staging. American Urological Association risk levels were used. Poisson regression was performed to estimate the population-based and sample-specific estimation for prevalence rate and corresponding 95% confidence interval. SAS 9.4 was used for data analysis. Results: Compared to PUF, KiTS19 and TCGA KIRC oversampled Caucasian by 9.5% (95% CI, −3.7 to 22.7%) and 15.1% (95% CI, 1.5 to 28.8%), undersampled African American by −6.7% (95% CI, −10% to −3.3%), and −5.5% (95% CI, −9.3% to −1.8%). Tertiary also undersampled African American by −6.6% (95% CI, −8.7% to −4.6%). The tertiary cohort largely undersampled aggressive cancers by −14.7% (95% CI, −20.9% to −8.4%). No statistically significant difference was found among PUF, TCGA, and KiTS19 in aggressive rate; however, heterogeneities in risk are notable. Conclusion: Heterogeneities between cohorts need to be considered in future AI training and cross-validation for renal masses.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Oncology	Publication Date: Dec 15, 2023
Citations: 1	License type: CC BY-NC 4.0

R Discovery Prime

R Discovery Prime

Diversity in Renal Mass Data Cohorts: Implications for Urology AI Researchers

Abstract

Talk to us

Similar Papers

More From: Oncology

Lead the way for us

Similar Papers

Sequencing of Renal Mass Biopsy and Ablation: Results from the National Cancer Database.
Annemarie Uhlig ... Andrew Lenis
Urology practice | VOL. 8
Annemarie Uhlig, et. al.Annemarie Uhlig ... Andrew Lenis
24 Jun 2021
Urology practice | VOL. 8

Kidney-specific cadherin, a specific marker for the distal portion of the nephron and related renal neoplasms
Steven S Shen ... Luan D Truong
Modern Pathology | VOL. 18
Steven S Shen, et. al.Steven S Shen ... Luan D Truong
01 Jul 2005
Modern Pathology | VOL. 18

Focus on kidney cancer
W.Marston Linehan ... Berton Zbar
Cancer Cell | VOL. 6
W.Marston Linehan, et. al.W.Marston Linehan ... Berton Zbar
01 Sep 2004
Cancer Cell | VOL. 6

Activity of tivozanib (AV-951) in patients (Pts) with different histologic subtypes of renal cell carcinoma (RCC).
P Bhargava ... R T Chacko
Journal of Clinical Oncology | VOL. 29
P Bhargava, et. al.P Bhargava ... R T Chacko
01 Mar 2011
Journal of Clinical Oncology | VOL. 29

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Diversity in Renal Mass Data Cohorts: Implications for Urology AI Researchers

Abstract

Talk to us

Similar Papers

More From: Oncology