Exploring the Interplay of Dataset Size and Imbalance on CNN Performance in Healthcare: Using X-rays to Identify COVID-19 Patients.

Moshe Davidian,Adi Lahav,Ben-Zion Joshua,Ori Wand,Yotam Lurie,Shlomo Mark

doi:10.3390/diagnostics14161727

Abstract

Convolutional Neural Network (CNN) systems in healthcare are influenced by unbalanced datasets and varying sizes. This article delves into the impact of dataset size, class imbalance, and their interplay on CNN systems, focusing on the size of the training set versus imbalance-a unique perspective compared to the prevailing literature. Furthermore, it addresses scenarios with more than two classification groups, often overlooked but prevalent in practical settings. Initially, a CNN was developed to classify lung diseases using X-ray images, distinguishing between healthy individuals and COVID-19 patients. Later, the model was expanded to include pneumonia patients. To evaluate performance, numerous experiments were conducted with varied data sizes and imbalance ratios for both binary and ternary classifications, measuring various indices to validate the model's efficacy. The study revealed that increasing dataset size positively impacts CNN performance, but this improvement saturates beyond a certain size. A novel finding is that the data balance ratio influences performance more significantly than dataset size. The behavior of three-class classification mirrored that of binary classification, underscoring the importance of balanced datasets for accurate classification. This study emphasizes the fact that achieving balanced representation in datasets is crucial for optimal CNN performance in healthcare, challenging the conventional focus on dataset size. Balanced datasets improve classification accuracy, both in two-class and three-class scenarios, highlighting the need for data-balancing techniques to improve model reliability and effectiveness. Our study is motivated by a scenario with 100 patient samples, offering two options: a balanced dataset with 200 samples and an unbalanced dataset with 500 samples (400 healthy individuals). We aim to provide insights into the optimal choice based on the interplay between dataset size and imbalance, enriching the discourse for stakeholders interested in achieving optimal model performance. Recognizing a single model's generalizability limitations, we assert that further studies on diverse datasets are needed.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Exploring the Interplay of Dataset Size and Imbalance on CNN Performance in Healthcare: Using X-rays to Identify COVID-19 Patients.

Abstract

Talk to us

Similar Papers

More From: Diagnostics (Basel, Switzerland)

Lead the way for us

Journal: Diagnostics (Basel, Switzerland)	Publication Date: Aug 8, 2024
License type: CC BY 4.0

Similar Papers

Comparison of clinical utility of deep learning-based systems for small-bowel capsule endoscopy reading.
...
Journal of Gastroenterology and Hepatology | VOL. 39
, et. al. ...
13 Oct 2023
Journal of Gastroenterology and Hepatology | VOL. 39

Comparison of diagnostic performance between convolutional neural networks and human endoscopists for diagnosis of colorectal polyp: A systematic review and meta-analysis.
Yixin Xu ... Yulin Tan
PloS one | VOL. 16
Yixin Xu, et. al.Yixin Xu ... Yulin Tan
16 Feb 2021
PloS one | VOL. 16

Comparison of diagnostic performance between convolutional neural networks and human endoscopists for diagnosis of colorectal polyp: A systematic review and meta-analysis
Dapeng Wu ... Yixin Xu
-
Dapeng Wu, et. al.Dapeng Wu ... Yixin Xu
16 Feb 2021
16 Feb 2021

Disease surveillance evaluation of primary small-bowel follicular lymphoma using capsule endoscopy images based on a deep convolutional neural network (with video)
Akihiko Sumioka ... Shinji Tanaka
Gastrointestinal Endoscopy | VOL. 98
Akihiko Sumioka, et. al.Akihiko Sumioka ... Shinji Tanaka
22 Jul 2023
Gastrointestinal Endoscopy | VOL. 98

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Exploring the Interplay of Dataset Size and Imbalance on CNN Performance in Healthcare: Using X-rays to Identify COVID-19 Patients.

Abstract

Talk to us

Similar Papers

More From: Diagnostics (Basel, Switzerland)