K Nearest Neighbor OveRsampling approach: An open source python package for data augmentation

Ashhadul Islam,Halima Bensmail,Samir Brahim Belhaouari,Atiq Ur Rehman

doi:10.1016/j.simpa.2022.100272

Ashhadul Islam, Halima Bensmail + Show 2 more

Open Access

https://doi.org/10.1016/j.simpa.2022.100272

Copy DOI

Journal: Software Impacts	Publication Date: May 1, 2022
Citations: 3	License type: cc-by

Affiliation: Hamad bin Khalifa University

Abstract

Data is present in abundance, but the problem of imbalanced dataset crops up time and again, vexing classifiers and reducing accuracy. This paper introduces K Nearest Neighbor OveRsampling (KNNOR) Algorithm — a novel data augmentation technique that considers the distribution of data and takes into account the k nearest neighbors while generating artificial data points. The KNNOR algorithm has outperformed the state-of-the-art augmentation algorithms by enabling classifiers to achieve much higher accuracy after injecting artificial minority datapoints into imbalanced datasets. This method is useful especially in health datasets where an imbalance is common and can even be applied to images of lower dimensions.

Full Text