Random kernel k-nearest neighbors regression.

Patchanok Srisuradetchai,Korn Suksrikran

doi:10.3389/fdata.2024.1402384

Abstract

The k-nearest neighbors (KNN) regression method, known for its nonparametric nature, is highly valued for its simplicity and its effectiveness in handling complex structured data, particularly in big data contexts. However, this method is susceptible to overfitting and fit discontinuity, which present significant challenges. This paper introduces the random kernel k-nearest neighbors (RK-KNN) regression as a novel approach that is well-suited for big data applications. It integrates kernel smoothing with bootstrap sampling to enhance prediction accuracy and the robustness of the model. This method aggregates multiple predictions using random sampling from the training dataset and selects subsets of input variables for kernel KNN (K-KNN). A comprehensive evaluation of RK-KNN on 15 diverse datasets, employing various kernel functions including Gaussian and Epanechnikov, demonstrates its superior performance. When compared to standard KNN and the random KNN (R-KNN) models, it significantly reduces the root mean square error (RMSE) and mean absolute error, as well as improving R-squared values. The RK-KNN variant that employs a specific kernel function yielding the lowest RMSE will be benchmarked against state-of-the-art methods, including support vector regression, artificial neural networks, and random forests.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Random kernel k-nearest neighbors regression.

Abstract

Talk to us

Similar Papers

More From: Frontiers in big data

Lead the way for us

Journal: Frontiers in big data	Publication Date: Jul 1, 2024
License type: CC BY 4.0

Similar Papers

Hybrid Particle Swarm and Gray Wolf optimization for Prediction of Appliances in Low-Energy Houses
El-Sayed M El-Kenawy ... Said H Abd Elkhalik
-
El-Sayed M El-Kenawy, et. al.El-Sayed M El-Kenawy ... Said H Abd Elkhalik
26 Jul 2022
26 Jul 2022

Presenting a soft sensor for monitoring and controlling well health and pump performance using machine learning, statistical analysis, and Petri net modeling.
Mohammad Hossein Amini ... Maliheh Arab
Environmental Science and Pollution Research | VOL. -
Mohammad Hossein Amini, et. al.Mohammad Hossein Amini ... Maliheh Arab
10 Feb 2021
Environmental Science and Pollution Research | VOL. -

Inclusion of fractal dimension in four machine learning algorithms improves the prediction accuracy of mean weight diameter of soil
Abhradip Sarkar ... Arti Bhatia
Ecological Informatics | VOL. 74
Abhradip Sarkar, et. al.Abhradip Sarkar ... Arti Bhatia
17 Dec 2022
Ecological Informatics | VOL. 74

AB1138 APPLICATION OF MACHINE LEARNING METHODS TO PREDICT SF-36 MENTAL HEALTH DOMAIN FOR SYSTEMIC LUPUS ERYTHEMATOSUS PATIENTS
E Gaydukova
Annals of the Rheumatic Diseases | VOL. 83
E GaydukovaE Gaydukova
01 Jun 2024
AB1138 APPLICATION OF MACHINE LEARNING METHODS TO PREDICT SF-36 MENTAL HEALTH DOMAIN FOR SYSTEMIC LUPUS ERYTHEMATOSUS PATIENTS
E Gaydukova

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Random kernel k-nearest neighbors regression.

Abstract

Talk to us

Similar Papers

More From: Frontiers in big data