Effects of Training Set Size on Supervised Machine-Learning Land-Cover Classification of Large-Area High-Resolution Remotely Sensed Data

Christopher A Ramezan,Aaron E Maxwell,Bradley S Price,Timothy A Warner

doi:10.3390/rs13030368

Christopher A Ramezan, Aaron E Maxwell + Show 2 more

Open Access

PDF Available

https://doi.org/10.3390/rs13030368

Copy DOI

Export

Save

Cite

Journal: Remote Sensing	Publication Date: Jan 21, 2021
Citations: 105	License type: CC BY 4.0

Affiliation: West Virginia University

Abstract
Highlights/Summary
Full-Text PDF
Similar Papers

Abstract

Listen

The size of the training data set is a major determinant of classification accuracy. Nevertheless, the collection of a large training data set for supervised classifiers can be a challenge, especially for studies covering a large area, which may be typical of many real-world applied projects. This work investigates how variations in training set size, ranging from a large sample size (n = 10,000) to a very small sample size (n = 40), affect the performance of six supervised machine-learning algorithms applied to classify large-area high-spatial-resolution (HR) (1–5 m) remotely sensed data within the context of a geographic object-based image analysis (GEOBIA) approach. GEOBIA, in which adjacent similar pixels are grouped into image-objects that form the unit of the classification, offers the potential benefit of allowing multiple additional variables, such as measures of object geometry and texture, thus increasing the dimensionality of the classification input data. The six supervised machine-learning algorithms are support vector machines (SVM), random forests (RF), k-nearest neighbors (k-NN), single-layer perceptron neural networks (NEU), learning vector quantization (LVQ), and gradient-boosted trees (GBM). RF, the algorithm with the highest overall accuracy, was notable for its negligible decrease in overall accuracy, 1.0%, when training sample size decreased from 10,000 to 315 samples. GBM provided similar overall accuracy to RF; however, the algorithm was very expensive in terms of training time and computational resources, especially with large training sets. In contrast to RF and GBM, NEU, and SVM were particularly sensitive to decreasing sample size, with NEU classifications generally producing overall accuracies that were on average slightly higher than SVM classifications for larger sample sizes, but lower than SVM for the smallest sample sizes. NEU however required a longer processing time. The k-NN classifier saw less of a drop in overall accuracy than NEU and SVM as training set size decreased; however, the overall accuracies of k-NN were typically less than RF, NEU, and SVM classifiers. LVQ generally had the lowest overall accuracy of all six methods, but was relatively insensitive to sample size, down to the smallest sample sizes. Overall, due to its relatively high accuracy with small training sample sets, and minimal variations in overall accuracy between very large and small sample sets, as well as relatively short processing time, RF was a good classifier for large-area land-cover classifications of HR remotely sensed data, especially when training data are scarce. However, as performance of different supervised classifiers varies in response to training set size, investigating multiple classification algorithms is recommended to achieve optimal accuracy for a project.

Highlights

The highest average overall accuracy was 99.8%, for the random forests (RF) classifications trained from the 10,000-sample set, while the lowest average overall accuracy was 87.4% for the NEU
The highest average overall accuracy was 99.8%, for the RF classifications trained from the 10,00012 of 27 sample set, while the lowest average overall accuracy was 87.4% for the NEU classifications trained from 40 samples
An investigation on training set size and classifier response that incorporates multiple datasets, study areas, and sensor types would be valuable. This analysis explored the effects of the number of training samples, varying from to 10,000, on six supervised machine-learning algorithms, support vector machines (SVM), RF, k-nearest neighbors (k-NN), NEU, learning vector quantization (LVQ), GBM, to classify a large-area HR remotely sensed dataset

Summary

Introduction

In circumstances where the number of training data is limited, or where constraints in processing power or time limit the number of training samples that can be processed, it would be advantageous to know the relative dependence of machine-learning classifiers on sample size. Most previous studies comparing supervised machine-learning classifier accuracy have used a single, fixed training sample size [2,3,4], and have ignored the effects of variation in sample size. Investigations that have examined the effects of sample size, for example [1,5,6], have generally focused on a single classifier, making it difficult to compare the relative dependence of machine-learning classifiers on sample size

Methods

Results

Discussion

Conclusion