A parallel Hyper-Surface Classifier for high dimensional data

Qing He,Zhong-Zhi Shi,Xu-Dong Ma,Qun Wang Qun Wang,Chang-Ying Du Chang-Ying Du

doi:10.1109/kam.2010.5646172

Qing He, Zhong-Zhi Shi + Show 3 more

https://doi.org/10.1109/kam.2010.5646172

Copy DOI

Export

Save

Cite

Publication Date: Oct 1, 2010

Citations: 3

Affiliation: Institute of Computing Technology

Abstract
Full-Text
Similar Papers

Abstract

Listen

The enlarging volumes of data resources produced in real world makes classification of very large scale data a challenging task. Therefore, parallel process of very large high dimensional data is very important. Hyper-Surface Classification (HSC) is approved to be an effective and efficient classification algorithm to handle two and three dimensional data. Though HSC can be extended to deal with high dimensional data with dimension reduction or ensemble techniques, it is not trivial to tackle high dimensional data directly. Inspired by the decision tree idea, an improvement of HSC is proposed to deal with high dimensional data directly in this work. Furthermore, we parallelize the improved HSC algorithm (PHSC) to handle large scale high dimensional data based on MapReduce framework, which is a current and powerful parallel programming technique used in many fields. Experimental results show that the parallel improved HSC algorithm not only can directly deal with high dimensional data, but also can handle large scale data set. Furthermore, the evaluation criterions of scaleup, speedup and sizeup validate its efficiency.

Full Text