Simple-random-sampling-based multiclass text classification algorithm.

Wuying Liu,Lin Wang,Mianzhu Yi

doi:10.1155/2014/517498

Simple-random-sampling-based multiclass text classification algorithm.

Wuying Liu, Lin Wang + Show 1 more

Open Access

https://doi.org/10.1155/2014/517498

Copy DOI

Journal: TheScientificWorldJournal	Publication Date: Jan 1, 2014
Citations: 14	License type: CC BY 3.0

Affiliation: National University of Defense Technology

#Multiclass Text Classification #Era Of Big Data + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

Multiclass text classification (MTC) is a challenging issue and the corresponding MTC algorithms can be used in many applications. The space-time overhead of the algorithms must be concerned about the era of big data. Through the investigation of the token frequency distribution in a Chinese web document collection, this paper reexamines the power law and proposes a simple-random-sampling-based MTC (SRSMTC) algorithm. Supported by a token level memory to store labeled documents, the SRSMTC algorithm uses a text retrieval approach to solve text classification problems. The experimental results on the TanCorp data set show that SRSMTC algorithm can achieve the state-of-the-art performance at greatly reduced space-time requirements.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: TheScientificWorldJournal

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.