Employing Structural and Textual Feature Extraction for Semistructured Document Classification

Mohammad Khabbaz,Keivan Kianmehr,Reda Alhajj

doi:10.1109/tsmcc.2012.2208102

Abstract

This paper addresses XML document classification by considering both structural and content-based features of the documents. This approach leads to better constructing a set of informative feature vectors that represents both structural and textual aspects of XML documents. For this purpose, we integrate soft clustering of words and feature reduction into the process. To extract structural information, we employ an existing frequent tree-mining algorithm combined with an information gain filter to retrieve the most informative substructures from XML documents. However, for extracting content information, we propose soft clustering of words using each cluster as a textual feature. We have conducted extensive experiments on a benchmark dataset, namely 20NewsGroups, and an XML documents dataset given in LOGML that describes the web-server logs of user sessions. With regards to the classifier built only using our textual features, the results show that it outperforms a naive support-vector-machine (SVM)-based classifier, as well as an information retrieval classifier (IRC). We further demonstrate the effectiveness of incorporating both structural and content information into the process of learning, by comparing our classifier model and several XML document classifiers. In particular, by applying SVM and decision tree algorithms using our feature vector representation of XML documents dataset, we have achieved 85.79% and 87.04% classification accuracy, respectively, which are higher than accuracy achieved by XRules, a well-known structural-based XML document classifier.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Employing Structural and Textual Feature Extraction for Semistructured Document Classification

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews)

Lead the way for us

Journal: IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews)	Publication Date: Nov 1, 2012
Citations: 13

Similar Papers

Class-specific GMM based intermediate matching kernel for classification of varying length patterns of long duration speech using support vector machines
A.D Dileep ... C Chandra Sekhar
Speech Communication | VOL. 57
A.D Dileep, et. al.A.D Dileep ... C Chandra Sekhar
07 Oct 2013
Speech Communication | VOL. 57

GMM-based intermediate matching kernel for classification of varying length patterns of long duration speech using support vector machines.
Aroor Dinesh Dileep ... Chellu Chandra Sekhar
IEEE Transactions on Neural Networks and Learning Systems | VOL. 25
Aroor Dinesh Dileep, et. al.Aroor Dinesh Dileep ... Chellu Chandra Sekhar
01 Aug 2014
IEEE Transactions on Neural Networks and Learning Systems | VOL. 25

Speaker recognition using pyramid match kernel based support vector machines
A D Dileep ... C Chandra Sekhar
International Journal of Speech Technology | VOL. 15
A D Dileep, et. al.A D Dileep ... C Chandra Sekhar
12 Jun 2012
International Journal of Speech Technology | VOL. 15

Performance Analysis of Classifiers with Feature Selection and Optimization in CBIR System for Biological Images
Shajee Mohan ... K S Shanthini
-
Shajee Mohan, et. al.Shajee Mohan ... K S Shanthini
01 Jan 2014
01 Jan 2014

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Employing Structural and Textual Feature Extraction for Semistructured Document Classification

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews)