Contextual feature selection for text classification

Francois Paradis,Jian-Yun Nie

doi:10.1016/j.ipm.2006.07.006

Contextual feature selection for text classification

Francois Paradis, Jian-Yun Nie

https://doi.org/10.1016/j.ipm.2006.07.006

Copy DOI

Journal: Information Processing and Management	Publication Date: Oct 24, 2006
Citations: 12

Affiliation: Université de Montréal

#Feature Selection For Text Classification #In-house Collection + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

We present a simple approach for the classification of “noisy” documents using bigrams and named entities. The approach combines conventional feature selection with a contextual approach to filter out passages around selected features. Originally designed for call for tender documents, the method can be useful for other web collections that also contain non-topical contents. Experiments are conducted on our in-house collection as well as on the 4-Universities data set, Reuters 21578 and 20 Newsgroups. We find a significant improvement on our collection and the 4-Universities data set (10.9% and 4.1%, respectively). Although the best results are obtained by combining bigrams and named entities, the impact of the latter is not found to be significant.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: Information Processing and Management

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.