Italian Text Categorization with Lemmatization and Support Vector Machines

Francesco Camastra,Gennaro Razi

doi:10.1007/978-981-13-8950-4_5

Abstract

The paper describes an Italian language text categorizer by Lemmatization and support vector machines. The categorizer is composed of six modules. The first module performs the tokenization, removing the punctuation signs; the second and third ones carry out stopping and lemmatization, respectively; the fourth module implements the bag-of-words approach; the fifth one performs feature dimensionality reduction eliminating poor discriminant features; finally, the last module does the classification. The Italian text categorizer has been validated on a database composed of more than 1100 articles, extracted from online edition of three Italian language newspapers, belonging to eight different categories. The work is highly novel, since to the best our knowledge, there are no works in literature on Italian text categorization.

Full Text