
A main challenge for Web content classification is how to model the input data. This paper discusses the application of two text modeling approaches, latent semantic analysis (LSA) and latent Dirichlet allocation (LDA), in the Web page classification task. We report results on a comparison of these two approaches using different vocabularies consisting of links and text. Both models are evaluated using different numbers of latent topics. Finally, we evaluate a hybrid latent variable model that combines the latent topics resulting from both LSA and LDA. This new approach turns out to be superior to the basic LSA and LDA models. In our experiments with categories and pages obtained from the ODP Web directory the hybrid model achieves an averaged F-measure value of 0.852 and an averaged ROC value of 0.96.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call