A strategy for identification of Web query interfaces using supervised learning

Heidy M Marin-Castro,Ivan Lopez-Arevalo,Victor J Sosa-Sosa

doi:10.1109/nwesp.2011.6088183

Abstract

The Deep Web is an enormous source of information constantly growing. It comprises a large amount of databases on the Web that are accessed through Web query interfaces related to different domains. The content of the Deep Web can not be reachable by traditional search engines, what makes almost impossible for common users to get access to this useful information. There are several problems related to the search for content in the Deep Web. One of them is the automatic identification of Web query interfaces, being this a mean to access the information in the Deep Web. The task of classify HTML forms contained inside Web page as Web query interface is challenging due to their enormous heterogeneity. This paper introduce a strategy that automatically identifies Web query interfaces independent of their domain. We make an adequate selection of HTML elements and use them appropriately to build characteristic vectors that are used as input of a supervised classifier to determine if a Web page contains or not a Web query interface. The experimental results show that the proposed strategy is efficient and accurate, achieving better classification results than works previously reported.

Full Text