Crowdsourcing large scale wrapper inference

Valter Crescenzi,Paolo Merialdo,Disheng Qiu

doi:10.1007/s10619-014-7163-9

Valter Crescenzi, Paolo Merialdo + Show 1 more

Open Access

https://doi.org/10.1007/s10619-014-7163-9

Copy DOI

Journal: Distributed and Parallel Databases	Publication Date: Oct 29, 2014
Citations: 53	License type: other-oa

Affiliation: Roma Tre University

Abstract

We present a crowdsourcing system for large-scale production of accurate wrappers to extract data from data-intensive websites. Our approach is based on supervised wrapper inference algorithms which demand the burden of generating training data to workers recruited on a crowdsourcing platform. Workers are paid for answering simple queries carefully chosen by the system. We present two algorithms: a single worker algorithm (\({\textsc {alf}}_{\eta }\)) and a multiple workers algorithm (alfred). Both the algorithms deal with the inherent uncertainty of the workers’ responses and use an active learning approach to select the most informative queries. alfred estimates the workers’ error rate to decide at runtime how many workers should be recruited to achieve a quality target. The system has been fully implemented and tested: the experimental evaluation conducted with both synthetic workers and real workers recruited on a crowdsourcing platform show that our approach is able to produce accurate wrappers at a low cost, even in presence of workers with a significant error rate.

Full Text