Web spam detection based on improved tri-training

Hailong Li

doi:10.1109/pic.2014.6972296

Abstract

Web spamming is the deliberate manipulation of search engine indexes to make a page get high ranking than which it deserved considering its true value. Since the evolution of web spam, a new based on machine learning algorithm web spam detection method which has self-learning ability has emerged. Web spam detection is viewed as a binary classification learning problem. Because labeled training examples are fairly expensive to obtain which need the participation of experts in this field and labor costs, how to fully utilize a large number of unlabeled web page examples on the web is a challenge faced by web spam detection. In this paper, we present a web spam detection algorithm according to improve tri-training. It uses a small amount of labeled examples and a large number of unlabeled examples to train classifiers, which can reduce the cost of labeled examples and improve the learning performance. Both web page content features and link features are used in this paper.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Web spam detection based on improved tri-training

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Word Sense Disambiguation by Learning Decision Trees from Unlabeled Data
Seong-Bae Park
Applied Intelligence | VOL. 19
Seong-Bae ParkSeong-Bae Park
01 Jan 2003
Applied Intelligence | VOL. 19

Gaussian process versus margin sampling active learning
Jin Zhou ... Shiliang Sun
Neurocomputing | VOL. 167
Jin Zhou, et. al.Jin Zhou ... Shiliang Sun
14 May 2015
Neurocomputing | VOL. 167

SemiBoost: Boosting for Semi-Supervised Learning
P.K Mallapragada ... Rong Jin
IEEE Transactions on Pattern Analysis and Machine Intelligence | VOL. 31
P.K Mallapragada, et. al.P.K Mallapragada ... Rong Jin
01 Nov 2009
IEEE Transactions on Pattern Analysis and Machine Intelligence | VOL. 31

Query-Based Video Event Definition Using Rough Set Theory and High-Dimensional Representation
Kimiaki Shirahama ... Kuniaki Uehara
-
Kimiaki Shirahama, et. al.Kimiaki Shirahama ... Kuniaki Uehara
01 Jan 2009
01 Jan 2009

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Web spam detection based on improved tri-training

Abstract

Talk to us

Similar Papers