Ed-Join

Chuan Xiao,Xuemin Lin,Wei Wang

doi:10.14778/1453856.1453957

Abstract

There has been considerable interest in similarity join in the research community recently. Similarity join is a fundamental operation in many application areas, such as data integration and cleaning, bioinformatics, and pattern recognition. We focus on efficient algorithms for similarity join with edit distance constraints. Existing approaches are mainly based on converting the edit distance constraint to a weaker constraint on the number of matching q -grams between pair of strings. In this paper, we propose the novel perspective of investigating mismatching q -grams. Technically, we derive two new edit distance lower bounds by analyzing the locations and contents of mismatching q -grams. A new algorithm, Ed-Join, is proposed that exploits the new mismatch-based filtering methods; it achieves substantial reduction of the candidate sizes and hence saves computation time. We demonstrate experimentally that the new algorithm outperforms alternative methods on large-scale real datasets under a wide range of parameter settings.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Ed-Join

Abstract

Talk to us

Similar Papers

More From: Proceedings of the VLDB Endowment

Lead the way for us

Journal: Proceedings of the VLDB Endowment	Publication Date: Aug 1, 2008
Citations: 273

Similar Papers

VChunkJoin: An Efficient Algorithm for Edit Similarity Joins
Wei Wang ... Chuan Xiao
IEEE Transactions on Knowledge and Data Engineering | VOL. 25
Wei Wang, et. al.Wei Wang ... Chuan Xiao
01 Aug 2013
IEEE Transactions on Knowledge and Data Engineering | VOL. 25

Similarity joins for uncertain strings
Manish Patil ... Rahul Shah
-
Manish Patil, et. al.Manish Patil ... Rahul Shah
18 Jun 2014
18 Jun 2014

Landmark-Join: Hash-Join Based String Similarity Joins with Edit Distance Constraints
Kazuyo Narita ... Shinji Nakadai
-
Kazuyo Narita, et. al.Kazuyo Narita ... Shinji Nakadai
01 Jan 2012
01 Jan 2012

A partition-based approach to structure similarity search
Xiang Zhao ... Qing Liu
Proceedings of the VLDB Endowment | VOL. 7
Xiang Zhao, et. al.Xiang Zhao ... Qing Liu
01 Nov 2013
Proceedings of the VLDB Endowment | VOL. 7

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Ed-Join

Abstract

Talk to us

Similar Papers

More From: Proceedings of the VLDB Endowment