The Smoothed Complexity of Edit Distance

Alexandr Andoni,Robert Krauthgamer

doi:10.1007/978-3-540-70575-8_30

Abstract

We initiate the study of the smoothed complexity of sequence alignment, by proposing a semi-random model of edit distance between two input strings, generated as follows. First, an adversary chooses two binary strings of length d and a longest common subsequence A of them. Then, every character is perturbed independently with probability p, except that A is perturbed in exactly the same way inside the two strings.We design two efficient algorithms that compute the edit distance on smoothed instances up to a constant factor approximation. The first algorithm runs in near-linear time, namely d 1 + ε for any fixed ε> 0. The second one runs in time sublinear in d, assuming the edit distance is not too small. These approximation and runtime guarantees are significantly better then the bounds known for worst-case inputs, e.g. near-linear time algorithm achieving approximation roughly d 1/3, due to Batu, Ergün, and Sahinalp [SODA 2006].Our technical contribution is twofold. First, we rely on finding matches between substrings in the two strings, where two substrings are considered a match if their edit distance is relatively small, a prevailing technique in commonly used heuristics, such as PatternHunter of Ma, Tromp and Li [Bioinformatics, 2002]. Second, we effectively reduce the smoothed edit distance to a simpler variant of (worst-case) edit distance, namely, edit distance on permutations (a.k.a. Ulam’s metric). We are thus able to build on algorithms developed for the Ulam metric, whose much better algorithmic guarantees usually do not carry over to general edit distance.KeywordsEdit DistanceOptimal AlignmentLonge Common SubsequenceSmooth ComplexityCandidate MatchThese keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

The Smoothed Complexity of Edit Distance

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

The smoothed complexity of edit distance
Alexandr Andoni ... Robert Krauthgamer
ACM Transactions on Algorithms | VOL. 8
Alexandr Andoni, et. al.Alexandr Andoni ... Robert Krauthgamer
01 Sep 2012
ACM Transactions on Algorithms | VOL. 8

How Compression and Approximation Affect Efficiency in String Distance Measures
Arun Ganesh ... Barna Saha
-
Arun Ganesh, et. al.Arun Ganesh ... Barna Saha
01 Jan 2021
01 Jan 2021

The Dyck Language Edit Distance Problem in Near-Linear Time
Barna Saha
-
Barna SahaBarna Saha
01 Oct 2014
01 Oct 2014

Improved Algorithms for Edit Distance and LCS: Beyond Worst Case
Mahdi Boroujeni ... Saeed Seddighin
-
Mahdi Boroujeni, et. al.Mahdi Boroujeni ... Saeed Seddighin
01 Jan 2020
01 Jan 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

The Smoothed Complexity of Edit Distance

Abstract

Talk to us

Similar Papers