Poisson Approximation for the Number of Repeats in a Stationary Markov Chain

Narjiss Touyar,Sophie Schbath,Dominique Cellier,Hélène Dauchel

doi:10.1017/s0021900200004344

Abstract

Detection of repeated sequences within complete genomes is a powerful tool to help understanding genome dynamics and species evolutionary history. To distinguish significant repeats from those that can be obtained just by chance, statistical methods have to be developed. In this paper we show that the distribution of the number of long repeats in long sequences generated by stationary Markov chains can be approximated by a Poisson distribution with explicit parameter. Thanks to the Chen-Stein method we provide a bound for the approximation error; this bound converges to 0 as soon as the length n of the sequence tends to ∞ and the length t of the repeats satisfies n 2ρ t = O(1) for some 0 &lt; ρ &lt; 1. Using this Poisson approximation, p-values can then be easily calculated to determine if a given genome is significantly enriched in repeats of length t.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Poisson Approximation for the Number of Repeats in a Stationary Markov Chain

Abstract

Talk to us

Similar Papers

More From: Journal of Applied Probability

Lead the way for us

Journal: Journal of Applied Probability	Publication Date: Jun 1, 2008
License type: mit

Similar Papers

Poisson Approximation for the Number of Repeats in a Stationary Markov Chain
Narjiss Touyar ... Hélène Dauchel
Journal of Applied Probability | VOL. 45
Narjiss Touyar, et. al.Narjiss Touyar ... Hélène Dauchel
01 Jun 2008
Journal of Applied Probability | VOL. 45

Some Thoughts About Data Type, Distribution, and Statistical Significance
D Scot Malay
The Journal of Foot and Ankle Surgery | VOL. 45
D Scot MalayD Scot Malay
01 Nov 2006
The Journal of Foot and Ankle Surgery | VOL. 45

Testing the Adequacy of Markov Chain and Mover-Stayer Models as Representations of Credit Behavior
Halina Frydman ... Jarl G Kallberg
Operations Research | VOL. 33
Halina Frydman, et. al.Halina Frydman ... Jarl G Kallberg
01 Dec 1985
Operations Research | VOL. 33

Analysis of Longitudinal Data of Epileptic Seizure Counts – A Two‐State Hidden Markov Regression Approach
Peiming Wang ... Martin L Puterman
Biometrical Journal | VOL. 43
Peiming Wang, et. al.Peiming Wang ... Martin L Puterman
01 Dec 2001
Biometrical Journal | VOL. 43

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Poisson Approximation for the Number of Repeats in a Stationary Markov Chain

Abstract

Talk to us

Similar Papers

More From: Journal of Applied Probability