Abstract

While the identification of small variants in panel sequencing data can be considered a solved problem, the identification of larger, multi-exon copy number variants (CNVs) still poses a considerable challenge. Thus, CNV calling has not been established in all laboratories performing panel sequencing. At the same time, such laboratories have accumulated large datasets and thus have the need to identify CNVs on their data to close the diagnostic gap. In this article, we present our method clearCNV that addresses this need in two ways. First, it helps laboratories to properly assign datasets to enrichment kits. Based on homogeneous subsets of data, clearCNV identifies CNVs affecting the targeted regions. Using real-world datasets and validation, we show that our method is highly competitive with previous methods and preferable in terms of specificity. The software is available for free under a permissible license at https://github.com/bihealth/clear-cnv. Supplementary data are available at Bioinformatics online.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call