Identity and compatibility of reference genome resources.

Michał Stolarczyk,Bingjie Xue,Nathan C Sheffield

doi:10.1093/nargab/lqab036

Michał Stolarczyk, Bingjie Xue + Show 1 more

Open Access

https://doi.org/10.1093/nargab/lqab036

Copy DOI

Journal: NAR genomics and bioinformatics	Publication Date: Apr 9, 2021
Citations: 9	License type: CC BY-NC 4.0

Affiliation: University of Virginia

Abstract

Genome analysis relies on reference data like sequences, feature annotations, and aligner indexes. These data can be found in many versions from many sources, making it challenging to identify and assess compatibility among them. For example, how can you determine which indexes are derived from identical raw sequence files, or which annotations share a compatible coordinate system? Here, we describe a novel approach to establish identity and compatibility of reference genome resources. We approach this with three advances: first, we derive unique identifiers for each resource; second, we record parent–child relationships among resources; and third, we describe recursive identifiers that determine identity as well as compatibility of coordinate systems and sequence names. These advances facilitate portability, reproducibility, and re-use of genome reference data. Available athttps://refgenie.databio.org.

Full Text