Abstract

Long non-coding RNAs (lncRNAs) have begun to receive overdue attention for their regulatory roles in gene expression and other cellular processes. Although most lncRNAs are lowly expressed and tissue-specific, notable exceptions include MALAT1 and its genomic neighbor NEAT1, two highly and ubiquitously expressed oncogenes with roles in transcriptional regulation and RNA splicing. Previous studies have suggested that NEAT1 is found only in mammals, while MALAT1 is present in all gnathostomes (jawed vertebrates) except birds. Here we show that these assertions are incomplete, likely due to the challenges associated with properly identifying these two lncRNAs. Using phylogenetic analysis and structure-aware annotation of publicly available genomic and RNA-seq coverage data, we show that NEAT1 is a common feature of tetrapod genomes except birds and squamates. Conversely, we identify MALAT1 in representative species of all major gnathostome clades, including birds. Our in-depth examination of MALAT1, NEAT1, and their genomic context in a wide range of vertebrate species allows us to reconstruct the series of events that led to the formation of the locus containing these genes in taxa from cartilaginous fish to mammals. This evolutionary history includes the independent loss of NEAT1 in birds and squamates, since NEAT1 is found in the closest living relatives of both clades (crocodilians and tuataras, respectively). These data clarify the origins and relationships of MALAT1 and NEAT1 and highlight an opportunity to study the change and continuity in lncRNA structure and function over deep evolutionary time.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call