Enhancing mitogenomic phylogeny and resolving the relationships of extinct megafaunal placental mammals
Enhancing mitogenomic phylogeny and resolving the relationships of extinct megafaunal placental mammals
- Research Article
43
- 10.1098/rstb.2009.0046
- Aug 12, 2009
- Philosophical Transactions of the Royal Society B: Biological Sciences
The genome sequence is an icon of early twenty-first century biology. Genomes of nearly 2000 cellular organisms, and from many thousands of organelles and viruses, are now in the public domain. For biological research in individual species, the genome sequence increasingly provides the common
- Research Article
366
- 10.1038/nature14249
- Mar 18, 2015
- Nature
No large group of recently extinct placental mammals remains as evolutionarily cryptic as the approximately 280 genera grouped as 'South American native ungulates'. To Charles Darwin, who first collected their remains, they included perhaps the 'strangest animal[s] ever discovered'. Today, much like 180 years ago, it is no clearer whether they had one origin or several, arose before or after the Cretaceous/Palaeogene transition 66.2million years ago, or are more likely to belong with the elephants and sirenians of superorder Afrotheria than with the euungulates (cattle, horses, and allies) of superorder Laurasiatheria. Morphology-based analyses have proved unconvincing because convergences are pervasive among unrelated ungulate-like placentals. Approaches using ancient DNA have also been unsuccessful, probably because of rapid DNA degradation in semitropical and temperate deposits. Here we apply proteomic analysis to screen bone samples of the Late Quaternary South American native ungulate taxa Toxodon (Notoungulata) and Macrauchenia (Litopterna) for phylogenetically informative protein sequences. For each ungulate, we obtain approximately 90% direct sequence coverage of type I collagen α1- and α2-chains, representing approximately 900 of 1,140 amino-acid residues for each subunit. A phylogeny is estimated from an alignment of these fossil sequences with collagen (I) gene transcripts from available mammalian genomes or mass spectrometrically derived sequence data obtained for this study. The resulting consensus tree agrees well with recent higher-level mammalian phylogenies. Toxodon and Macrauchenia form a monophyletic group whose sister taxon is not Afrotheria or any of its constituent clades as recently claimed, but instead crown Perissodactyla (horses, tapirs, and rhinoceroses). These results are consistent with the origin of at least some South American native ungulates from 'condylarths', a paraphyletic assembly of archaic placentals. With ongoing improvements in instrumentation and analytical procedures, proteomics may produce a revolution in systematics such as that achieved by genomics, but with the possibility of reaching much further back in time.
- Research Article
41
- 10.3389/fevo.2020.00094
- Apr 15, 2020
- Frontiers in Ecology and Evolution
Key insights into the evolutionary history of recently extinct or critically endangered species can be obtained through analysis of genomic data collected using high-throughput sequencing and ancient DNA from museum specimens, particularly where specimens are rare. For instance, the evolutionary history of the critically endangered Puebla deer mouse, Peromyscus mekisturus, remains unclear due to discordance between morphological and molecular phylogenetic analyses. However, previous molecular analyses were based on PCR and Sanger sequencing of only a few mitochondrial genes. Here, we used ancient DNA from historical museum specimens followed by target enrichment and high-throughput sequencing of several thousand nuclear ultraconserved elements and whole mitochondrial genomes to test the validity of the previous phylogenetic placement of P. mekisturus. Based on UCEs and mitogenomes, our results revealed that P. mekisturus forms a well-supported distinct lineage outside the clade containing all other members of the Peromyscus melanophrys group. Additionally, the mitogenome phylogeny further supports the placement of P. mekisturus as the sister species of the genus Reithrodontomys. This conflicts with the previous mtDNA phylogenetic reconstruction, in which P. mekisturus was nested within the species P. melanophrys. Our study demonstrates that high-throughput sequencing of ancient DNA, appropriately controlling for contamination and degradation, can provide a robust resolution of the evolutionary history and taxonomic status of species for which few or no modern genetic samples exist. In light of our results and pending further analysis with denser taxon sampling and the addition of morphological data, a re-evaluation of the taxonomy and conservation management plans of P. mekisturus is needed to ensure that the evolutionary distinctiveness of this species is recognized in future conservation efforts.
- Research Article
124
- 10.1007/s10539-010-9208-4
- May 27, 2010
- Biology & Philosophy
The ‘Tree of Life’ is intended to represent the pattern of evolutionary processes that result in bifurcating species lineages. Often justified in reference to Darwin’s discussions of trees, the Tree of Life has run up against numerous challenges especially in regard to prokaryote evolution. This special issue examines scientific, historical and philosophical aspects of debates about the Tree of Life, with the aim of turning these criticisms towards a reconstruction of prokaryote phylogeny and even some aspects of the standard evolutionary understanding of eukaryotes. These discussions have arisen out of a multidisciplinary collaboration of people with an interest in the Tree of Life, and we suggest that this sort of focused engagement enables a practical understanding of the relationships between biology, philosophy and history.
- Conference Article
6
- 10.1109/icpads.2005.240
- Jul 20, 2005
Accurate reconstruction of phylogenetic trees very often involves solving hard optimization problems, particularly the maximum parsimony (MP) and maximum likelihood (ML) problems. Various heuristics have been devised for solving these two problems; however, they obtain good results within reasonable time only on small datasets. This has been a major impediment for large-scale phylogeny reconstruction, particularly for the effort to assemble the Tree of Life - the evolutionary relationship of all organisms on earth. Roshan et al. recently introduced Rec-I-DCM3, an efficient and accurate meta-method for solving the MP problem on large datasets of up to 14,000 taxa. Nonetheless, a drastic improvement in Rec-I-DCM3's performance is still needed in order to achieve similar (or better) accuracy on datasets at the scale of the Tree of Life. In this paper, we improve the performance of Rec-I-DCM3 via parallelization. Experimental results demonstrate that our parallel method, PRec-I-DCM3, achieves significant improvements, both in speed and accuracy, over its sequential counterpart
- Research Article
4
- 10.1360/n972015-01355
- Feb 22, 2016
- Chinese Science Bulletin
Since the concept of a “Tree of Life” was raised by Charles Darwin, researches in this field have not only contributed to our understanding of phylogenetic relationships among taxa, but also significantly accelerated the development of related subjects in biological science. Evolutionary biologist Dobzhansky once remarked that “nothing makes sense in biology except in the light of evolution”, which has been largely echoed by later biologists. Indeed, reconstruction of an accurate phylogeny of the living world is very important for biological classification and nomenclature, and also crucial to elucidate the origin and diversification of life. We have experienced three major phases for Tree of Life reconstruction in the past century. Prior to the 1990s, taxonomists published classification systems that were largely dependent on morphological characters. DNA sequencing technology facilitated by the development of polymerase chain reaction (PCR) techniques has allowed systematists to reconstruct phylogenetic relationships using molecular data. More recently, the rapid development of next-generation sequencing tools has brought the Tree of Life to a phylogenomic era by enabling the construction of phylogenies using hundreds or thousands of loci from organellar and nuclear genomes. However, significant conflicts have been detected in phylogenies of various organisms with the large increase in the number of loci used for phylogenetic analyses. Given the level of conflict in some data sets, some researchers have begun to doubt the accuracy and congruence of the Tree of Life and its applications in related biological fields. So, will there ever be a Tree of Life that systematists can agree on?
- Research Article
52
- 10.1093/sysbio/syab027
- Apr 10, 2021
- Systematic Biology
Six-state amino acid recoding strategies are commonly applied to combat the effects of compositional heterogeneity and substitution saturation in phylogenetic analyses. While these methods have been endorsed from a theoretical perspective, their performance has never been extensively tested. Here, we test the effectiveness of six-state recoding approaches by comparing the performance of analyses on recoded and non-recoded data sets that have been simulated under gradients of compositional heterogeneity or saturation. In our simulation analyses, non-recoding approaches consistently outperform six-state recoding approaches. Our results suggest that six-state recoding strategies are not effective in the face of high saturation. Furthermore, while recoding strategies do buffer the effects of compositional heterogeneity, the loss of information that accompanies six-state recoding outweighs its benefits. In addition, we evaluate recoding schemes with 9, 12, 15, and 18 states and show that these consistently outperform six-state recoding. Our analyses of other recoding schemes suggest that under conditions of very high compositional heterogeneity, it may be advantageous to apply recoding using more than six states, but we caution that applying any recoding should include sufficient justification. Our results have important implications for the more than 90 published papers that have incorporated six-state recoding, many of which have significant bearing on relationships across the tree of life. [Compositional heterogeneity; Dayhoff 6-state recoding; S&R 6-state recoding; six-state amino acid recoding; substitution saturation.]
- Research Article
58
- 10.1007/s00239-005-0216-y
- Oct 1, 2006
- Journal of Molecular Evolution
Evolution operates on whole genomes through direct rearrangements of genes, such as inversions, transpositions, and inverted transpositions, as well as through operations, such as duplications, losses, and transfers, that also affect the gene content of the genomes. Because these events are rare relative to nucleotide substitutions, gene order data offer the possibility of resolving ancient branches in the tree of life; the combination of gene order data with sequence data also has the potential to provide more robust phylogenetic reconstructions, since each can elucidate evolution at different time scales. Distance corrections greatly improve the accuracy of phylogeny reconstructions from DNA sequences, enabling distance-based methods to approach the accuracy of the more elaborate methods based on parsimony or likelihood at a fraction of the computational cost. This paper focuses on developing distance correction methods for phylogeny reconstruction from whole genomes. The main question we investigate is how to estimate evolutionary histories from whole genomes with equal gene content, and we present a technique, the empirically derived estimator (EDE), that we have developed for this purpose. We study the use of EDE on whole genomes with identical gene content, and we explore the accuracy of phylogenies inferred using EDE with the neighbor joining and minimum evolution methods under a wide range of model conditions. Our study shows that tree reconstruction under these two methods is much more accurate when based on EDE distances than when based on other distances previously suggested for whole genomes.
- Research Article
154
- 10.1098/rspb.2014.2671
- May 7, 2015
- Proceedings of the Royal Society B: Biological Sciences
Since the late eighteenth century, fossils of bizarre extinct creatures have been described from the Americas, revealing a previously unimagined chapter in the history of mammals. The most bizarre of these are the ‘native’ South American ungulates thought to represent a group of mammals that evolved in relative isolation on South America, but with an uncertain affinity to any particular placental lineage. Many authors have considered them descended from Laurasian ‘condylarths’, which also includes the probable ancestors of perissodactyls and artiodactyls, whereas others have placed them either closer to the uniquely South American xenarthrans (anteaters, armadillos and sloths) or the basal afrotherians (e.g. elephants and hyraxes). These hypotheses have been debated owing to conflicting morphological characteristics and the hitherto inability to retrieve molecular information. Of the ‘native’ South American mammals, only the toxodonts and litopterns persisted until the Late Pleistocene–Early Holocene. Owing to known difficulties in retrieving ancient DNA (aDNA) from specimens from warm climates, this research presents a molecular phylogeny for both Macrauchenia patachonica (Litopterna) and Toxodon platensis (Notoungulata) recovered using proteomics-based (liquid chromatography–tandem mass spectrometry) sequencing analyses of bone collagen. The results place both taxa in a clade that is monophyletic with the perissodactyls, which today are represented by horses, rhinoceroses and tapirs.
- Research Article
52
- 10.1371/journal.pone.0139611
- Nov 5, 2015
- PLOS ONE
For over 200 years, fossils of bizarre extinct creatures have been described from the Americas that have ranged from giant ground sloths to the ‘native’ South American ungulates, groups of mammals that evolved in relative isolation on South America. Ground sloths belong to the South American xenarthrans, a group with modern although morphologically and ecologically very different representatives (anteaters, armadillos and sloths), which has been proposed to be one of the four main eutherian clades. Recently, proteomics analyses of bone collagen have recently been used to yield a molecular phylogeny for a range of mammals including the unusual ‘Malagasy aardvark’ shown to be most closely related to the afrotherian tenrecs, and the south American ungulates supporting their morphological association with condylarths. However, proteomics results generate partial sequence information that could impact upon the phylogenetic placement that has not been appropriately tested. For comparison, this paper examines the phylogenetic potential of proteomics-based sequencing through the analysis of collagen extracted from two extinct giant ground sloths, Lestodon and Megatherium. The ground sloths were placed as sister taxa to extant sloths, but with a closer relationship between Lestodon and the extant sloths than the basal Megatherium. These results highlight that proteomics methods could yield plausible phylogenies that share similarities with other methods, but have the potential to be more useful in fossils beyond the limits of ancient DNA survival.
- Research Article
22
- 10.1016/j.ympev.2019.106576
- Aug 2, 2019
- Molecular Phylogenetics and Evolution
Ancient DNA from a 2,500-year-old Caribbean fossil places an extinct bird (Caracara creightoni) in a phylogenetic context
- Research Article
73
- 10.1016/j.ympev.2018.12.019
- Dec 17, 2018
- Molecular Phylogenetics and Evolution
New patellogastropod mitogenomes help counteracting long-branch attraction in the deep phylogeny of gastropod mollusks
- Research Article
15
- 10.1016/j.ympev.2004.10.020
- Dec 15, 2004
- Molecular Phylogenetics and Evolution
Reconstructing evolutionary relationships from functional data: a consistent classification of organisms based on translation inhibition response
- Abstract
5
- 10.1186/1471-2105-15-s3-a8
- Feb 1, 2014
- BMC Bioinformatics
Background The way gene families and genomes evolve can be understood in detail only when the location of gene duplication episodes in the tree of life can be deciphered. Since most genes belong to larger gene families, the analysis of the gene family histories thus plays an important role in the study of genome evolution. Empirically, one frequently observes that the tree that describes the evolution of species, the species tree, is inconsistent with the tree that is obtained from a group of genes of a gene family (the gene tree). Goodman et al. deduced that this inconsistency might be the result of mistaking paralogs for orthologs. Orthologous genes refer to copies of genes that reveal the phylogeny of species, while paralogous genes have been created by duplication events. Phylogeny reconstruction can help to understand how gene families evolved and to identify the chronology of duplications within a gene family of a single species. Several software tools, including GeneTree, DupTree, NOTUNG, and AUGIST have been developed for this task. There is, however, lack of both test data and evaluation procedures to test, compare, and benchmark their performance and results. Results We present here a simulation environment designed to generate large gene families with complex duplication histories on which reconstruction algorithms can be tested and software tools can be benchmarked. The simulation of gene family histories starts with the generation of species trees. Within these rooted bifurcating trees the nodes represent species and edges their relation. Specifically, internal nodes represent ancient species whereas leaf nodes represent extant species. Given a number of species N, we generate a random tree T under the Age Model described in Keller-Schmidt et al. This model starts with a rooted tree with two leaves. In an iterative process one of the leaves is selected and two new leaves are attached to it until the tree has N leaves. This model makes use of the idea that the longer a leaf has not been involved in a speciation, the less likely it will be in the future. The user will introduce n number of genes (gene families), which will be placed at the root of the generated species tree T. T will then be traversed in a depth first order. For each visited edge a number of events is sampled from a stochastic Poisson Process Pλ ,l where λ is the probability of the event to happen and l the branch length. The process may generate none, one or a series of these events: one gene gets duplicated (gene duplication), a group of genes gets duplicated (cluster duplication), the whole group of genes gets duplicated (genome duplication) and one gene of the species gets lost (gene loss). After each gene duplication, one of the copies will be lost with a user defined probability θ, based on the fact that when there is a gene duplication, one of the copies might be lost or become nonfunctional. In the case of a cluster or genome duplication, we apply this probability to every gene in the group, since it is known that in the wake of multiple gene duplications and in particular for genome duplications we have to expect that many duplicated genes are rapidly lost again through the formation of pseudogenes. A small example of a gene family history generated by our simulation is shown in Fig. 1. We also show the gene tree generated from the gene family history embedded in the species tree. Each leaf node represents a gene and each internal node represents an event (speciation or duplication). This tree is typically depicted as the reconciled tree as in Fig. 2. Finally, the algorithm will generate one gene tree for each species, i.e. the pruned reconciled tree containing only genes of a certain species. Furthermore, for each gene family the orthology and homology matrices are computed. To generate the orthology matrix, we say that two genes are orthologous if their lowest common
- Research Article
126
- 10.1006/mpev.1997.0404
- Jun 1, 1997
- Molecular Phylogenetics and Evolution
Correlation of Functional Domains and Rates of Nucleotide Substitution in Cytochromeb