Using ggtree to Visualize Data on Tree-Like Structures.
Ggtree is an R/Bioconductor package for visualizing tree-like structures and associated data. After 5 years of continual development, ggtree has been evolved as a package suite that contains treeio for tree data input and output, tidytree for tree data manipulation, and ggtree for tree data visualization. Ggtree was originally designed to work with phylogenetic trees, and has been expanded to support other tree-like structures, which extends the application of ggtree to present tree data in other disciplines. This article contains five basic protocols describing how to visualize trees using the grammar of graphics syntax, how to visualize hierarchical clustering results with associated data, how to estimate bootstrap values and visualize the values on the tree, how to estimate continuous and discrete ancestral traits and visualize ancestral states on the tree, and how to visualize a multiple sequence alignment with a phylogenetic tree. The ggtree package is freely available at https://www.bioconductor.org/packages/ggtree. © 2020 by John Wiley & Sons, Inc. Basic Protocol 1: Using grammar of graphics for visualizing trees Basic Protocol 2: Visualizing hierarchical clustering using ggtree Basic Protocol 3: Visualizing bootstrap values as symbolic points Basic Protocol 4: Visualizing ancestral status Basic Protocol 5: Visualizing a multiple sequence alignment with a phylogenetic tree.
- Research Article
12
- 10.1002/cpz1.70016
- Oct 30, 2024
- Current protocols
Multiple sequence alignments and phylogenetic trees are rich in biological information and are fundamental to research in biology. PhyKIT is a tool for processing and analyzing the information content of multiple sequence alignments and phylogenetic trees. Here, we describe how to use PhyKIT for diverse analyses, including (i) constructing a phylogenomic supermatrix, (ii) detecting errors in orthology inference, (iii) quantifying biases in phylogenomic data sets, (iv) identifying radiation events or lack of resolution using gene support frequencies, and (v) conducting evolution-based screens to facilitate gene function prediction. Several PhyKIT functions that streamline multiple sequence alignment and phylogenetic processing-such as renaming FASTA entries or tree tips-are also discussed. These protocols demonstrate how simple command-line operations in the unified framework of PhyKIT facilitate diverse phylogenomic data analysis and processing, from supermatrix construction and diagnosis to gaining clues about gene function. © 2024 The Author(s). Current Protocols published by Wiley Periodicals LLC. Basic Protocol 1: Installing PhyKIT and syntax for usage Basic Protocol 2: Constructing a phylogenomic supermatrix Basic Protocol 3: Detecting anomalies in orthology relationships Basic Protocol 4: Quantifying biases in phylogenomic data matrices and related measures Basic Protocol 5: Identifying polytomies Basic Protocol 6: Assessing gene-gene coevolution as a genetic screen.
- Research Article
- 10.1002/cpz1.671
- Feb 1, 2023
- Current Protocols
Gene-centric analysis is commonly used to chart the structure, function, and activity of microbial communities in natural and engineered environments. A common approach is to create custom ad hoc reference marker gene sets, but these come with the typical disadvantages of inaccuracy and limited utility beyond assigning query sequences taxonomic labels. The Tree-based Sensitive and Accurate Phylogenetic Profiler (TreeSAPP) software package standardizes analysis of phylogenetic and functional marker genes and improves predictive performance using a classification algorithm that leverages information-rich reference packages consisting of a multiple sequence alignment, a profile hidden Markov model, taxonomic lineage information, and a phylogenetic tree. Here, we provide a set of protocols that link the various analysis modules in TreeSAPP into a coherent process that both informs and directs the user experience. This workflow, initiated from a collection of candidate reference sequences, progresses through construction and refinement of a reference package to marker identification and normalized relative abundance calculations for homologous sequences in metagenomic and metatranscriptomic datasets. The alpha subunit of methyl-coenzyme M reductase (McrA) involved in biological methane cycling is presented as a use case given its dual role as a phylogenetic and functional marker gene driving an ecologically relevant process. These protocols fill several gaps in prior TreeSAPP documentation and provide best practices for reference package construction and refinement, including manual curation steps from trusted sources in support of reproducible gene-centric analysis. © 2023 The Authors. Current Protocols published by Wiley Periodicals LLC. Basic Protocol 1: Creating reference packages Support Protocol 1: Installing TreeSAPP Support Protocol 2: Annotating traits within a phylogenetic context Basic Protocol 2: Updating reference packages Basic Protocol 3: Calculating relative abundance of genes in metagenomic and metatranscriptomic datasets.
- Research Article
8
- 10.1007/s10682-019-10011-6
- Oct 8, 2019
- Evolutionary Ecology
One major challenge of using the phylogenetic comparative method (PCM) is the analysis of the evolution of interrelated continuous and discrete traits in a single multivariate statistical framework. In addition, more intricate parameters such as branch-specific directional selection have rarely been integrated into such multivariate PCM frameworks. Here, originally motivated to analyze the complex evolutionary trajectories of group size (continuous variable) and social systems (discrete variable) in African subterranean rodents, we develop a flexible approach using approximate Bayesian computation (ABC). Specifically, our multivariate ABC-PCM method allows the user to flexibly model an underlying latent evolutionary function between continuous and discrete traits. The ABC-PCM also simultaneously incorporates complex evolutionary parameters such as branch-specific selection. This study highlights the flexibility of ABC-PCMs in analyzing the evolution of phenotypic traits interrelated in a complex manner.
- Research Article
350
- 10.1099/00207713-51-3-731
- May 1, 2001
- International Journal of Systematic and Evolutionary Microbiology
The Win95/98/NT program FreeTree for computation of distance matrices and construction of phylogenetic or phenetic trees on the basis of random amplified polymorphic DNA (RAPD), RFLP and allozyme data is presented. In contrast to other similar software, the program FreeTree (available at http://www.natur.cuni.cz/~flegr/programs/freetree or http://ijs.sgmjournals.org/content/vol51/issue3/) can also assess the robustness of the tree topology by bootstrap, jackknife or operational taxonomic unit-jackknife analysis. Moreover, the program can be also used for the analysis of data obtained in several independent experiments performed with non-identical subsets of taxa. The function of the program was demonstrated by an analysis of RAPD data from 42 strains of 10 species of trichomonads. On the phylogenetic tree constructed using FreeTree, the high bootstrap values and short terminal branches for the Tritrichomonas foetus/suis 14-strain branch suggested relatively recent and probably clonal radiation of this species. At the same time, the relatively lower bootstrap values and long terminal branches for the Trichomonas vaginalis 20-strain branch suggested more ancient radiation of this species and the possible existence of genetic recombination (sexual reproduction) in this human pathogen. The low bootstrap values and the star-like topology of the whole Trichomonadidae tree confirm that the RAPD method is not suitable for phylogenetic analysis of protozoa at the level of higher taxa. It is proposed that the repeated bootstrap analysis should be an obligatory part of any RAPD study. It makes it possible to assess the reliability of the tree obtained and to adjust the amount of collected data (the number of random primers) to the amount of phylogenetic signals in the RAPD data of the taxon analysed. The FreeTree program makes such analysis possible.
- Research Article
24
- 10.1002/cpz1.30
- Feb 1, 2021
- Current Protocols
Protein evolution and protein engineering techniques are of great interest in basic science and industrial applications such as pharmacology, medicine, or biotechnology. Ancestral sequence reconstruction (ASR) is a powerful technique for probing evolutionary relationships and engineering robust proteins with good thermostability and broad substrate specificity. The following protocol describes the setting up and execution of an automated FireProtASR workflow using a dedicated web site. The service allows for inference of ancestral proteins automatically, from a single protein sequence. Once a protein sequence is submitted, the server will build a dataset of homology sequences, perform a multiple sequence alignment (MSA), build a phylogenetic tree, and reconstruct ancestral nodes. The protocol is also highly flexible and allows for multiple forms of input, advanced settings, and the ability to start jobs from: (i) a single sequence, (ii) a set of homologous sequences, (iii) an MSA, and (iv) a phylogenetic tree. This approach automates all necessary steps and offers a way for novices with limited exposure to ASR techniques to improve the properties of a protein of interest. The technique can even be used to introduce catalytic promiscuity into an enzyme. A web server for accessing the fully automated workflow is freely accessible at https://loschmidt.chemi.muni.cz/fireprotasr/. © 2021 Wiley Periodicals LLC. Basic Protocol: ASR using the Web Server FireProtASR.
- Research Article
16
- 10.1002/cpns.104
- Sep 27, 2020
- Current Protocols in Neuroscience
MagellanMapper is a software suite designed for visual inspection and end-to-end automated processing of large-volume, 3D brain imaging datasets in a memory-efficient manner. The rapidly growing number of large-volume, high-resolution datasets necessitates visualization of raw data at both macro- and microscopic levels to assess the quality of data, as well as automated processing to quantify data in an unbiased manner for comparison across a large number of samples. To facilitate these analyses, MagellanMapper provides both a graphical user interface for manual inspection and a command-line interface for automated image processing. At the macroscopic level, the graphical interface allows researchers to view full volumetric images simultaneously in each dimension and to annotate anatomical label placements. At the microscopic level, researchers can inspect regions of interest at high resolution to build ground truth data of cellular locations such as nuclei positions. Using the command-line interface, researchers can automate cell detection across volumetric images, refine anatomical atlas labels to fit underlying histology, register these atlases to sample images, and perform statistical analyses by anatomical region. MagellanMapper leverages established open-source computer vision libraries and is itself open source and freely available for download and extension. © 2020 Wiley Periodicals LLC. Basic Protocol 1: MagellanMapper installation Alternate Protocol: Alternative methods for MagellanMapper installation Basic Protocol 2: Import image files into MagellanMapper Basic Protocol 3: Region of interest visualization and annotation Basic Protocol 4: Explore an atlas along all three dimensions and register to a sample brain Basic Protocol 5: Automated 3D anatomical atlas construction Basic Protocol 6: Whole-tissue cell detection and quantification by anatomical label Support Protocol: Import a tiled microscopy image in proprietary format into MagellanMapper.
- Peer Review Report
- 10.7554/elife.82538.sa0
- Oct 31, 2022
By producing one hundred SARS-CoV-2 phylogenetic trees for different geographical and time scales, the maximum likelihood-based phylodynamic method implemented here enabled to robustly describe SARS-CoV-2 geographic spread through France, Europe, and worldwide in 2020.
- Peer Review Report
- 10.7554/elife.82538.sa2
- Mar 5, 2023
Full text Figures and data Side by side Abstract Editor's evaluation Introduction Results Discussion Materials and methods Data availability References Decision letter Author response Article and author information Metrics Abstract Although France was one of the most affected European countries by the COVID-19 pandemic in 2020, the dynamics of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) movement within France, but also involving France in Europe and in the world, remain only partially characterized in this timeframe. Here, we analyzed GISAID deposited sequences from January 1 to December 31, 2020 (n = 638,706 sequences at the time of writing). To tackle the challenging number of sequences without the bias of analyzing a single subsample of sequences, we produced 100 subsamples of sequences and related phylogenetic trees from the whole dataset for different geographic scales (worldwide, European countries, and French administrative regions) and time periods (from January 1 to July 25, 2020, and from July 26 to December 31, 2020). We applied a maximum likelihood discrete trait phylogeographic method to date exchange events (i.e., a transition from one location to another one), to estimate the geographic spread of SARS-CoV-2 transmissions and lineages into, from and within France, Europe, and the world. The results unraveled two different patterns of exchange events between the first and second half of 2020. Throughout the year, Europe was systematically associated with most of the intercontinental exchanges. SARS-CoV-2 was mainly introduced into France from North America and Europe (mostly by Italy, Spain, the United Kingdom, Belgium, and Germany) during the first European epidemic wave. During the second wave, exchange events were limited to neighboring countries without strong intercontinental movement, but Russia widely exported the virus into Europe during the summer of 2020. France mostly exported B.1 and B.1.160 lineages, respectively, during the first and second European epidemic waves. At the level of French administrative regions, the Paris area was the main exporter during the first wave. But, for the second epidemic wave, it equally contributed to virus spread with Lyon area, the second most populated urban area after Paris in France. The main circulating lineages were similarly distributed among the French regions. To conclude, by enabling the inclusion of tens of thousands of viral sequences, this original phylodynamic method enabled us to robustly describe SARS-CoV-2 geographic spread through France, Europe, and worldwide in 2020. Editor's evaluation This paper is a comprehensive, quantitative, and robust overview of the global, European, and French genomic epidemiology of SARS-CoV-2 in the first year of the pandemic. It contributes methodological advances in maximum likelihood phylogeography, using multiple scales and providing a simulation-based validation. The results show two distinct patterns of SARS-CoV-2 exchange events between the first and second half of 2020, with Europe being involved in most intercontinental exchanges: France experienced viral introductions primarily from North America and Europe during the first wave, while the second wave saw limited intercontinental movement and a significant contribution of the virus from Russia into Europe. https://doi.org/10.7554/eLife.82538.sa0 Decision letter Reviews on Sciety eLife's review process Introduction On December 1, 2019, an outbreak of severe respiratory disease was identified in the city of Wuhan, China (Huang et al., 2020). The severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) was rapidly identified as the agent of the disease (Zhu et al., 2020), responsible for the ongoing global pandemic of coronavirus disease 2019 (COVID-19). By the end of 2020, the virus caused over 1.8 million deaths worldwide including ~65,000 deaths in France, concomitantly with social and economic devastations in many regions of the world (Mofijur et al., 2021; Santomauro et al., 2021). Since the beginning of COVID-19 pandemic, the scientific community has thoroughly characterized the virus, including its pathogenesis, the monitoring of its circulation in human populations, and the development of several treatments or vaccines (Cevik et al., 2020; Krammer, 2020). Epidemiological models have been particularly helpful to quantify viral spread both in the short and long terms and to inform public health decisions (Hoertel et al., 2020; Kissler et al., 2020). In addition to clinical and epidemiological insights, viral whole-genome sequencing has become a powerful and invaluable tool to better understand infection dynamics (Volz et al., 2013), including the COVID-19 pandemic. The number of available SARS-CoV-2 whole-genome sequences has rapidly grown thanks to the efforts of scientists and researchers gathered via international networks such as the Global Initiative on Sharing All Influenza Data, GISAID (https://www.gisaid.org/; Khare et al., 2021). These genomic sequences are essential to effectively reconstruct the global viral spread and the origins of variants. Genomic data have become a strong asset in addition to epidemiological data to inform governments and help public health decisions (Attwood et al., 2022; Rife et al., 2017). However, due to the computational time required for many analyses, existing phylogenetic tools are limited for studying large amounts of data such as those generated by widespread viral sequencing. Therefore, it is still necessary to develop methods to analyze large datasets while optimizing computational calculation times. Producing appropriate subsamples through several replicates may be an efficient approach in this matter. In France, the first COVID-19 suspected case was identified in late December 2019 (Deslandes et al., 2020), and the first confirmed cases of SARS-CoV-2 infection were detected on January 24, 2020, in individuals who had recently traveled in China (Bernard Stoecklin et al., 2020). COVID-19 cases remained scarce until the end of February, when the national incidence curve of new SARS-CoV-2 infections started to rise (Figure 1). By the end of February, reinforced measures were announced, including social distancing, cessation of passenger flights to France, school closure, and finally, a complete lockdown across the entire country from March 17 to May 10, 2020. The reported daily incidence and numbers of severe cases peaked at the beginning of April 2020 before decreasing steadily until August 2020. However, after the relaxation of social distancing measures in June, a second wave of infections occurred in early September peaking at more than 100,000 positive cases and 1300 confirmed deaths in a single day on November 2, 2020 (Figure 1). After this peak, daily incidence and severe COVID-19 cases gradually diminished down to a number of positive daily cases varying between 2000 and 25,000 at the end of 2020 thanks to a second national lockdown applied between October 29 and December 15, 2020. Epidemiological trends were similar in most European countries except for Russia or Romania, where high rates of SARS-CoV-2-related deaths were reported even in the summer of 2020. Of note, the other continents showed different patterns of virus circulation: compared to Europe, the number of deaths increased about 2 weeks later in North America and remained high throughout 2020; and from early May, Asia and South America were also highly impacted by the pandemic (Figure 1—figure supplement 1). Figure 1 with 2 supplements see all Download asset Open asset Timeline of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)-related deaths and stringency index in France, 2020. Key events are indicated on the timeline. Official lockdowns included stay home orders and closure of schools and daycares. Based on SARS-CoV-2-related deaths, the two first French epidemic waves are, respectively, dated from March to July 2020, and September to December 2020. SARS-CoV-2-related deaths are displayed as the daily number of deaths (light blue area) and as the weekly average of daily number of deaths (dark blue curve). The stringency (Oxford) index is a composite measure based on different response indicators including school and workplace closures and travel bans, rescaled to a value from 0 to 100 (100 = strictest) (Hale et al., 2021). Elucidating the SARS-CoV-2 dynamic throughout the various phases of the pandemic is paramount to better anticipate how to limit virus circulation for future viral epidemics (Rife et al., 2017). Here, we analyzed GISAID deposited sequences to elucidate the origins and spread of the virus in France, Europe, and the world from January 1 to December 31, 2020. Through a maximum likelihood discrete trait phylogeographic method, we estimated the main geographical areas that contributed to viral introduction into France and Europe, the countries/continents to which France exported SARS-CoV-2 the most and the contribution of the different French regions to the national circulation of the virus. The main exchanged lineages were also investigated. We looked at the differences in virus circulation during each of the two European epidemic waves of 2020 independently. Given France's central geographic location in Europe and the high proportion of international travelers visiting this country before the pandemic, we aimed to explore the role that France played in SARS-CoV-2 exchanges both in Europe and worldwide. Results Defining appropriate subsamples using simulations From January 1 to December 31, 2020, a total of 638,706 sequences were retained in our study. Inferring a phylogenic tree with such a large number of sequences would require very long calculation times. To overcome this limit, we constructed smaller datasets by randomly choosing subsamples (with replacement) of the sequences. The number of sequences for each country at each week was chosen to be proportional to the number of SARS-CoV-2-related deaths per country and per week with a 2-week shift to account for the time between infection and death. As a proof of concept, we conducted an extensive simulation study to estimate the accuracy of the discrete trait phylogeographic inference for rates of transitions between two distinct locations. First, to evaluate the precision of such inference on a tree of 1000 leaves, we simulated a two-states model with different combinations of transition rates in 50 replicates. Parameters were correctly estimated with limited variability across the 50 replicates. The median parameters across replicates gave a very accurate estimation (Figure 2A). Figure 2 with 4 supplements see all Download asset Open asset Estimating variability in transition rates using simulations. (A) Estimated versus true parameters in the simulation study of the two-states model. The two panels show the two transition rates. For each set of parameters, 50 replicates were conducted. The large red dot is the median of the replicates. The red cross is the true parameter value, on the bisector. (B) Estimated rate of transition in subsampled trees. For each replicate (n = 50), one point is the result of one subsampled smaller phylogenetic tree (from a large phylogenetic tree). The big dot shows the median for each replicate. The horizontal red line is the overall median (of the medians), across replicates. The horizontal dashed gray line is the true rate. Only one of the two rate parameters is shown. (C) Log-median error in parameter estimation as a function of the log number of replicates, when inference is conducted on truly independent replicate evolutionary histories, on a tree of 1000 leaves. The points are the data, the dashed line shows the line of slope '−1' which is the expectation as the replicates are truly independent. (D) Log-median error as a function of log number of subsamples used for the inference done on subsampled phylogenetic trees. The colored points and lines show the inference done on 50 distinct realizations of the evolutionary process on the whole tree. The dashed line is the overall regression line with a slope of −0.7. We tested between 1 and 10 subsampled trees (x-axis). Next, to evaluate how independent parameter estimates are done on randomly subsampled trees of the same larger phylogeny, we inferred parameters on 50 100-leaves trees randomly subsampled from a 10,000-leaves SARS-CoV-2 phylogenetic tree. For each resulting subtree, we conducted inferences on 50 replicates corresponding to 50 realizations of the stochastic process of evolution of the discrete character – as done in the first simulation – on the whole tree of 10,000 leaves. For each replicate, we observed some error on the estimation of the parameter, because one replicate only corresponds to one possible realization of the evolutionary process, although the overall median of inferred parameters across subsampled trees was closer to the true parameter values (Figure 2B). Different estimations of the transition rates conducted on different subsampled trees are not expected to be fully independent because the subtrees partly share the same evolutionary history. Therefore, we estimated the level of independence of these estimations. When several estimates are perfectly independent from one another and are averaged to obtain the final estimate of the quantity of interest, we expect the error in parameter estimation to converge to 0 with a 1/N (N−1) scaling, where N is the number of replicates. This is indeed what we observed when we calculated the error on estimation of the parameter as a function of the chosen number of replicates N in the first set of simulations. Here, the replicates were truly independent replicate realizations of the evolutionary history and inference was conducted on the whole tree of 1000 leaves (Figure 2C). On the contrary, when estimates are perfectly dependent, error on the averaged parameter estimate is expected to not decrease with N. When evaluating the error on parameter estimates across subsamples of the large tree, we expected the scaling of error as a function of number of subsamples N to be intermediate between non-independence (~N0 scaling) and perfect independence (~N−1 scaling). Using the relationship between log(error) as a function of log(N), we estimated a slope of −0.7 (Figure 2D). Thus, inferences conducted on subsamples of the same phylogenetic tree are partly independent. The precise degree of independence is expected to depend on the shape of the phylogenetic tree, but the coefficient was similar when doing the same study on a randomly generated tree instead of the SARS-CoV-2 tree. We finally conducted another round of simulations to evaluate the error on what we considered as exchange between multiple locations when using sparse subsampling. For that, a 1,000,000-leaves tree was simulated with a five-states discrete trait representing geographical units. Then, 100 subsampled 1000-leaves trees from the whole phylogenetic tree were produced and the ancestry for the discrete trait was reconstructed from the leaf data only. We estimated the number of transitions (exchanges) of each type and compared them with the one obtained from the main tree, finding a mean error rate of 2.7% over the 100 subsamples (Figure 2—figure supplement 1). Altogether, these simulations suggested that using subsamples of 1000 sequences from a large dataset and performing partially independent replicates seems to be sufficient to accurately estimate transition events. Description of the datasets and global diversity of SARS-CoV-2 sequences We defined 100 subsamples of sequences proportionally to COVID-19 deaths across geographic locations and time for different geographic scales (worldwide, Europe, and French regions) and time periods (from January 1 to July 25, 2020, and from July 26 to December 31, 2020, respectively, covering the first and second European epidemic waves). We chose the sampling intensity guided by the weekly number of SARS-CoV-2-related deaths reported by public health organizations. Here, the number of SARS-CoV-2-related deaths was used rather than the number of detected cases because the latter was biased due to variable ascertainment rates across countries and time. For example, the larger number of PCR tests conducted in the second epidemic wave could wrongly suggest that the virus circulated much more during the second half of 2020 (Figure 1—figure supplement 2). For each geographic scale and time period, there was a positive correlation between the weekly number of SARS-CoV-2-related deaths and the weekly number of sequences we included for a subsample (Spearman's rank correlation, p < 0.001; r = 0.94 for the lowest correlation). We also confirmed that the number of sequences per territory was, on average, properly temporally distributed within each time period (Figure 2—figure supplement 2). Some countries and French administrative regions were however discarded in the analyses because they were not sufficiently represented in the GISAID database. Overall, a total of 39,288 and 39,755 distinct SARS-CoV-2 sequences were included across the 100 sampled phylogenies for the worldwide dataset, respectively, for the first and the second time periods (Table 1). At the European scale, 26,757 and 27,658 different SARS-CoV-2 sequences covering 11 countries were analyzed across the 100 subsamples (Table 1). Focusing on French administrative regions, sequences available on the GISAID database were very sparse. The Provence-Alpes-Côte d'Azur (PACA, Marseille area) was the only region that highly sequenced SARS-CoV-2 in 2020. Île-de-France (IDF, Paris area), Auvergne-Rhône-Alpes (ARA, Lyon area), Occitanie (OCC, Toulouse and Montpellier area), and Bretagne (BRE, Rennes area) have sequenced much less than PACA, but provided sufficient data to investigate SARS-CoV-2 geographic exchange events in France. The remaining French administrative regions were discarded since too few sequences were available to properly match the number of weekly SARS-CoV-2 deaths (Figure 2—figure supplement 3). We thus considered 2543 unique sequences across the 100 subsamples between January 1 and July 25, 2020, and 3124 unique sequences between July 26 and December 31, 2020 (Table 1). Table 1 Number of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) sequences investigated for each dataset. DatasetGeographiesPeriod investigatedAverage number of sequences sampled in a subsampleTotal number of sequencesWorldAfrica, Asia, Europe, France, North America, Oceania, South AmericaJanuary 1 to July 25, 202084639,288July 26 to December 31, 202077739,755EuropeBelgium, France, Germany, Italy, The Netherlands, Poland, Romania, Russia, Spain, Sweden, United KingdomJanuary 1 to July 25, 202090426,757July 26 to December 31, 202087227,658FranceAuvergne-Rhône-Alpes (ARA), Bretagne (BRE), Île-de-France (IDF), Occitanie (OCC), Provence-Alpes-Côte d'Azur (PACA)January 1 to July 25, 20204162543July 26 to December 31, 20204333124 The genomic diversity of circulating SARS-CoV-2 in the different continents, countries, and French regions was found to be similar (Figure 2—figure supplement 4). Overall, genomes showed high sequence conservation compared to the Wuhan-Hu-1 reference in 2020 (mean and median of ~13 single nucleotide polymorphisms (SNPs) with 95% of the distribution comprised between 4 and 25 SNPs). Which continents exchanged SARS-CoV-2 with Europe and France? Through 100 distinct, dated and ancestrally reconstructed phylogenetic trees, we first studied SARS-CoV-2 exchanges worldwide for each of the time periods studied. Between January 1 and July 25, 2020 (covering the first European epidemic wave), we found that Europe (excluding France) accounted for 57.3% of the total number of exportation events, and was the main source of SARS-CoV-2 exportations toward the other continents in all of the subsamples (Figure 3A–D and Figure 3—figure supplement 1). North America also highly participated in virus exportation during this period time (24.3%). South America and Asia were each associated with 7.1% of the total number of exportation events, consistent with a later circulation of the virus in these continents (Figure 1—figure supplement 1). France was estimated to have contributed of the total exportation events, that France was not the European source of SARS-CoV-2 at the international level between January 1 and July 25, 2020. The exportation events from France were mostly toward Europe to a to North America, South America, and Asia (Figure and Figure 3—figure supplement 2). These events mostly of the B.1 and lineages (Figure North America a large proportion of SARS-CoV-2 from other continents of the introduction by South America and Europe (Figure average of of all SARS-CoV-2 introductions were into France, and from North America and Europe (Figure 3—figure supplement 2). These introductions of the B.1 and lineages (Figure The first introductions into France were detected at the beginning of February, and increased to a before the lockdown from March 2020 (Figure Only South America and Asia were associated with a in SARS-CoV-2 introductions after this because such measures were there and the circulation of the virus remained limited in these regions. Figure with 2 supplements see all Download asset Open asset acute respiratory syndrome coronavirus 2 (SARS-CoV-2) exchange worldwide. events were inferred with 100 subsampled phylogenies between January 1 and July 25, 2020, and between July 26 and December 31, 2020. (A) Number of introduction and exportation events for each subsample and for each and France. (B) SARS-CoV-2 exchange between continents and France during the two time periods investigated. In these of a location to the and with an at the is proportional to the exchange (C) Number of exportation and (D) introduction events per territory over time. The mean number of exchanges over the subsamples and for each week was the of the complete lockdowns in France. of lineages exported from France and introduced into France. with a proportion were into the From July 26 to December 31, 2020 European epidemic wave), we observed exchange events worldwide compared to the first half of 2020. we showed the of analyzing several as there was a large in the total number of exportation or introduction events, in Europe (Figure Europe was, as between January 1 and July 25, 2020, the main source of exchanges with a total of of the exportation events across by North America Asia and South America (Figure 3A–D and Figure 3—figure supplement 1). of the events occurred during the summer period to August 2020), corresponding to the summer in most countries of the world. France accounted for of the exportation events, but they were toward other European countries and overall detected from August to November 2020 (Figure and Figure 3—figure supplement consistent with the SARS-CoV-2 incidence in this period in France (Figure 1—figure supplement 2). The B.1.160 accounted for all the exportation events from France (Figure In a similar SARS-CoV-2 introductions into France mostly from Europe (Figure and were detected at a rate from April 2020, at a but limited rate from 2020, and at a strong level in September and October 2020 (Figure 3—figure supplement 2). These SARS-CoV-2 introductions into France in of B.1.160 B.1 and lineages (Figure the virus spread in We aimed to a more of SARS-CoV-2 exchanges between France and other European countries with the same Here, we only on European countries associated with a high incidence and without due to a of data on GISAID (Table 1). By the of introduction and exportation events between January 1 and July 25, 2020 across the we observed that was the to virus exportation toward other European countries, with an average of of the total number of exportation events. The United Kingdom, France, and also highly participated in virus and of and of the total number of exportation events, (Figure and Figure supplement 1). These are in line with epidemiological data, since was the first country in Europe to be affected by the and France, the United Kingdom, and were the other European countries associated with the number of SARS-CoV-2-related deaths during the first wave (Figure 1—figure supplement 1). The number of all exportation events however after the of lockdowns in the different countries (with the first one in on March (Figure France mostly exported SARS-CoV-2 toward and the United and a less toward and the All of these events occurred before the lockdown in France (Figure and Figure supplement and of the B.1 and lineages (Figure The rate of SARS-CoV-2 exportations from France until the second European epidemic wave, as it was also the case for other European countries except Russia (Figure For all introduction events, the were more the United accounted for a of the total number of events, while Russia, Belgium, Germany, Italy, Spain, France, and the represented between and of the total number of events (Figure In France, a high rate of introduction events was observed in and March before the lockdown and mostly from the United and (Figure supplement 2). These introductions in of the and B.1 lineages (Figure Figure 4 with 2 supplements see all Download asset Open asset acute respiratory syndrome coronavirus 2 (SARS-CoV-2) exchanges on the European events were calculated by the results from 100 subsampled phylogenies between January 1 and July 25, 2020, and between July 26 and December 31, 2020. (A) Number of introduction and exportation events for each subsample and for each European (B) SARS-CoV-2 exchange between European countries during the two time periods investigated. In these of a location to the and with an at is proportional to the exchange (C) Number of exportation and (D) introduction events per territory over time. The mean number of exchanges over subsamples and for each week was the of the complete lockdowns in France. of lineages exported from France and introduced into France. with a proportion were into the The second time period (from July 26 to December 31, showed a different of exchanges. Here, we estimated exchanges compared to the first half of 2020. Russia accounted for most of the exportation events (Figure These events were estimated to during the the relaxation of measures in most European and the summer periods (Figure This result was expected since Russia was the European country to a high number of SARS-CoV-2-related deaths during this period (Figure 1—figure supplement 1). France and the United also highly participated in virus exportation (Figure and Figure supplement 1). of these events were detected between August and October 2020 (Figure and before the second lockdown in most European countries first one in on October 2020). these are consistent with epidemiological as was the first country in European to be associated with a of SARS-CoV-2-related deaths, rapidly by France. France mostly exported the virus toward and (Figure and Figure supplement and mostly the B.1.160 (Figure Focusing on introduction events, accounted for of the total number of by the United and France For the remaining European countries, the proportion of introduction events was comprised between and
- Discussion
5
- 10.1016/s1471-4906(03)00165-0
- Jun 25, 2003
- Trends in Immunology
Response to Shields: Molecular evolution of CXC chemokines and receptors
- Research Article
- 10.24237/asj.02.02.695b
- Apr 29, 2024
- Academic Science Journal
The outbreak of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), which causes coronavirus disease 2019 (COVID-19), has spread worldwide. Therefore, this study aimed to build a phylogenetic tree of complete genomes of SARS-CoV-2 and other species of Gamma coronavirus to explore the possibility of finding the evolutionary relationships between them and wished to analyze them in order to forecast the best trees illustrating the sequences' evolutionary relationships and obtain a well-supported phylogenetic tree by using the Neighbor-Joining (NJ) and Maximum Likelihood (ML) methods after performing multiple sequence alignment (MSA). This study utilized 16 isolates of Gamma coronavirus species and SARS-CoV-2 retrieved from the NCBI (National Center for Biotechnology Information) database for this investigation. The experimental outcomes when applying the two methods to the same dataset show that a well-supported and trustworthy phylogenetic tree was obtained with a bootstrapping value of 100% for all branches of the tree when applying the ML method. Additionally, a well-supported and fast-constructing phylogenetic tree was obtained through the NJ method for all branches except one, where the bootstrapping value appeared to be 56%. The research was conducted in 2022 at the College of Science, Diyala University.
- Research Article
2
- 10.1002/cpz1.619
- Nov 1, 2022
- Current protocols
ConVarT (https://convart.org/) is a search engine for searching for conjugate variants between humans and other species. The search engine is based on matching conjugate variants called MatchVars between species. Matching equivalent variants requires correct alignment of orthologous proteins with the use of multiple sequence alignments (MSA). Indeed, the ConVarT pipeline has performed over a million MSAs and integrated variants and variant-specific annotations (pathogenicity, phenotypic variants; etc.) into the corresponding positions on MSAs. When a clinically relevant variant is discovered whose functional relevance is unknown, ConVarT offers clinician scientists the possibility to search for a MatchVar in other species and to look for functional data on that variant. Fortunately, ConVarT enables users to paste a protein sequence in FASTA format to search for human orthologous proteins. A pairwise sequence alignment (PSA) is then performed between the provided protein sequence and the human orthologous protein, allowing users to visualize human variants on the PSA. Here, we describe the step-by-step usage of ConVarT. © 2022 Wiley Periodicals LLC. Basic Protocol 1: Searching matching variants (MatchVar) with gene/protein identifiers. Basic Protocol 2: Searching with a FASTA sequence. Alternate Protocol: Search with gene name in multiple species. Basic Protocol 3: Search genes associated with a disease.
- Book Chapter
21
- 10.1007/978-3-030-19318-8_10
- Jan 1, 2019
Molecular phylogeny is used to study the relationships among the set of objects by generating phylogenetic or evolutionary tree. The objects in the study can be organisms or biomolecules such as gene or protein. The evolutionary history hidden in the biomolecules establishes the evolutionary patterns in the form of a tree when a suitable data, data substitution models, and tree construction methods are used. These evolutionary patterns are used to study the relationships among the objects. These patterns sometimes make it difficult to infer the relationship among the objects. In addition, different tree construction methods like unweighted pair group method with arithmetic mean (UPGMA), neighbor joining, minimum evolution, Fitch-Margoliash, maximum parsimony, maximum likelihood, Monte Carlo’s simulation, Bayes, and so on and types of data used in the analysis make it much more complicated to infer the relationships. The above tree construction methods follow different principles to construct a phylogenetic tree. Most often, the tree topologies generated by different methods for the same data will be the same, whereas in some cases the tree topologies may be different in their internal branching. These differences in the tree topologies may make it difficult to assess the confidence of the phylogenetic tree. Further, combination of the tree construction methods and data used by phylogeny program packages such as MEGA, Molphy, Phylip, PAML, and PAUP also make it difficult to assess the confidence of the phylogenetic tree. Molecular phylogeny has a wide range of applications such as affiliating taxonomy of an organism, studying reproductive biology in lower organisms, assessing the process of cryptic speciation in a species, understanding the history of life, resolving controversial history of life, reconstructing the paths of infection in an epidemiology, classifying proteins or genes into families, and many more. If the interpretation of the evolutionary patterns is not appropriate, then the inference of the study may be misleading. Thus, interpretation of the tree and relationships among the organisms is always dependent on assessing the confidence of the phylogenetic tree. Literature review shows that sampling methods such as bootstrapping, jackknifing, and Bayesian simulation and statistical methods such as Kishino-Hasegawa test and Shimodaira-Hasegawa test are used to assess the confidence of the phylogenetic tree. Thus, this chapter reviews the applications, construction, and assessment of phylogenetic tree.
- Research Article
103
- 10.3201/eid1911.130908
- Nov 1, 2013
- Emerging Infectious Diseases
The Portuguese Foundation for Science and Technology supported the doctoral fellowship of A.M.L. (SFRH/BD/78738/2011) and the postdoctoral fellowship of J.A. (SFRH/ BPD/73512/2010). The Portuguese Foundation for Science and Technology Projects PTDC/CVT/108490/2008 and FCT-ANR/ BIA-BIC/0043/2012 supported this work.
- Research Article
- 10.11646/zootaxa.4881.3.5
- Nov 20, 2020
- Zootaxa
A complementary description of Panonychus caricae Hatzinikolis, 1984, is presented based on the morphology of adult female and male individuals collected from fig trees (Ficus sp., Moraceae) in Greece. Morphological differences between Panonychus caricae and two closely related species, Panonychus ulmi (Koch, 1836) and Panonychus hadzhibejliae (Reck, 1947), are discussed. Panonychus caricae can be separated from two other Panonychus species using the length of the female dorsal setae in combination with the ratio between the length of female dorsal opisthosomal setae f2 and h1, and the ratio between the length of dorsal setae sc1 and h1. A phylogenetic maximum likelihood tree was constructed based on the cytochrome c oxidase subunit I (COI) gene of mitochondrial DNA (mtDNA) from 10 species of the subgenus Panonychus s.str. (including the re-described species P. caricae) and the only two species of the subgenus Sasanychus. The phylogenetic tree indicates that these 12 species are clearly separated from each other. The two subgenera, Panonychus s.str. and Sasanychus, comprise strongly supported monophyletic clades with 98% bootstrap values. The convergence of molecular and morphological data (dorsal setae set on tubercles or not, number of tactile setae on tibiae I and II, and patterns of the dorsocentral striae) suggests that Sasanychus should not be classified under the genus Panonychus. Consequently, molecular and morphological evidence supports the resurrection of the genus Sasanychus, which contains two species, S. akitanus (Ehara) and S. pusillus Ehara Gotoh, as distinct from Panonychus. A key to the world species of Panonychus and Sasanychus is also provided.
- Supplementary Content
181
- 10.1038/emboj.2008.189
- Sep 25, 2008
- The EMBO Journal
Co-evolution has an important function in the evolution of species and it is clearly manifested in certain scenarios such as host–parasite and predator–prey interactions, symbiosis and mutualism. The extrapolation of the concepts and methodologies developed for the study of species co-evolution at the molecular level has prompted the development of a variety of computational methods able to predict protein interactions through the characteristics of co-evolution. Particularly successful have been those methods that predict interactions at the genomic level based on the detection of pairs of protein families with similar evolutionary histories (similarity of phylogenetic trees: mirrortree). Future advances in this field will require a better understanding of the molecular basis of the co-evolution of protein families. Thus, it will be important to decipher the molecular mechanisms underlying the similarity observed in phylogenetic trees of interacting proteins, distinguishing direct specific molecular interactions from other general functional constraints. In particular, it will be important to separate the effects of physical interactions within protein complexes (‘co-adaptation') from other forces that, in a less specific way, can also create general patterns of co-evolution.