Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Improving the Accuracy of Demographic and Molecular Clock Model Comparison While Accommodating Phylogenetic Uncertainty

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Recent developments in marginal likelihood estimation for model selection in the field of Bayesian phylogenetics and molecular evolution have emphasized the poor performance of the harmonic mean estimator (HME). Although these studies have shown the merits of new approaches applied to standard normally distributed examples and small real-world data sets, not much is currently known concerning the performance and computational issues of these methods when fitting complex evolutionary and population genetic models to empirical real-world data sets. Further, these approaches have not yet seen widespread application in the field due to the lack of implementations of these computationally demanding techniques in commonly used phylogenetic packages. We here investigate the performance of some of these new marginal likelihood estimators, specifically, path sampling (PS) and stepping-stone (SS) sampling for comparing models of demographic change and relaxed molecular clocks, using synthetic data and real-world examples for which unexpected inferences were made using the HME. Given the drastically increased computational demands of PS and SS sampling, we also investigate a posterior simulation-based analogue of Akaike's information criterion (AIC) through Markov chain Monte Carlo (MCMC), a model comparison approach that shares with the HME the appealing feature of having a low computational overhead over the original MCMC analysis. We confirm that the HME systematically overestimates the marginal likelihood and fails to yield reliable model classification and show that the AICM performs better and may be a useful initial evaluation of model choice but that it is also, to a lesser degree, unreliable. We show that PS and SS sampling substantially outperform these estimators and adjust the conclusions made concerning previous analyses for the three real-world data sets that we reanalyzed. The methods used in this article are now available in BEAST, a powerful user-friendly software package to perform Bayesian evolutionary analyses.

Similar Papers
  • Research Article
  • Cite Count Icon 637
  • 10.1093/molbev/mss243
Accurate Model Selection of Relaxed Molecular Clocks in Bayesian Phylogenetics
  • Feb 1, 2012
  • Molecular Biology and Evolution
  • Guy Baele + 4 more

Recent implementations of path sampling (PS) and stepping-stone sampling (SS) have been shown to outperform the harmonic mean estimator (HME) and a posterior simulation-based analog of Akaike's information criterion through Markov chain Monte Carlo (AICM), in bayesian model selection of demographic and molecular clock models. Almost simultaneously, a bayesian model averaging approach was developed that avoids conditioning on a single model but averages over a set of relaxed clock models. This approach returns estimates of the posterior probability of each clock model through which one can estimate the Bayes factor in favor of the maximum a posteriori (MAP) clock model; however, this Bayes factor estimate may suffer when the posterior probability of the MAP model approaches 1. Here, we compare these two recent developments with the HME, stabilized/smoothed HME (sHME), and AICM, using both synthetic and empirical data. Our comparison shows reassuringly that MAP identification and its Bayes factor provide similar performance to PS and SS and that these approaches considerably outperform HME, sHME, and AICM in selecting the correct underlying clock model. We also illustrate the importance of using proper priors on a large set of empirical data sets.

  • Research Article
  • Cite Count Icon 36
  • 10.1371/journal.pone.0106990
Marginal likelihood estimate comparisons to obtain optimal species delimitations in Silene sect. Cryptoneurae (Caryophyllaceae).
  • Sep 12, 2014
  • PLoS ONE
  • Zeynep Aydin + 3 more

Coalescent-based inference of phylogenetic relationships among species takes into account gene tree incongruence due to incomplete lineage sorting, but for such methods to make sense species have to be correctly delimited. Because alternative assignments of individuals to species result in different parametric models, model selection methods can be applied to optimise model of species classification. In a Bayesian framework, Bayes factors (BF), based on marginal likelihood estimates, can be used to test a range of possible classifications for the group under study. Here, we explore BF and the Akaike Information Criterion (AIC) to discriminate between different species classifications in the flowering plant lineage Silene sect. Cryptoneurae (Caryophyllaceae). We estimated marginal likelihoods for different species classification models via the Path Sampling (PS), Stepping Stone sampling (SS), and Harmonic Mean Estimator (HME) methods implemented in BEAST. To select among alternative species classification models a posterior simulation-based analog of the AIC through Markov chain Monte Carlo analysis (AICM) was also performed. The results are compared to outcomes from the software BP&P. Our results agree with another recent study that marginal likelihood estimates from PS and SS methods are useful for comparing different species classifications, and strongly support the recognition of the newly described species S. ertekinii.

  • Research Article
  • Cite Count Icon 139
  • 10.1093/sysbio/syv083
Genealogical Working Distributions for Bayesian Model Testing with Phylogenetic Uncertainty.
  • Nov 1, 2015
  • Systematic Biology
  • Guy Baele + 2 more

Marginal likelihood estimates to compare models using Bayes factors frequently accompany Bayesian phylogenetic inference. Approaches to estimate marginal likelihoods have garnered increased attention over the past decade. In particular, the introduction of path sampling (PS) and stepping-stone sampling (SS) into Bayesian phylogenetics has tremendously improved the accuracy of model selection. These sampling techniques are now used to evaluate complex evolutionary and population genetic models on empirical data sets, but considerable computational demands hamper their widespread adoption. Further, when very diffuse, but proper priors are specified for model parameters, numerical issues complicate the exploration of the priors, a necessary step in marginal likelihood estimation using PS or SS. To avoid such instabilities, generalized SS (GSS) has recently been proposed, introducing the concept of "working distributions" to facilitate--or shorten--the integration process that underlies marginal likelihood estimation. However, the need to fix the tree topology currently limits GSS in a coalescent-based framework. Here, we extend GSS by relaxing the fixed underlying tree topology assumption. To this purpose, we introduce a "working" distribution on the space of genealogies, which enables estimating marginal likelihoods while accommodating phylogenetic uncertainty. We propose two different "working" distributions that help GSS to outperform PS and SS in terms of accuracy when comparing demographic and evolutionary models applied to synthetic data and real-world examples. Further, we show that the use of very diffuse priors can lead to a considerable overestimation in marginal likelihood when using PS and SS, while still retrieving the correct marginal likelihood using both GSS approaches. The methods used in this article are available in BEAST, a powerful user-friendly software package to perform Bayesian evolutionary analyses.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 124
  • 10.1186/1471-2105-14-85
Make the most of your samples: Bayes factor estimators for high-dimensional models of sequence evolution
  • Mar 6, 2013
  • BMC Bioinformatics
  • Guy Baele + 2 more

BackgroundAccurate model comparison requires extensive computation times, especially for parameter-rich models of sequence evolution. In the Bayesian framework, model selection is typically performed through the evaluation of a Bayes factor, the ratio of two marginal likelihoods (one for each model). Recently introduced techniques to estimate (log) marginal likelihoods, such as path sampling and stepping-stone sampling, offer increased accuracy over the traditional harmonic mean estimator at an increased computational cost. Most often, each model’s marginal likelihood will be estimated individually, which leads the resulting Bayes factor to suffer from errors associated with each of these independent estimation processes.ResultsWe here assess the original ‘model-switch’ path sampling approach for direct Bayes factor estimation in phylogenetics, as well as an extension that uses more samples, to construct a direct path between two competing models, thereby eliminating the need to calculate each model’s marginal likelihood independently. Further, we provide a competing Bayes factor estimator using an adaptation of the recently introduced stepping-stone sampling algorithm and set out to determine appropriate settings for accurately calculating such Bayes factors, with context-dependent evolutionary models as an example. While we show that modest efforts are required to roughly identify the increase in model fit, only drastically increased computation times ensure the accuracy needed to detect more subtle details of the evolutionary process.ConclusionsWe show that our adaptation of stepping-stone sampling for direct Bayes factor calculation outperforms the original path sampling approach as well as an extension that exploits more samples. Our proposed approach for Bayes factor estimation also has preferable statistical properties over the use of individual marginal likelihood estimates for both models under comparison. Assuming a sigmoid function to determine the path between two competing models, we provide evidence that a single well-chosen sigmoid shape value requires less computational efforts in order to approximate the true value of the (log) Bayes factor compared to the original approach. We show that the (log) Bayes factors calculated using path sampling and stepping-stone sampling differ drastically from those estimated using either of the harmonic mean estimators, supporting earlier claims that the latter systematically overestimate the performance of high-dimensional models, which we show can lead to erroneous conclusions. Based on our results, we argue that highly accurate estimation of differences in model fit for high-dimensional models requires much more computational effort than suggested in recent studies on marginal likelihood estimation.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 7
  • 10.3390/w11081579
Making Steppingstones out of Stumbling Blocks: A Bayesian Model Evidence Estimator with Application to Groundwater Transport Model Selection
  • Jul 30, 2019
  • Water
  • Ahmed S Elshall + 1 more

Bayesian model evidence (BME) is a measure of the average fit of a model to observation data given all the parameter values that the model can assume. By accounting for the trade-off between goodness-of-fit and model complexity, BME is used for model selection and model averaging purposes. For strict Bayesian computation, the theoretically unbiased Monte Carlo based numerical estimators are preferred over semi-analytical solutions. This study examines five BME numerical estimators and asks how accurate estimation of the BME is important for penalizing model complexity. The limiting cases for numerical BME estimators are the prior sampling arithmetic mean estimator (AM) and the posterior sampling harmonic mean (HM) estimator, which are straightforward to implement, yet they result in underestimation and overestimation, respectively. We also consider the path sampling methods of thermodynamic integration (TI) and steppingstone sampling (SS) that sample multiple intermediate distributions that link the prior and the posterior. Although TI and SS are theoretically unbiased estimators, they could have a bias in practice arising from numerical implementation. For example, sampling errors of some intermediate distributions can introduce bias. We propose a variant of SS, namely the multiple one-steppingstone sampling (MOSS) that is less sensitive to sampling errors. We evaluate these five estimators using a groundwater transport model selection problem. SS and MOSS give the least biased BME estimation at an efficient computational cost. If the estimated BME has a bias that covariates with the true BME, this would not be a problem because we are interested in BME ratios and not their absolute values. On the contrary, the results show that BME estimation bias can be a function of model complexity. Thus, biased BME estimation results in inaccurate penalization of more complex models, which changes the model ranking. This was less observed with SS and MOSS as with the three other methods.

  • Research Article
  • Cite Count Icon 668
  • 10.1534/genetics.109.112532
Unified Framework to Evaluate Panmixia and Migration Direction Among Multiple Sampling Locations
  • May 1, 2010
  • Genetics
  • Peter Beerli + 1 more

For many biological investigations, groups of individuals are genetically sampled from several geographic locations. These sampling locations often do not reflect the genetic population structure. We describe a framework using marginal likelihoods to compare and order structured population models, such as testing whether the sampling locations belong to the same randomly mating population or comparing unidirectional and multidirectional gene flow models. In the context of inferences employing Markov chain Monte Carlo methods, the accuracy of the marginal likelihoods depends heavily on the approximation method used to calculate the marginal likelihood. Two methods, modified thermodynamic integration and a stabilized harmonic mean estimator, are compared. With finite Markov chain Monte Carlo run lengths, the harmonic mean estimator may not be consistent. Thermodynamic integration, in contrast, delivers considerably better estimates of the marginal likelihood. The choice of prior distributions does not influence the order and choice of the better models when the marginal likelihood is estimated using thermodynamic integration, whereas with the harmonic mean estimator the influence of the prior is pronounced and the order of the models changes. The approximation of marginal likelihood using thermodynamic integration in MIGRATE allows the evaluation of complex population genetic models, not only of whether sampling locations belong to a single panmictic population, but also of competing complex structured population models.

  • Research Article
  • Cite Count Icon 96
  • 10.1093/bioinformatics/btt340
Bayesian evolutionary model testing in the phylogenomics era: matching model complexity with computational efficiency
  • Jun 12, 2013
  • Bioinformatics
  • Guy Baele + 1 more

The advent of new sequencing technologies has led to increasing amounts of data being available to perform phylogenetic analyses, with genomic data giving rise to the field of phylogenomics. High-performance computing is becoming an indispensable research tool to fit complex evolutionary models, which take into account specific genomic properties, to large datasets. Here, we perform an extensive Bayesian phylogenetic model selection study, comparing codon and nucleotide substitution models, including codon position partitioning for nucleotide data as well gene-specific substitution models for both data types. For the best fitting partitioned models, we also compare independent partitioning with standard diffuse prior specification to conditional partitioning via hierarchical prior specification. To compare the different models, we use state-of-the-art marginal likelihood estimation techniques, including path sampling and stepping-stone sampling. We show that a full codon model best describes the features of a whole mitochondrial genome dataset, consisting of 12 protein-coding genes, but only when each gene is allowed to evolve under a separate codon model. However, when using hierarchical prior specification for the partition-specific parameters instead of independent diffuse priors, codon position partitioned nucleotide models can still outperform standard codon models. We demonstrate the feasibility of fitting such a combination of complex models using the BEAGLE library for BEAST in combination with recent graphics cards. We argue that development and use of such models needs to be accompanied by state-of-the-art marginal likelihood estimators because the more traditional and computationally less demanding estimators do not offer adequate accuracy.

  • Research Article
  • Cite Count Icon 13
  • 10.1089/cmb.2010.0139
Improved Harmonic Mean Estimator for Phylogenetic Model Evidence
  • Mar 13, 2012
  • Journal of Computational Biology
  • Serena Arima + 1 more

Bayesian phylogenetic methods are generating noticeable enthusiasm in the field of molecular systematics. Many phylogenetic models are often at stake, and different approaches are used to compare them within a Bayesian framework. The Bayes factor, defined as the ratio of the marginal likelihoods of two competing models, plays a key role in Bayesian model selection. We focus on an alternative estimator of the marginal likelihood whose computation is still a challenging problem. Several computational solutions have been proposed, none of which can be considered outperforming the others simultaneously in terms of simplicity of implementation, computational burden and precision of the estimates. Practitioners and researchers, often led by available software, have privileged so far the simplicity of the harmonic mean (HM) estimator. However, it is known that the resulting estimates of the Bayesian evidence in favor of one model are biased and often inaccurate, up to having an infinite variance so that the reliability of the corresponding conclusions is doubtful. We consider possible improvements of the generalized harmonic mean (GHM) idea that recycle Markov Chain Monte Carlo (MCMC) simulations from the posterior, share the computational simplicity of the original HM estimator, but, unlike it, overcome the infinite variance issue. We show reliability and comparative performance of the improved harmonic mean estimators comparing them to approximation techniques relying on improved variants of the thermodynamic integration.

  • Abstract
  • Cite Count Icon 3
  • 10.1093/ve/vez002.065
A66 Tracing the evolutionary history of an emerging Salmonella 4,[5],12:i:- clone in the United States
  • Aug 1, 2019
  • Virus Evolution
  • E Elnekave + 5 more

Salmonellosis is one of the leading causes of foodborne disease worldwide, with an estimated one million cases a year in the United States. Salmonella 4,[5],12: i:-, a monophasic variant of Salmonella typhimurium, is an emerging serovar that has been associated with multiple foodborne outbreaks throughout the world, mostly attributed to pig and pig products. Recently, we have demonstrated that two distinct groups of Salmonella 4,[5],12:i:- circulate in the USA and Europe, with the majority of isolates recovered during recent years belonging to an emerging multidrug-resistant clade (Elnekave et al. 2018). We applied Bayesian phylodynamic reconstruction to uncover the evolutionary history of this clade. We used a dataset of whole-genome sequences of 1446 4,[5],12:i:- isolates from different sources (livestock, human, food products, and others) from the USA (n = 752) and Europe (n = 694), collected between 2008 and 2017 and belonging to the Multilocus Subtype 34, which was predominant in the emerging clade (Elnekave et al. 2018). A subset (n = 110) of Salmonella 4,[5],12:i: isolates was then randomly selected after stratifying by location and year of isolation in order to achieve balanced sampling. Evidence of temporal signal was confirmed by looking at root-to-tip divergences using TempEst. Evolutionary hypotheses using strict and relaxed-clock models were tested using BEAST for a variety of demographic models and assuming a general time reversible substitution model. Model selection was performed by estimating Bayes Factors using path sampling and stepping-stone sampling. The selected model was then used for applying discrete trait models comparing different scenarios of transmission between locations (i.e. bidirectional symmetric/asymmetric or unidirectional). Our preliminary phylodynamic inference results indicate that the origin of this subtype was in Europe and dates back to 1990 (HPD 95%: 1984–2001). We report an exponential growth rate of 0.362 per year, which corresponds to a doubling time of 1.43 years. Our results suggest that this subtype was introduced to the US in the year 2000 (HPD 95%: 1994–2006). Phylodynamic analysis suggests that the recent increase in isolation of Salmonella 4,[5],12:i:- from different sources in the USA may be due to the exponential expansion of an emerging clone which originated in Europe and then expanded to the USA. The emergence and expansion of this serovar is of great public health importance due to the high prevalence of multidrug resistance traits found in USA isolates from this group and especially due to the presence of plasmid-mediated resistance genes for quinolones and extended spectrum cephalosporins, key antimicrobials used for the treatment of invasive Salmonella infections.

  • Research Article
  • Cite Count Icon 23
  • 10.5897/ajb2010.000-3145
Estimation of divergence times for major lineages of galliform birds: Evidence from complete mitochondrial genome sequences
  • May 24, 2010
  • AFRICAN JOURNAL OF BIOTECHNOLOGY
  • Xianzhao Kan + 10 more

Determining an absolute timescale for avian evolutionary history has been recently challenged by the relaxed molecular clock methods, that rates of molecular evolution can vary significantly among organisms. In this study, we used relaxed molecular clocks to date the divergence of major lineages of Galliformes based on complete mitochondrial genomes. A nucleotide dataset of 13 concatenated protein-coding genes from 22 species of Galliformes was used to investigate the evolutionary divergences within the group. Using Gallus bravardi, Schaubortyx and Gallinuloides fossils as calibration points, divergence times analyses were performed with four relaxed molecular clock methods as follows: (1) Bayesian method of Multidivtime; (2) Bayesian Markov chain Monte Carlo (MCMC) analysis of the Bayesian evolutionary analysis by sampling trees (BEAST); (3) local rate minimum deformation method (LRMD) of TREEFINDER; and (4) nonparametric rate smoothing (NPRS) of TREEFINDER. The various relaxed clock methods all indicated that (1) Megapodiidae originated in the Late Cretaceous; (2) Numididae, Phasianidae, Arborophilinae and Coturnicinae originated in the Eocene of Palaeogene; (3) Pavoninae and Gallininae originated at the Eocene-Oligocene boundary; (4) Phasianinae and Meleagridinae originated in the Oligocene; (5) divergence times estimation among most genera of Phasianidae were much older than those of the previous studies. Our results might provide a more likely time scale for evolutionary history of the galliform birds.

  • Research Article
  • Cite Count Icon 1
  • 10.17485/ijst/2018/v11i16/118701
The efficiency of multiple imputation and maximum likelihood methods for estimating missing values
  • Apr 1, 2018
  • Indian Journal of Science and Technology
  • Tlhalitshi Volition Montshiwa + 2 more

Objectives: This study investigated the efficiency of Multiple Imputation (MI) and Maximum Likelihood (ML) methods for estimating missing values. The study was set to use the findings to make recommendations for future studies about the impact of missing data imputation on the accuracy of Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC). Methods: The completedset (with no missing values) used in this study was collected in 2010/11 through the Income and Expenditure Survey (IES) and had 25328 observations. Missing data were generated by randomly deleting 10%, 20%, 30%, 40% and 50% of the values from the complete dataset. The missing values in each of the five datasets were imputed using MI and ML methods. Subsequently, absolute error values of AIC and BIC from multiple regression analysis were computed for each dataset. The study then compared the absolute errors for each missing value imputation method. Findings: The findings of the study revealed that AIC and BIC are more accurate when missing values are estimated by the Full Information Maximum Likelihood (FIML) of the ML algorithm, provided 10% of the data are missing. For all datasets, AIC and BIC were least accurate when missing values were imputed by Expectation Maximisation (EM) of the ML algorithm. The findings also showed that AIC and BIC are more accurate when the rate of MISSINGNESS gets large provided missing values were estimated using either the Fully Conditional Specification (FCS) or Markov Chain Monte Carlo (MCMC), MI algorithms. Application: When the rate of MISSINGNESS is small (at most 10%), FIML should be used to handle missing data if AIC and BIC are going to be used. Also both FCS and MCMC should be considered over EM algorithms when the rate of MISSINGNESS is high (at least 40% missing). Keywords: Maximum Likelihood Imputation, Multiple Imputation, AIC, BIC

  • Research Article
  • 10.1007/s11222-026-10875-z
Approximating evidence via bounded harmonic means
  • Jan 1, 2026
  • Statistics and Computing
  • Dana Naderi + 3 more

Efficient Bayesian model selection relies on the model evidence or marginal likelihood, whose computation often requires evaluating an intractable integral. The harmonic mean estimator (HME) has long been a standard method of approximating the evidence. While computationally simple, the version introduced by Newton and Raftery (1994) potentially suffers from infinite variance. To overcome this issue, Gelfand and Dey (1994) defined a standardized representation of the estimator based on an instrumental function and Robert and Wraith (2009) later proposed to use higher posterior density (HPD) indicators as instrumental functions. Following this approach, a practical method is proposed, based on an elliptical covering of the HPD region with non-overlapping ellipsoids. The resulting estimator, called the Elliptical Covering Marginal Likelihood Estimator (ECMLE), not only eliminates the infinite-variance issue of the original HME and allows exact volume computations, but is also able to be used in multimodal settings. Through several examples, we illustrate that ECMLE outperforms other recent methods such as THAMES and its improved version (Metodiev et al. 2025). Moreover, ECMLE demonstrates lower variance-a key challenge that subsequent HME variants have sought to address-and provides more stable evidence approximations, even in challenging settings.Supplementary InformationThe online version contains supplementary material available at 10.1007/s11222-026-10875-z.

  • Research Article
  • Cite Count Icon 61
  • 10.1371/journal.pone.0048380
Major radiations in the evolution of Caviid rodents: reconciling fossils, ghost lineages, and relaxed molecular clocks.
  • Oct 29, 2012
  • PLoS ONE
  • María Encarnación Pérez + 1 more

BackgroundCaviidae is a diverse group of caviomorph rodents that is broadly distributed in South America and is divided into three highly divergent extant lineages: Caviinae (cavies), Dolichotinae (maras), and Hydrochoerinae (capybaras). The fossil record of Caviidae is only abundant and diverse since the late Miocene. Caviids belongs to Cavioidea sensu stricto (Cavioidea s.s.) that also includes a diverse assemblage of extinct taxa recorded from the late Oligocene to the middle Miocene of South America (“eocardiids”).ResultsA phylogenetic analysis combining morphological and molecular data is presented here, evaluating the time of diversification of selected nodes based on the calibration of phylogenetic trees with fossil taxa and the use of relaxed molecular clocks. This analysis reveals three major phases of diversification in the evolutionary history of Cavioidea s.s. The first two phases involve two successive radiations of extinct lineages that occurred during the late Oligocene and the early Miocene. The third phase consists of the diversification of Caviidae. The initial split of caviids is dated as middle Miocene by the fossil record. This date falls within the 95% higher probability distribution estimated by the relaxed Bayesian molecular clock, although the mean age estimate ages are 3.5 to 7 Myr older. The initial split of caviids is followed by an obscure period of poor fossil record (refered here as the Mayoan gap) and then by the appearance of highly differentiated modern lineages of caviids, which evidentially occurred at the late Miocene as indicated by both the fossil record and molecular clock estimates.ConclusionsThe integrated approach used here allowed us identifying the agreements and discrepancies of the fossil record and molecular clock estimates on the timing of the major events in cavioid evolution, revealing evolutionary patterns that would not have been possible to gather using only molecular or paleontological data alone.

  • Research Article
  • Cite Count Icon 298
  • 10.1093/sysbio/syt069
Species Delimitation Using Bayes Factors: Simulations and Application to the Sceloporus scalaris Species Group (Squamata: Phrynosomatidae)
  • Dec 25, 2013
  • Systematic Biology
  • Jared A Grummer + 2 more

Current molecular methods of species delimitation are limited by the types of species delimitation models and scenarios that can be tested. Bayes factors allow for more flexibility in testing non-nested species delimitation models and hypotheses of individual assignment to alternative lineages. Here, we examined the efficacy of Bayes factors in delimiting species through simulations and empirical data from the Sceloporus scalaris species group. Marginal-likelihood scores of competing species delimitation models, from which Bayes factor values were compared, were estimated with four different methods: harmonic mean estimation (HME), smoothed harmonic mean estimation (sHME), path-sampling/thermodynamic integration (PS), and stepping-stone (SS) analysis. We also performed model selection using a posterior simulation-based analog of the Akaike information criterion through Markov chain Monte Carlo analysis (AICM). Bayes factor species delimitation results from the empirical data were then compared with results from the reversible-jump MCMC (rjMCMC) coalescent-based species delimitation method Bayesian Phylogenetics and Phylogeography (BP&P). Simulation results show that HME and sHME perform poorly compared with PS and SS marginal-likelihood estimators when identifying the true species delimitation model. Furthermore, Bayes factor delimitation (BFD) of species showed improved performance when species limits are tested by reassigning individuals between species, as opposed to either lumping or splitting lineages. In the empirical data, BFD through PS and SS analyses, as well as the rjMCMC method, each provide support for the recognition of all scalaris group taxa as independent evolutionary lineages. Bayes factor species delimitation and BP&P also support the recognition of three previously undescribed lineages. In both simulated and empirical data sets, harmonic and smoothed harmonic mean marginal-likelihood estimators provided much higher marginal-likelihood estimates than PS and SS estimators. The AICM displayed poor repeatability in both simulated and empirical data sets, and produced inconsistent model rankings across replicate runs with the empirical data. Our results suggest that species delimitation through the use of Bayes factors with marginal-likelihood estimates via PS or SS analyses provide a useful and complementary alternative to existing species delimitation methods.

  • Research Article
  • Cite Count Icon 38
  • 10.1016/j.ympev.2014.07.005
Exploring new dating approaches for parasites: The worldwide Apodanthaceae (Cucurbitales) as an example
  • Jul 22, 2014
  • Molecular Phylogenetics and Evolution
  • Sidonie Bellot + 1 more

Exploring new dating approaches for parasites: The worldwide Apodanthaceae (Cucurbitales) as an example

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant