Dataset Information

Relative Efficiencies of Simple and Complex Substitution Models in Estimating Divergence Times in Phylogenomics.

ABSTRACT: The conventional wisdom in molecular evolution is to apply parameter-rich models of nucleotide and amino acid substitutions for estimating divergence times. However, the actual extent of the difference between time estimates produced by highly complex models compared with those from simple models is yet to be quantified for contemporary data sets that frequently contain sequences from many species and genes. In a reanalysis of many large multispecies alignments from diverse groups of taxa, we found that the use of the simplest models can produce divergence time estimates and credibility intervals similar to those obtained from the complex models applied in the original studies. This result is surprising because the use of simple models underestimates sequence divergence for all the data sets analyzed. We found three fundamental reasons for the observed robustness of time estimates to model complexity in many practical data sets. First, the estimates of branch lengths and node-to-tip distances under the simplest model show an approximately linear relationship with those produced by using the most complex models applied on data sets with many sequences. Second, relaxed clock methods automatically adjust rates on branches that experience considerable underestimation of sequence divergences, resulting in time estimates that are similar to those from complex models. And, third, the inclusion of even a few good calibrations in an analysis can reduce the difference in time estimates from simple and complex models. The robustness of time estimates to model complexity in these empirical data analyses is encouraging, because all phylogenomics studies use statistical models that are oversimplified descriptions of actual evolutionary substitution processes.

SUBMITTER: Tao Q

PROVIDER: S-EPMC7253201 | biostudies-literature | 2020 Jun

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Relative Efficiencies of Simple and Complex Substitution Models in Estimating Divergence Times in Phylogenomics.

Tao Qiqing Q Barba-Montoya Jose J Huuki Louise A LA Durnan Mary Kathleen MK Kumar Sudhir S

Molecular biology and evolution 20200601 6

The conventional wisdom in molecular evolution is to apply parameter-rich models of nucleotide and amino acid substitutions for estimating divergence times. However, the actual extent of the difference between time estimates produced by highly complex models compared with those from simple models is yet to be quantified for contemporary data sets that frequently contain sequences from many species and genes. In a reanalysis of many large multispecies alignments from diverse groups of taxa, we fo ...[more]

PMID: 32119075

Similar Datasets

Project description:BackgroundEstimating relative causal effects (i.e., "substitution effects") is a common aim of nutritional research. In observational data, this is usually attempted using 1 of 2 statistical modeling approaches: the leave-one-out model and the energy partition model. Despite their widespread use, there are concerns that neither approach is well understood in practice.ObjectivesWe aimed to explore and illustrate the theory and performance of the leave-one-out and energy partition models for estimating substitution effects in nutritional epidemiology.MethodsMonte Carlo data simulations were used to illustrate the theory and performance of both the leave-one-out model and energy partition model, by considering 3 broad types of causal effect estimands: 1) direct substitutions of the exposure with a single component, 2) inadvertent substitutions of the exposure with several components, and 3) average relative causal effects of the exposure instead of all other dietary sources. Models containing macronutrients, foods measured in calories, and foods measured in grams were all examined.ResultsThe leave-one-out and energy partition models both performed equally well when the target estimand involved substituting a single exposure with a single component, provided all variables were measured in the same units. Bias occurred when the substitution involved >1 substituting component. Leave-one-out models that examined foods in mass while adjusting for total energy intake evaluated obscure estimands.ConclusionsRegardless of the approach, substitution models need to be constructed from clearly defined causal effect estimands. Estimands involving a single exposure and a single substituting component are typically estimated more accurately than estimands involving more complex substitutions. The practice of examining foods measured in grams or portions while adjusting for total energy intake is likely to deliver obscure relative effect estimands with unclear interpretations.

Project description:An absolute timescale for evolution is essential if we are to associate evolutionary phenomena, such as adaptation or speciation, with potential causes, such as geological activity or climatic change. Timescales in most phylogenetic studies use geologically dated fossils or phylogeographic events as calibration points, but more recently, it has also become possible to use experimentally derived estimates of the mutation rate as a proxy for substitution rates. The large radiation of drosophilid taxa endemic to the Hawaiian islands has provided multiple calibration points for the Drosophila phylogeny, thanks to the "conveyor belt" process by which this archipelago forms and is colonized by species. However, published date estimates for key nodes in the Drosophila phylogeny vary widely, and many are based on simplistic models of colonization and coalescence or on estimates of island age that are not current. In this study, we use new sequence data from seven species of Hawaiian Drosophila to examine a range of explicit coalescent models and estimate substitution rates. We use these rates, along with a published experimentally determined mutation rate, to date key events in drosophilid evolution. Surprisingly, our estimate for the date for the most recent common ancestor of the genus Drosophila based on mutation rate (25-40 Ma) is closer to being compatible with independent fossil-derived dates (20-50 Ma) than are most of the Hawaiian-calibration models and also has smaller uncertainty. We find that Hawaiian-calibrated dates are extremely sensitive to model choice and give rise to point estimates that range between 26 and 192 Ma, depending on the details of the model. Potential problems with the Hawaiian calibration may arise from systematic variation in the molecular clock due to the long generation time of Hawaiian Drosophila compared with other Drosophila and/or uncertainty in linking island formation dates with colonization dates. As either source of error will bias estimates of divergence time, we suggest mutation rate estimates be used until better models are available.

Dataset Information

Relative Efficiencies of Simple and Complex Substitution Models in Estimating Divergence Times in Phylogenomics.

Publications

Relative Efficiencies of Simple and Complex Substitution Models in Estimating Divergence Times in Phylogenomics.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets