Unknown

Dataset Information

0

Detecting and locating whole genome duplications on a phylogeny: a probabilistic approach.


ABSTRACT: Whole genome duplications (WGDs) followed by massive gene loss occurred in the evolutionary history of many groups. WGDs are usually inferred from the age distribution of paralogs (Ks-based methods) or from gene collinearity data (synteny). However, Ks-based methods are restricted to detect the recent WGDs due to saturation effects and the difficulty to date old duplicates, and synteny is difficult to reconstruct for distantly related species. Recently, Jiao et al. (Jiao Y, Wickett N, Ayyampalayam S, Chanderbali AS, Landherr L, Ralph PE, Tomsho LP, Hu Y, Liang H, Soltis PS, et al. 2011. Ancestral polyploidy in seed plants and angiosperms. Nature 473:97-100) introduced an empirical method that aims to detect a peak in duplication ages among nodes selected from a previous phylogenetic analysis. In this context, we present here two rigorous methods based on data from multiple gene families and on a new probabilistic model. Our model assumes that all gene lineages are instantaneously duplicated at the WGD event with a possible almost-immediate loss of some extra copies. Our reconciliation method relies on aligned molecular sequences, whereas our gene count method relies only on gene count data across species. We show, using extensive simulations, that both methods have a good detection power. Surprisingly, the gene count method enjoys no loss of power compared with the reconciliation method, despite the fact that sequence information is not used. We finally illustrate the performance of our methods on a benchmark yeast data set. Both methods are able to detect the well-known WGD in the Saccharomyces cerevisiae clade and agree on a small retention rate at the WGD, as established by synteny-based methods.

SUBMITTER: Rabier CE 

PROVIDER: S-EPMC4038794 | biostudies-other | 2014 Mar

REPOSITORIES: biostudies-other

altmetric image

Publications

Detecting and locating whole genome duplications on a phylogeny: a probabilistic approach.

Rabier Charles-Elie CE   Ta Tram T   Ané Cécile C  

Molecular biology and evolution 20131219 3


Whole genome duplications (WGDs) followed by massive gene loss occurred in the evolutionary history of many groups. WGDs are usually inferred from the age distribution of paralogs (Ks-based methods) or from gene collinearity data (synteny). However, Ks-based methods are restricted to detect the recent WGDs due to saturation effects and the difficulty to date old duplicates, and synteny is difficult to reconstruct for distantly related species. Recently, Jiao et al. (Jiao Y, Wickett N, Ayyampalay  ...[more]

Similar Datasets

| S-EPMC6225891 | biostudies-other
| S-EPMC7116244 | biostudies-literature
| S-EPMC3880059 | biostudies-literature
| S-EPMC6680281 | biostudies-literature
| S-EPMC8190143 | biostudies-literature
| S-EPMC10899794 | biostudies-literature
| S-EPMC4125410 | biostudies-literature
| S-EPMC1839080 | biostudies-other
| S-EPMC11199601 | biostudies-literature
| S-EPMC3464658 | biostudies-literature