Dataset Information

SWAMP: Sliding Window Alignment Masker for PAML.

ABSTRACT: With the greater availability of genetic data, large genome-wide scans for positive selection increasingly incorporate data from a range of sources. These data sets may be derived from different sequencing methods, each of which has potential sources of error. Sequencing errors, compounded by alignment errors, greatly increase the number of false positives in tests for adaptive evolution. Genome-wide analyses often fail to fully address these issues or to provide sufficient detail on postalignment masking/filtering. Here, we introduce a Sliding Window Alignment Masker for Phylogenetic Analysis by Maximum Likelihood (SWAMP) that scans multiple-sequence alignments for short regions enriched with unreasonably high rates of nonsynonymous substitutions caused, for example, by sequence or alignment errors. SWAMP prevents their inclusion in downstream evolutionary analyses and therefore increases the reliability of downstream analyses. It is able to effectively mask short stretches of erroneous sequence, particularly prevalent in low-coverage genomes, which may not be detected by existing methods based on filtering by sitewise conservation or alignment confidence. SWAMP offers a flexible masking approach, and the user can apply different masking regimens to specific branches or sequences in the phylogeny allowing the stringency of masking to vary according to branch length, expected divergence levels, or assembly quality. We exemplify SWAMPs effectiveness on a dataset of 6,379 protein-coding genes from primate species, including data of variable quality. Full reporting of the software parameters will further improve the reproducibility of genome-wide analyses, as well as reduce false-positive rates.

SUBMITTER: Harrison PW

PROVIDER: S-EPMC4251194 | biostudies-literature | 2014

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

SWAMP: Sliding Window Alignment Masker for PAML.

Harrison Peter W PW Jordan Gregory E GE Montgomery Stephen H SH

Evolutionary bioinformatics online 20141201

With the greater availability of genetic data, large genome-wide scans for positive selection increasingly incorporate data from a range of sources. These data sets may be derived from different sequencing methods, each of which has potential sources of error. Sequencing errors, compounded by alignment errors, greatly increase the number of false positives in tests for adaptive evolution. Genome-wide analyses often fail to fully address these issues or to provide sufficient detail on postalignme ...[more]

PMID: 25525323

Dataset Information

SWAMP: Sliding Window Alignment Masker for PAML.

Publications

SWAMP: Sliding Window Alignment Masker for PAML.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Parametric Dependencies of Sliding Window Correlation.
| S-EPMC6538081 | biostudies-literature

Sliding window correlation analysis: Modulating window shape for dynamic brain connectivity in resting state.
| S-EPMC6513676 | biostudies-literature

SWAV: a web-based visualization browser for sliding window analysis.
| S-EPMC6954255 | biostudies-literature

An average sliding window correlation method for dynamic functional connectivity.
| S-EPMC6865616 | biostudies-literature

Single-scale time-dependent window-sizes in sliding-window dynamic functional connectivity analysis: A validation study.
| S-EPMC7594665 | biostudies-literature

PPP Sliding Window Algorithm and Its Application in Deformation Monitoring.
| S-EPMC4886509 | biostudies-other

Real-Time Stress Assessment Using Sliding Window Based Convolutional Neural Network.
| S-EPMC7472011 | biostudies-literature

GNSS spoofing detection using a maximum likelihood-based sliding window method.
| S-EPMC7454987 | biostudies-literature

Sliding window functional connectivity inference with nonstationary autocorrelations and cross-correlations.
| S-EPMC11212997 | biostudies-literature

A sentence sliding window approach to extract protein annotations from biomedical articles.
| S-EPMC1869011 | biostudies-literature