Unknown

Dataset Information

0

Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement.


ABSTRACT: Advances in modern sequencing technologies allow us to generate sufficient data to analyze hundreds of bacterial genomes from a single machine in a single day. This potential for sequencing massive numbers of genomes calls for fully automated methods to produce high-quality assemblies and variant calls. We introduce Pilon, a fully automated, all-in-one tool for correcting draft assemblies and calling sequence variants of multiple sizes, including very large insertions and deletions. Pilon works with many types of sequence data, but is particularly strong when supplied with paired end data from two Illumina libraries with small e.g., 180 bp and large e.g., 3-5 Kb inserts. Pilon significantly improves draft genome assemblies by correcting bases, fixing mis-assemblies and filling gaps. For both haploid and diploid genomes, Pilon produces more contiguous genomes with fewer errors, enabling identification of more biologically relevant genes. Furthermore, Pilon identifies small variants with high accuracy as compared to state-of-the-art tools and is unique in its ability to accurately identify large sequence variants including duplications and resolve large insertions. Pilon is being used to improve the assemblies of thousands of new genomes and to identify variants from thousands of clinically relevant bacterial strains. Pilon is freely available as open source software.

SUBMITTER: Walker BJ 

PROVIDER: S-EPMC4237348 | biostudies-literature | 2014

REPOSITORIES: biostudies-literature

altmetric image

Publications

Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement.

Walker Bruce J BJ   Abeel Thomas T   Shea Terrance T   Priest Margaret M   Abouelliel Amr A   Sakthikumar Sharadha S   Cuomo Christina A CA   Zeng Qiandong Q   Wortman Jennifer J   Young Sarah K SK   Earl Ashlee M AM  

PloS one 20141119 11


Advances in modern sequencing technologies allow us to generate sufficient data to analyze hundreds of bacterial genomes from a single machine in a single day. This potential for sequencing massive numbers of genomes calls for fully automated methods to produce high-quality assemblies and variant calls. We introduce Pilon, a fully automated, all-in-one tool for correcting draft assemblies and calling sequence variants of multiple sizes, including very large insertions and deletions. Pilon works  ...[more]

Similar Datasets

| S-EPMC7889865 | biostudies-literature
| S-EPMC10403694 | biostudies-literature
| S-EPMC7671403 | biostudies-literature
2013-11-15 | GSE52324 | GEO
2013-11-15 | E-GEOD-52324 | biostudies-arrayexpress
| S-EPMC8627028 | biostudies-literature
| S-EPMC8483015 | biostudies-literature
| S-EPMC4184257 | biostudies-literature
2014-03-05 | GSE55576 | GEO
2014-03-05 | E-GEOD-55576 | biostudies-arrayexpress