Unknown

Dataset Information

0

CoreCruncher: Fast and Robust Construction of Core Genomes in Large Prokaryotic Data Sets.


ABSTRACT: The core genome represents the set of genes shared by all, or nearly all, strains of a given population or species of prokaryotes. Inferring the core genome is integral to many genomic analyses, however, most methods rely on the comparison of all the pairs of genomes; a step that is becoming increasingly difficult given the massive accumulation of genomic data. Here, we present CoreCruncher; a program that robustly and rapidly constructs core genomes across hundreds or thousands of genomes. CoreCruncher does not compute all pairwise genome comparisons and uses a heuristic based on the distributions of identity scores to classify sequences as orthologs or paralogs/xenologs. Although it is much faster than current methods, our results indicate that our approach is more conservative than other tools and less sensitive to the presence of paralogs and xenologs. CoreCruncher is freely available from: https://github.com/lbobay/CoreCruncher. CoreCruncher is written in Python 3.7 and can also run on Python 2.7 without modification. It requires the python library Numpy and either Usearch or Blast. Certain options require the programs muscle or mafft.

SUBMITTER: Harris CD 

PROVIDER: S-EPMC7826169 | biostudies-literature |

REPOSITORIES: biostudies-literature

Similar Datasets

| S-EPMC10726161 | biostudies-literature
| S-EPMC7212266 | biostudies-literature
| S-EPMC3439324 | biostudies-literature
| S-EPMC3258164 | biostudies-literature
| S-EPMC2878000 | biostudies-literature
| S-EPMC1087476 | biostudies-literature
| S-EPMC5389551 | biostudies-literature
| S-EPMC2515346 | biostudies-literature
| S-EPMC3544756 | biostudies-literature
| S-EPMC4625461 | biostudies-literature