Dataset Information

PaSiT: a novel approach based on short-oligonucleotide frequencies for efficient bacterial identification and typing.

ABSTRACT:

Motivation

One of the most widespread methods used in taxonomy studies to distinguish between strains or taxa is the calculation of average nucleotide identity. It requires a computationally expensive alignment step and is therefore not suitable for large-scale comparisons. Short oligonucleotide-based methods do offer a faster alternative but at the expense of accuracy. Here, we aim to address this shortcoming by providing a software that implements a novel method based on short-oligonucleotide frequencies to compute inter-genomic distances.

Results

Our tetranucleotide and hexanucleotide implementations, which were optimized based on a taxonomically well-defined set of over 200 newly sequenced bacterial genomes, are as accurate as the short oligonucleotide-based method TETRA and average nucleotide identity, for identifying bacterial species and strains, respectively. Moreover, the lightweight nature of this method makes it applicable for large-scale analyses.

Availability and implementation

The method introduced here was implemented, together with other existing methods, in a dependency-free software written in C, GenDisCal, available as source code from https://github.com/LM-UGent/GenDisCal. The software supports multithreading and has been tested on Windows and Linux (CentOS). In addition, a Java-based graphical user interface that acts as a wrapper for the software is also available.

Supplementary information

Supplementary data are available at Bioinformatics online.

SUBMITTER: Goussarov G

PROVIDER: S-EPMC7178395 | biostudies-literature | 2020 Apr

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

PaSiT: a novel approach based on short-oligonucleotide frequencies for efficient bacterial identification and typing.

Goussarov Gleb G Cleenwerck Ilse I Mysara Mohamed M Leys Natalie N Monsieurs Pieter P Tahon Guillaume G Carlier Aurélien A Vandamme Peter P Van Houdt Rob R

Bioinformatics (Oxford, England) 20200401 8

<h4>Motivation</h4>One of the most widespread methods used in taxonomy studies to distinguish between strains or taxa is the calculation of average nucleotide identity. It requires a computationally expensive alignment step and is therefore not suitable for large-scale comparisons. Short oligonucleotide-based methods do offer a faster alternative but at the expense of accuracy. Here, we aim to address this shortcoming by providing a software that implements a novel method based on short-oligonuc ...[more]

PMID: 31899493

Similar Datasets

Project description:BackgroundThe increasing number of sequenced prokaryotic genomes contains a wealth of genomic data that needs to be effectively analysed. A set of statistical tools exists for such analysis, but their strengths and weaknesses have not been fully explored. The statistical methods we are concerned with here are mainly used to examine similarities between archaeal and bacterial DNA from different genomes. These methods compare observed genomic frequencies of fixed-sized oligonucleotides with expected values, which can be determined by genomic nucleotide content, smaller oligonucleotide frequencies, or be based on specific statistical distributions. Advantages with these statistical methods include measurements of phylogenetic relationship with relatively small pieces of DNA sampled from almost anywhere within genomes, detection of foreign/conserved DNA, and homology searches. Our aim was to explore the reliability and best suited applications for some popular methods, which include relative oligonucleotide frequencies (ROF), di- to hexanucleotide zero'th order Markov methods (ZOM) and 2.order Markov chain Method (MCM). Tests were performed on distant homology searches with large DNA sequences, detection of foreign/conserved DNA, and plasmid-host similarity comparisons. Additionally, the reliability of the methods was tested by comparing both real and random genomic DNA.ResultsOur findings show that the optimal method is context dependent. ROFs were best suited for distant homology searches, whilst the hexanucleotide ZOM and MCM measures were more reliable measures in terms of phylogeny. The dinucleotide ZOM method produced high correlation values when used to compare real genomes to an artificially constructed random genome with similar %GC, and should therefore be used with care. The tetranucleotide ZOM measure was a good measure to detect horizontally transferred regions, and when used to compare the phylogenetic relationships between plasmids and hosts, significant correlation (R2 = 0.4) was found with genomic GC content and intra-chromosomal homogeneity.ConclusionThe statistical methods examined are fast, easy to implement, and powerful for a number of different applications involving genomic sequence comparisons. However, none of the measures examined were superior in all tests, and therefore the choice of the statistical method should depend on the task at hand.

Project description:BackgroundThe metagenomic analysis of microbial communities holds the potential to improve our understanding of the role of microbes in clinical conditions. Recent, dramatic improvements in DNA sequencing throughput and cost will enable such analyses on individuals. However, such advances in throughput generally come at the cost of shorter read-lengths, limiting the discriminatory power of each read. In particular, classifying the microbial content of samples by sequencing the < 1,600 bp 16S rRNA gene will be affected by such limitations.ResultsWe describe a method for identifying the phylogenetic content of bacterial samples using high-throughput Pyrosequencing targeted at the 16S rRNA gene. Our analysis is adapted to the shorter read-lengths of such technology and uses a database of 16S rDNA to determine the most specific phylogenetic classification for reads, resulting in a weighted phylogenetic tree characterizing the content of the sample. We present results for six samples obtained from the human vagina during pregnancy that corroborates previous studies using conventional techniques.Next, we analyze the power of our method to classify reads at each level of the phylogeny using simulation experiments. We assess the impacts of read-length and database completeness on our method, and predict how we do as technology improves and more bacteria are sequenced. Finally, we study the utility of targeting specific 16S variable regions and show that such an approach considerably improves results for certain types of microbial samples. Using simulation, our method can be used to determine the most informative variable region.ConclusionThis study provides positive validation of the effectiveness of targeting 16S metagenomes using short-read sequencing technology. Our methodology allows us to infer the most specific assignment of the sequence reads within the phylogeny, and to identify the most discriminative variable region to target. The analysis of high-throughput Pyrosequencing on human flora samples will accelerate the study of the relationship between the microbial world and ourselves.

Project description:Pathogen typing is pivotal to detecting the emergence of high-risk clones in hospital settings and to limit their spread. Unfortunately, the most commonly used typing methods (i.e., pulsed-field gel electrophoresis [PFGE], multilocus sequence typing [MLST], and whole-genome sequencing [WGS]) are expensive or time-consuming, limiting their application to real-time surveillance. High-resolution melting (HRM) can be applied to perform cost-effective and fast pathogen typing, but developing highly discriminatory protocols is challenging. Here, we present hypervariable-locus melting typing (HLMT), a novel approach to HRM-based typing that enables the development of more effective and portable typing protocols. HLMT types the strains by assigning them to melting types (MTs) on the basis of a reference data set (HLMT-assignment) and/or by clustering them using melting temperatures (HLMT-clustering). We applied the HLMT protocol developed on the capsular gene wzi for Klebsiella pneumoniae on 134 strains collected during surveillance programs in four hospitals. Then, we compared the HLMT results to those obtained using wzi, MLST, WGS, and PFGE typing. HLMT distinguished most of the K. pneumoniae high-risk clones with a sensitivity comparable to that of PFGE and MLST+wzi. It also drew surveillance epidemiological curves comparable to those obtained using MLST+wzi, PFGE, and WGS typing. Furthermore, the results obtained using HLMT-assignment were consistent with those of wzi typing for 95% of the typed strains, with a Jaccard index value of 0.9. HLMT is a fast and scalable approach for pathogen typing, suitable for real-time hospital microbiological surveillance. HLMT is also inexpensive, and thus, it is applicable for infection control programs in low- and middle-income countries. IMPORTANCE In this work, we describe hypervariable-locus melting typing (HLMT), a novel fast approach to pathogen typing using the high-resolution melting (HRM) assay. The method includes a novel approach for gene target selection, primer design, and HRM data analysis. We successfully applied this method to distinguish the high-risk clones of Klebsiella pneumoniae, one of the most important nosocomial pathogens worldwide. We also compared HLMT to typing using WGS, the capsular gene wzi, MLST, and PFGE. Our results show that HLMT is a typing method suitable for real-time epidemiological investigation. The application of HLMT to hospital microbiology surveillance can help to rapidly detect outbreak emergence, improving the effectiveness of infection control strategies.

Dataset Information

PaSiT: a novel approach based on short-oligonucleotide frequencies for efficient bacterial identification and typing.

Motivation

Results

Availability and implementation

Supplementary information

Publications

PaSiT: a novel approach based on short-oligonucleotide frequencies for efficient bacterial identification and typing.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets