Unknown

Dataset Information

0

ROCker: accurate detection and quantification of target genes in short-read metagenomic data sets by modeling sliding-window bitscores.


ABSTRACT: Functional annotation of metagenomic and metatranscriptomic data sets relies on similarity searches based on e-value thresholds resulting in an unknown number of false positive and negative matches. To overcome these limitations, we introduce ROCker, aimed at identifying position-specific, most-discriminant thresholds in sliding windows along the sequence of a target protein, accounting for non-discriminative domains shared by unrelated proteins. ROCker employs the receiver operating characteristic (ROC) curve to minimize false discovery rate (FDR) and calculate the best thresholds based on how simulated shotgun metagenomic reads of known composition map onto well-curated reference protein sequences and thus, differs from HMM profiles and related methods. We showcase ROCker using ammonia monooxygenase (amoA) and nitrous oxide reductase (nosZ) genes, mediating oxidation of ammonia and the reduction of the potent greenhouse gas, N2O, to inert N2, respectively. ROCker typically showed 60-fold lower FDR when compared to the common practice of using fixed e-values. Previously uncounted ‘atypical’ nosZ genes were found to be two times more abundant, on average, than their typical counterparts in most soil metagenomes and the abundance of bacterial amoA was quantified against the highly-related particulate methane monooxygenase (pmoA). Therefore, ROCker can reliably detect and quantify target genes in short-read metagenomes.

SUBMITTER: Orellana LH 

PROVIDER: S-EPMC5388429 | biostudies-literature | 2017 Feb

REPOSITORIES: biostudies-literature

altmetric image

Publications

ROCker: accurate detection and quantification of target genes in short-read metagenomic data sets by modeling sliding-window bitscores.

Orellana Luis H LH   Rodriguez-R Luis M LM   Konstantinidis Konstantinos T KT  

Nucleic acids research 20170201 3


Functional annotation of metagenomic and metatranscriptomic data sets relies on similarity searches based on e-value thresholds resulting in an unknown number of false positive and negative matches. To overcome these limitations, we introduce ROCker, aimed at identifying position-specific, most-discriminant thresholds in sliding windows along the sequence of a target protein, accounting for non-discriminative domains shared by unrelated proteins. ROCker employs the receiver operating characteris  ...[more]

Similar Datasets

2024-07-10 | GSE271528 | GEO
2024-07-10 | GSE271530 | GEO
2024-07-10 | GSE271527 | GEO
2023-09-01 | GSE225380 | GEO
| S-EPMC3919567 | biostudies-literature
| PRJNA935371 | ENA
| S-EPMC3106329 | biostudies-literature
| PRJNA1132172 | ENA
| S-EPMC7993542 | biostudies-literature
2023-09-01 | GSE225377 | GEO