Unknown

Dataset Information

0

Utilizing Amino Acid Composition and Entropy of Potential Open Reading Frames to Identify Protein-Coding Genes.


ABSTRACT: One of the main steps in gene-finding in prokaryotes is determining which open reading frames encode for a protein, and which occur by chance alone. There are many different methods to differentiate the two; the most prevalent approach is using shared homology with a database of known genes. This method presents many pitfalls, most notably the catch that you only find genes that you have seen before. The four most popular prokaryotic gene-prediction programs (GeneMark, Glimmer, Prodigal, Phanotate) all use a protein-coding training model to predict protein-coding genes, with the latter three allowing for the training model to be created ab initio from the input genome. Different methods are available for creating the training model, and to increase the accuracy of such tools, we present here GOODORFS, a method for identifying protein-coding genes within a set of all possible open reading frames (ORFS). Our workflow begins with taking the amino acid frequencies of each ORF, calculating an entropy density profile (EDP), using KMeans to cluster the EDPs, and then selecting the cluster with the lowest variation as the coding ORFs. To test the efficacy of our method, we ran GOODORFS on 14,179 annotated phage genomes, and compared our results to the initial training-set creation step of four other similar methods (Glimmer, MED2, PHANOTATE, Prodigal). We found that GOODORFS was the most accurate (0.94) and had the best F1-score (0.85), while Glimmer had the highest precision (0.92) and PHANOTATE had the highest recall (0.96).

SUBMITTER: McNair K 

PROVIDER: S-EPMC7827183 | biostudies-literature | 2021 Jan

REPOSITORIES: biostudies-literature

altmetric image

Publications

Utilizing Amino Acid Composition and Entropy of Potential Open Reading Frames to Identify Protein-Coding Genes.

McNair Katelyn K   Ecale Zhou Carol L CL   Souza Brian B   Malfatti Stephanie S   Edwards Robert A RA  

Microorganisms 20210108 1


One of the main steps in gene-finding in prokaryotes is determining which open reading frames encode for a protein, and which occur by chance alone. There are many different methods to differentiate the two; the most prevalent approach is using shared homology with a database of known genes. This method presents many pitfalls, most notably the catch that you only find genes that you have seen before. The four most popular prokaryotic gene-prediction programs (GeneMark, Glimmer, Prodigal, Phanota  ...[more]

Similar Datasets

2019-07-03 | GSE125218 | GEO
| S-EPMC7085969 | biostudies-literature
2020-03-14 | GSE131650 | GEO
2021-04-28 | GSE154491 | GEO
| S-EPMC5082802 | biostudies-literature
| S-EPMC2527020 | biostudies-literature
| S-EPMC3486843 | biostudies-other
2014-09-11 | E-GEOD-60384 | biostudies-arrayexpress
2018-12-14 | GSE120762 | GEO
| S-EPMC2813248 | biostudies-literature