Unknown

Dataset Information

0

Accurate prediction of human essential genes using only nucleotide composition and association information.


ABSTRACT:

Motivation

Previously constructed classifiers in predicting eukaryotic essential genes integrated a variety of features including experimental ones. If we can obtain satisfactory prediction using only nucleotide (sequence) information, it would be more promising. Three groups recently identified essential genes in human cancer cell lines using wet experiments and it provided wonderful opportunity to accomplish our idea. Here we improved the Z curve method into the ?-interval form to denote nucleotide composition and association information and used it to construct the SVM classifying model.

Results

Our model accurately predicted human gene essentiality with an AUC higher than 0.88 both for 5-fold cross-validation and jackknife tests. These results demonstrated that the essentiality of human genes could be reliably reflected by only sequence information. We re-predicted the negative dataset by our Pheg server and 118 genes were additionally predicted as essential. Among them, 20 were found to be homologues in mouse essential genes, indicating that some of the 118 genes were indeed essential, however previous experiments overlooked them. As the first available server, Pheg could predict essentiality for anonymous gene sequences of human. It is also hoped the ?-interval Z curve method could be effectively extended to classification issues of other DNA elements.

Availability and implementation

http://cefg.uestc.edu.cn/Pheg.

Contact

fbguo@uestc.edu.cn.

Supplementary information

Supplementary data are available at Bioinformatics online.

SUBMITTER: Guo FB 

PROVIDER: S-EPMC7110051 | biostudies-literature | 2017 Jun

REPOSITORIES: biostudies-literature

altmetric image

Publications

Accurate prediction of human essential genes using only nucleotide composition and association information.

Guo Feng-Biao FB   Dong Chuan C   Hua Hong-Li HL   Liu Shuo S   Luo Hao H   Zhang Hong-Wan HW   Jin Yan-Ting YT   Zhang Kai-Yue KY  

Bioinformatics (Oxford, England) 20170601 12


<h4>Motivation</h4>Previously constructed classifiers in predicting eukaryotic essential genes integrated a variety of features including experimental ones. If we can obtain satisfactory prediction using only nucleotide (sequence) information, it would be more promising. Three groups recently identified essential genes in human cancer cell lines using wet experiments and it provided wonderful opportunity to accomplish our idea. Here we improved the Z curve method into the λ-interval form to deno  ...[more]

Similar Datasets

| S-EPMC4084796 | biostudies-literature
| S-EPMC4034769 | biostudies-literature
| S-EPMC5463619 | biostudies-literature
| S-EPMC5994942 | biostudies-literature
| S-EPMC4864458 | biostudies-literature
| S-EPMC4965633 | biostudies-literature
| S-EPMC5992449 | biostudies-literature
2016-02-26 | PXD003441 | Pride
| S-EPMC4605288 | biostudies-literature
| S-EPMC3614465 | biostudies-other