Dataset Information

Enhanced regulatory sequence prediction using gapped k-mer features.

ABSTRACT: Oligomers of length k, or k-mers, are convenient and widely used features for modeling the properties and functions of DNA and protein sequences. However, k-mers suffer from the inherent limitation that if the parameter k is increased to resolve longer features, the probability of observing any specific k-mer becomes very small, and k-mer counts approach a binary variable, with most k-mers absent and a few present once. Thus, any statistical learning approach using k-mers as features becomes susceptible to noisy training set k-mer frequencies once k becomes large. To address this problem, we introduce alternative feature sets using gapped k-mers, a new classifier, gkm-SVM, and a general method for robust estimation of k-mer frequencies. To make the method applicable to large-scale genome wide applications, we develop an efficient tree data structure for computing the kernel matrix. We show that compared to our original kmer-SVM and alternative approaches, our gkm-SVM predicts functional genomic regulatory elements and tissue specific enhancers with significantly improved accuracy, increasing the precision by up to a factor of two. We then show that gkm-SVM consistently outperforms kmer-SVM on human ENCODE ChIP-seq datasets, and further demonstrate the general utility of our method using a Naïve-Bayes classifier. Although developed for regulatory sequence analysis, these methods can be applied to any sequence classification problem.

SUBMITTER: Ghandi M

PROVIDER: S-EPMC4102394 | biostudies-literature | 2014 Jul

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Enhanced regulatory sequence prediction using gapped k-mer features.

Ghandi Mahmoud M Lee Dongwon D Mohammad-Noori Morteza M Beer Michael A MA

PLoS computational biology 20140717 7

Oligomers of length k, or k-mers, are convenient and widely used features for modeling the properties and functions of DNA and protein sequences. However, k-mers suffer from the inherent limitation that if the parameter k is increased to resolve longer features, the probability of observing any specific k-mer becomes very small, and k-mer counts approach a binary variable, with most k-mers absent and a few present once. Thus, any statistical learning approach using k-mers as features becomes sus ...[more]

PMID: 25033408

Similar Datasets

Project description:Small ribonucleic acid (sRNA) sequences are 50-500 nucleotide long, noncoding RNA (ncRNA) sequences that play an important role in regulating transcription and translation within a bacterial cell. As such, identifying sRNA sequences within an organism's genome is essential to understand the impact of the RNA molecules on cellular processes. Recently, numerous machine learning models have been applied to predict sRNAs within bacterial genomes. In this study, we considered the sRNA prediction as an imbalanced binary classification problem to distinguish minor positive sRNAs from major negative ones within imbalanced data and then performed a comparative study with six learning algorithms and seven assessment metrics. First, we collected numerical feature groups extracted from known sRNAs previously identified in Salmonella typhimurium LT2 (SLT2) and Escherichia coli K12 (E. coli K12) genomes. Second, as a preliminary study, we characterized the sRNA-size distribution with the conformity test for Benford's law. Third, we applied six traditional classification algorithms to sRNA features and assessed classification performance with seven metrics, varying positive-to-negative instance ratios, and utilizing stratified 10-fold cross-validation. We revisited important individual features and feature groups and found that classification with combined features perform better than with either an individual feature or a single feature group in terms of Area Under Precision-Recall curve (AUPR). We reconfirmed that AUPR properly measures classification performance on imbalanced data with varying imbalance ratios, which is consistent with previous studies on classification metrics for imbalanced data. Overall, eXtreme Gradient Boosting (XGBoost), even without exploiting optimal hyperparameter values, performed better than the other five algorithms with specific optimal parameter settings. As a future work, we plan to extend XGBoost further to a large amount of published sRNAs in bacterial genomes and compare its classification performance with recent machine learning models' performance.

Project description:ObjectiveTo develop a radiomics risk score based on dynamic contrast-enhanced (DCE) MRI for prognosis prediction in patients with glioblastoma.Materials and methodsOne hundred and fifty patients (92 male [61.3%]; mean age ± standard deviation, 60.5 ± 13.5 years) with glioblastoma who underwent preoperative MRI were enrolled in the study. Six hundred and forty-two radiomic features were extracted from volume transfer constant (Ktrans), fractional volume of vascular plasma space (Vp), and fractional volume of extravascular extracellular space (Ve) maps of DCE MRI, wherein the regions of interest were based on both T1-weighted contrast-enhancing areas and non-enhancing T2 hyperintense areas. Using feature selection algorithms, salient radiomic features were selected from the 642 features. Next, a radiomics risk score was developed using a weighted combination of the selected features in the discovery set (n = 105); the risk score was validated in the validation set (n = 45) by investigating the difference in prognosis between the "radiomics risk score" groups. Finally, multivariable Cox regression analysis for progression-free survival was performed using the radiomics risk score and clinical variables as covariates.Results16 radiomic features obtained from non-enhancing T2 hyperintense areas were selected among the 642 features identified. The radiomics risk score was used to stratify high- and low-risk groups in both the discovery and validation sets (both p < 0.001 by the log-rank test). The radiomics risk score and presence of isocitrate dehydrogenase (IDH) mutation showed independent associations with progression-free survival in opposite directions (hazard ratio, 3.56; p = 0.004 and hazard ratio, 0.34; p = 0.022, respectively).ConclusionWe developed and validated the "radiomics risk score" from the features of DCE MRI based on non-enhancing T2 hyperintense areas for risk stratification of patients with glioblastoma. It was associated with progression-free survival independently of IDH mutation status.

Dataset Information

Enhanced regulatory sequence prediction using gapped k-mer features.

Publications

Enhanced regulatory sequence prediction using gapped k-mer features.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets