Unknown

Dataset Information

0

A cost-sensitive online learning method for peptide identification.


ABSTRACT:

Background

Post-database search is a key procedure in peptide identification with tandem mass spectrometry (MS/MS) strategies for refining peptide-spectrum matches (PSMs) generated by database search engines. Although many statistical and machine learning-based methods have been developed to improve the accuracy of peptide identification, the challenge remains on large-scale datasets and datasets with a distribution of unbalanced PSMs. A more efficient learning strategy is required for improving the accuracy of peptide identification on challenging datasets. While complex learning models have larger power of classification, they may cause overfitting problems and introduce computational complexity on large-scale datasets. Kernel methods map data from the sample space to high dimensional spaces where data relationships can be simplified for modeling.

Results

In order to tackle the computational challenge of using the kernel-based learning model for practical peptide identification problems, we present an online learning algorithm, OLCS-Ranker, which iteratively feeds only one training sample into the learning model at each round, and, as a result, the memory requirement for computation is significantly reduced. Meanwhile, we propose a cost-sensitive learning model for OLCS-Ranker by using a larger loss of decoy PSMs than that of target PSMs in the loss function.

Conclusions

The new model can reduce its false discovery rate on datasets with a distribution of unbalanced PSMs. Experimental studies show that OLCS-Ranker outperforms other methods in terms of accuracy and stability, especially on datasets with a distribution of unbalanced PSMs. Furthermore, OLCS-Ranker is 15-85 times faster than CRanker.

SUBMITTER: Liang X 

PROVIDER: S-EPMC7183122 | biostudies-literature | 2020 Apr

REPOSITORIES: biostudies-literature

altmetric image

Publications

A cost-sensitive online learning method for peptide identification.

Liang Xijun X   Xia Zhonghang Z   Jian Ling L   Wang Yongxiang Y   Niu Xinnan X   Link Andrew J AJ  

BMC genomics 20200425 1


<h4>Background</h4>Post-database search is a key procedure in peptide identification with tandem mass spectrometry (MS/MS) strategies for refining peptide-spectrum matches (PSMs) generated by database search engines. Although many statistical and machine learning-based methods have been developed to improve the accuracy of peptide identification, the challenge remains on large-scale datasets and datasets with a distribution of unbalanced PSMs. A more efficient learning strategy is required for i  ...[more]

Similar Datasets

| S-EPMC4066940 | biostudies-other
| S-EPMC4857761 | biostudies-literature
| S-EPMC9057140 | biostudies-literature
| S-EPMC8725666 | biostudies-literature
| S-EPMC9735681 | biostudies-literature
| S-EPMC1682177 | biostudies-other
| S-EPMC9674908 | biostudies-literature
| S-EPMC7142864 | biostudies-literature
| S-EPMC8230879 | biostudies-literature
2022-09-25 | PXD010613 | Pride