Dataset Information

An Alignment-Free Algorithm in Comparing the Similarity of Protein Sequences Based on Pseudo-Markov Transition Probabilities among Amino Acids.

ABSTRACT: In this paper, we have proposed a novel alignment-free method for comparing the similarity of protein sequences. We first encode a protein sequence into a 440 dimensional feature vector consisting of a 400 dimensional Pseudo-Markov transition probability vector among the 20 amino acids, a 20 dimensional content ratio vector, and a 20 dimensional position ratio vector of the amino acids in the sequence. By evaluating the Euclidean distances among the representing vectors, we compare the similarity of protein sequences. We then apply this method into the ND5 dataset consisting of the ND5 protein sequences of 9 species, and the F10 and G11 datasets representing two of the xylanases containing glycoside hydrolase families, i.e., families 10 and 11. As a result, our method achieves a correlation coefficient of 0.962 with the canonical protein sequence aligner ClustalW in the ND5 dataset, much higher than those of other 5 popular alignment-free methods. In addition, we successfully separate the xylanases sequences in the F10 family and the G11 family and illustrate that the F10 family is more heat stable than the G11 family, consistent with a few previous studies. Moreover, we prove mathematically an identity equation involving the Pseudo-Markov transition probability vector and the amino acids content ratio vector.

SUBMITTER: Li Y

PROVIDER: S-EPMC5137889 | biostudies-literature | 2016

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

An Alignment-Free Algorithm in Comparing the Similarity of Protein Sequences Based on Pseudo-Markov Transition Probabilities among Amino Acids.

Li Yushuang Y Song Tian T Yang Jiasheng J Zhang Yi Y Yang Jialiang J

PloS one 20161205 12

In this paper, we have proposed a novel alignment-free method for comparing the similarity of protein sequences. We first encode a protein sequence into a 440 dimensional feature vector consisting of a 400 dimensional Pseudo-Markov transition probability vector among the 20 amino acids, a 20 dimensional content ratio vector, and a 20 dimensional position ratio vector of the amino acids in the sequence. By evaluating the Euclidean distances among the representing vectors, we compare the similarit ...[more]

PMID: 27918587

Dataset Information

An Alignment-Free Algorithm in Comparing the Similarity of Protein Sequences Based on Pseudo-Markov Transition Probabilities among Amino Acids.

Publications

An Alignment-Free Algorithm in Comparing the Similarity of Protein Sequences Based on Pseudo-Markov Transition Probabilities among Amino Acids.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Nonparametric tests for transition probabilities in nonhomogeneous Markov processes.
| S-EPMC7173284 | biostudies-literature

SigAlign: an alignment algorithm guided by explicit similarity criteria.
| S-EPMC11347165 | biostudies-literature

Confidence intervals for Markov chain transition probabilities based on next generation sequencing reads data.
| S-EPMC8277151 | biostudies-literature

A Procedure for Deriving Formulas to Convert Transition Rates to Probabilities for Multistate Markov Models.
| S-EPMC5582645 | biostudies-literature

Alignment-free similarity analysis for protein sequences based on fuzzy integral.
| S-EPMC6391537 | biostudies-literature

A hybrid landmark Aalen-Johansen estimator for transition probabilities in partially non-Markov multi-state models.
| S-EPMC8536588 | biostudies-literature

An algorithm for progressive multiple alignment of sequences with insertions.
| S-EPMC1180752 | biostudies-literature

Research on Components Assembly Platform of Biological Sequences Alignment Algorithm.
| S-EPMC7859483 | biostudies-literature

EpiAlign: an alignment-based bioinformatic tool for comparing chromatin state sequences.
| S-EPMC6648345 | biostudies-literature

Comparison of the flexible parametric survival model and Cox model in estimating Markov transition probabilities using real-world data.
| S-EPMC6104919 | biostudies-literature