Unknown

Dataset Information

0

NCMHap: a novel method for haplotype reconstruction based on Neutrosophic c-means clustering.


ABSTRACT:

Background

Single individual haplotype problem refers to reconstructing haplotypes of an individual based on several input fragments sequenced from a specified chromosome. Solving this problem is an important task in computational biology and has many applications in the pharmaceutical industry, clinical decision-making, and genetic diseases. It is known that solving the problem is NP-hard. Although several methods have been proposed to solve the problem, it is found that most of them have low performances in dealing with noisy input fragments. Therefore, proposing a method which is accurate and scalable, is a challenging task.

Results

In this paper, we introduced a method, named NCMHap, which utilizes the Neutrosophic c-means (NCM) clustering algorithm. The NCM algorithm can effectively detect the noise and outliers in the input data. In addition, it can reduce their effects in the clustering process. The proposed method has been evaluated by several benchmark datasets. Comparing with existing methods indicates when NCM is tuned by suitable parameters, the results are encouraging. In particular, when the amount of noise increases, it outperforms the comparing methods.

Conclusion

The proposed method is validated using simulated and real datasets. The achieved results recommend the application of NCMHap on the datasets which involve the fragments with a huge amount of gaps and noise.

SUBMITTER: Zamani F 

PROVIDER: S-EPMC7579908 | biostudies-literature | 2020 Oct

REPOSITORIES: biostudies-literature

altmetric image

Publications

NCMHap: a novel method for haplotype reconstruction based on Neutrosophic c-means clustering.

Zamani Fatemeh F   Olyaee Mohammad Hossein MH   Khanteymoori Alireza A  

BMC bioinformatics 20201022 1


<h4>Background</h4>Single individual haplotype problem refers to reconstructing haplotypes of an individual based on several input fragments sequenced from a specified chromosome. Solving this problem is an important task in computational biology and has many applications in the pharmaceutical industry, clinical decision-making, and genetic diseases. It is known that solving the problem is NP-hard. Although several methods have been proposed to solve the problem, it is found that most of them ha  ...[more]

Similar Datasets

| S-EPMC5112810 | biostudies-literature
| S-EPMC7310859 | biostudies-literature
| S-EPMC6931272 | biostudies-literature
| S-EPMC7094093 | biostudies-literature
| S-EPMC7924552 | biostudies-literature
| S-EPMC7046296 | biostudies-literature
| S-EPMC4711805 | biostudies-literature
| S-EPMC2952426 | biostudies-literature
| S-EPMC6248925 | biostudies-literature
| S-EPMC7987176 | biostudies-literature