Unknown

Dataset Information

SPRISS: approximating frequent k-mers by sampling reads, and applications.


ABSTRACT:

Motivation

The extraction of k-mers is a fundamental component in many complex analyses of large next-generation sequencing datasets, including reads classification in genomics and the characterization of RNA-seq datasets. The extraction of all k-mers and their frequencies is extremely demanding in terms of running time and memory, owing to the size of the data and to the exponential number of k-mers to be considered. However, in several applications, only frequent k-mers, which are k-mers appearing in a relatively high proportion of the data, are required by the analysis.

Results

In this work, we present SPRISS, a new efficient algorithm to approximate frequent k-mers and their frequencies in next-generation sequencing data. SPRISS uses a simple yet powerful reads sampling

SUBMITTER: Santoro D 

PROVIDER: S-EPMC9237683 | biostudies-literature | 2022 Jun

REPOSITORIES: biostudies-literature

altmetric image

Publications

Sorry, this publication's infomation has not been loaded in the Indexer, please go directly to PUBMED or Altmetric.

Similar Datasets