Unknown

Dataset Information

0

Nucleotide Archival Format (NAF) enables efficient lossless reference-free compression of DNA sequences.


ABSTRACT:

Summary

DNA sequence databases use compression such as gzip to reduce the required storage space and network transmission time. We describe Nucleotide Archival Format (NAF)-a new file format for lossless reference-free compression of FASTA and FASTQ-formatted nucleotide sequences. Nucleotide Archival Format compression ratio is comparable to the best DNA compressors, while providing dramatically faster decompression. We compared our format with DNA compressors: DELIMINATE and MFCompress, and with general purpose compressors: gzip, bzip2, xz, brotli and zstd.

Availability and implementation

NAF compressor and decompressor, as well as format specification are available at https://github.com/KirillKryukov/naf. Format specification is in public domain. Compressor and decompressor are open source under the zlib/libpng license, free for nearly any use.

Supplementary information

Supplementary data are available at Bioinformatics online.

SUBMITTER: Kryukov K 

PROVIDER: S-EPMC6761962 | biostudies-literature |

REPOSITORIES: biostudies-literature

Similar Datasets

| S-EPMC7265431 | biostudies-literature
| S-EPMC7079445 | biostudies-literature
| S-EPMC7165212 | biostudies-literature
| S-EPMC4481695 | biostudies-literature
| S-EPMC8756627 | biostudies-literature
| S-EPMC7517294 | biostudies-literature
| S-EPMC3083090 | biostudies-literature
| S-EPMC3592443 | biostudies-literature
| S-EPMC3606433 | biostudies-literature
| S-EPMC8271783 | biostudies-literature