Dataset Information

Hapo-G, haplotype-aware polishing of genome assemblies with accurate reads.

ABSTRACT: Single-molecule sequencing technologies have recently been commercialized by Pacific Biosciences and Oxford Nanopore with the promise of sequencing long DNA fragments (kilobases to megabases order) and then, using efficient algorithms, provide high quality assemblies in terms of contiguity and completeness of repetitive regions. However, the error rate of long-read technologies is higher than that of short-read technologies. This has a direct consequence on the base quality of genome assemblies, particularly in coding regions where sequencing errors can disrupt the coding frame of genes. In the case of diploid genomes, the consensus of a given gene can be a mixture between the two haplotypes and can lead to premature stop codons. Several methods have been developed to polish genome assemblies using short reads and generally, they inspect the nucleotide one by one, and provide a correction for each nucleotide of the input assembly. As a result, these algorithms are not able to properly process diploid genomes and they typically switch from one haplotype to another. Herein we proposed Hapo-G (Haplotype-Aware Polishing Of Genomes), a new algorithm capable of incorporating phasing information from high-quality reads (short or long-reads) to polish genome assemblies and in particular assemblies of diploid and heterozygous genomes.

SUBMITTER: Aury JM

PROVIDER: S-EPMC8092372 | biostudies-literature | 2021 Jun

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Hapo-G, haplotype-aware polishing of genome assemblies with accurate reads.

Aury Jean-Marc JM Istace Benjamin B

NAR genomics and bioinformatics 20210503 2

Single-molecule sequencing technologies have recently been commercialized by Pacific Biosciences and Oxford Nanopore with the promise of sequencing long DNA fragments (kilobases to megabases order) and then, using efficient algorithms, provide high quality assemblies in terms of contiguity and completeness of repetitive regions. However, the error rate of long-read technologies is higher than that of short-read technologies. This has a direct consequence on the base quality of genome assemblies, ...[more]

PMID: 33987534

Dataset Information

Hapo-G, haplotype-aware polishing of genome assemblies with accurate reads.

Publications

Hapo-G, haplotype-aware polishing of genome assemblies with accurate reads.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

The genome polishing tool POLCA makes fast and accurate corrections in genome assemblies.
| S-EPMC7347232 | biostudies-literature

Haplotype-aware diplotyping from noisy long reads.
| S-EPMC6547545 | biostudies-literature

Fast and accurate mapping of long reads to complete genome assemblies with VerityMap.
| S-EPMC9808623 | biostudies-literature

Haplotype threading: accurate polyploid phasing from long reads.
| S-EPMC7504856 | biostudies-literature

Repeat and haplotype aware error correction in nanopore sequencing reads with DeChat.
| S-EPMC11659559 | biostudies-literature

phasebook: haplotype-aware de novo assembly of diploid genomes from long reads.
| S-EPMC8549298 | biostudies-literature

Polypolish: Short-read polishing of long-read bacterial genome assemblies.
| S-EPMC8812927 | biostudies-literature

polishCLR: A Nextflow Workflow for Polishing PacBio CLR Genome Assemblies.
| S-EPMC9985148 | biostudies-literature

JASPER: A fast genome polishing tool that improves accuracy of genome assemblies.
| S-EPMC10096238 | biostudies-literature

ARAMIS: From systematic errors of NGS long reads to accurate assemblies.
| S-EPMC8574707 | biostudies-literature