Unknown

Dataset Information

0

Minimal absent words in four human genome assemblies.


ABSTRACT: Minimal absent words have been computed in genomes of organisms from all domains of life. Here, we aim to contribute to the catalogue of human genomic variation by investigating the variation in number and content of minimal absent words within a species, using four human genome assemblies. We compare the reference human genome GRCh37 assembly, the HuRef assembly of the genome of Craig Venter, the NA12878 assembly from cell line GM12878, and the YH assembly of the genome of a Han Chinese individual. We find the variation in number and content of minimal absent words between assemblies more significant for large and very large minimal absent words, where the biases of sequencing and assembly methodologies become more pronounced. Moreover, we find generally greater similarity between the human genome assemblies sequenced with capillary-based technologies (GRCh37 and HuRef) than between the human genome assemblies sequenced with massively parallel technologies (NA12878 and YH). Finally, as expected, we find the overall variation in number and content of minimal absent words within a species to be generally smaller than the variation between species.

SUBMITTER: Garcia SP 

PROVIDER: S-EPMC3248429 | biostudies-literature | 2011

REPOSITORIES: biostudies-literature

altmetric image

Publications

Minimal absent words in four human genome assemblies.

Garcia Sara P SP   Pinho Armando J AJ  

PloS one 20111229 12


Minimal absent words have been computed in genomes of organisms from all domains of life. Here, we aim to contribute to the catalogue of human genomic variation by investigating the variation in number and content of minimal absent words within a species, using four human genome assemblies. We compare the reference human genome GRCh37 assembly, the HuRef assembly of the genome of Craig Venter, the NA12878 assembly from cell line GM12878, and the YH assembly of the genome of a Han Chinese individ  ...[more]

Similar Datasets

| S-EPMC2375138 | biostudies-literature
| S-EPMC4514932 | biostudies-literature
| S-EPMC7807135 | biostudies-literature
| S-EPMC5017631 | biostudies-literature
| S-EPMC6972860 | biostudies-literature
| S-EPMC357027 | biostudies-literature
| S-EPMC4422153 | biostudies-literature
| S-EPMC7818936 | biostudies-literature
| S-EPMC3276541 | biostudies-literature
| PRJNA588796 | ENA