Unknown

Dataset Information

0

Frequency matrix approach demonstrates high sequence quality in avian BARCODEs and highlights cryptic pseudogenes.


ABSTRACT: The accuracy of DNA barcode databases is critical for research and practical applications. Here we apply a frequency matrix to assess sequencing errors in a very large set of avian BARCODEs. Using 11,000 sequences from 2,700 bird species, we show most avian cytochrome c oxidase I (COI) nucleotide and amino acid sequences vary within a narrow range. Except for third codon positions, nearly all (96%) sites were highly conserved or limited to two nucleotides or two amino acids. A large number of positions had very low frequency variants present in single individuals of a species; these were strongly concentrated at the ends of the barcode segment, consistent with sequencing error. In addition, a small fraction (0.1%) of BARCODEs had multiple very low frequency variants shared among individuals of a species; these were found to represent overlooked cryptic pseudogenes lacking stop codons. The calculated upper limit of sequencing error was 8 × 10(-5) errors/nucleotide, which was relatively high for direct Sanger sequencing of amplified DNA, but unlikely to compromise species identification. Our results confirm the high quality of the avian BARCODE database and demonstrate significant quality improvement in avian COI records deposited in GenBank over the past decade. This approach has potential application for genetic database quality control, discovery of cryptic pseudogenes, and studies of low-level genetic variation.

SUBMITTER: Stoeckle MY 

PROVIDER: S-EPMC3428349 | biostudies-literature | 2012

REPOSITORIES: biostudies-literature

altmetric image

Publications

Frequency matrix approach demonstrates high sequence quality in avian BARCODEs and highlights cryptic pseudogenes.

Stoeckle Mark Y MY   Kerr Kevin C R KC  

PloS one 20120827 8


The accuracy of DNA barcode databases is critical for research and practical applications. Here we apply a frequency matrix to assess sequencing errors in a very large set of avian BARCODEs. Using 11,000 sequences from 2,700 bird species, we show most avian cytochrome c oxidase I (COI) nucleotide and amino acid sequences vary within a narrow range. Except for third codon positions, nearly all (96%) sites were highly conserved or limited to two nucleotides or two amino acids. A large number of po  ...[more]

Similar Datasets

| S-EPMC6099809 | biostudies-literature
| S-EPMC5547596 | biostudies-literature
| S-EPMC10783946 | biostudies-literature
| S-EPMC4717347 | biostudies-literature
| S-EPMC2770531 | biostudies-literature
| S-EPMC6476762 | biostudies-literature
| S-EPMC6369775 | biostudies-literature
| S-EPMC122450 | biostudies-literature
| S-EPMC2359806 | biostudies-other
| S-EPMC7502337 | biostudies-literature