Unknown

Dataset Information

0

MultiDomainBenchmark: a multi-domain query and subject database suite.


ABSTRACT: BACKGROUND:Genetic sequence database retrieval benchmarks play an essential role in evaluating the performance of sequence searching tools. To date, all phylogenetically diverse benchmarks known to the authors include only query sequences with single protein domains. Domains are the primary building blocks of protein structure and function. Independently, each domain can fulfill a single function, but most proteins (>80% in Metazoa) exist as multi-domain proteins. Multiple domain units combine in various arrangements or architectures to create different functions and are often under evolutionary pressures to yield new ones. Thus, it is crucial to create gold standards reflecting the multi-domain complexity of real proteins to more accurately evaluate sequence searching tools. DESCRIPTION:This work introduces MultiDomainBenchmark (MDB), a database suite of 412 curated multi-domain queries and 227,512 target sequences, representing at least 5108 species and 1123 phylogenetically divergent protein families, their relevancy annotation, and domain location. Here, we use the benchmark to evaluate the performance of two commonly used sequence searching tools, BLAST/PSI-BLAST and HMMER. Additionally, we introduce a novel classification technique for multi-domain proteins to evaluate how well an algorithm recovers a domain architecture. CONCLUSION:MDB is publicly available at http://csc.columbusstate.edu/carroll/MDB/ .

SUBMITTER: Carroll HD 

PROVIDER: S-EPMC6376684 | biostudies-literature | 2019 Feb

REPOSITORIES: biostudies-literature

altmetric image

Publications

MultiDomainBenchmark: a multi-domain query and subject database suite.

Carroll Hyrum D HD   Spouge John L JL   Gonzalez Mileidy M  

BMC bioinformatics 20190214 1


<h4>Background</h4>Genetic sequence database retrieval benchmarks play an essential role in evaluating the performance of sequence searching tools. To date, all phylogenetically diverse benchmarks known to the authors include only query sequences with single protein domains. Domains are the primary building blocks of protein structure and function. Independently, each domain can fulfill a single function, but most proteins (>80% in Metazoa) exist as multi-domain proteins. Multiple domain units c  ...[more]

Similar Datasets

| S-EPMC6777547 | biostudies-literature
| S-EPMC4412149 | biostudies-literature
| S-EPMC5460524 | biostudies-other
| S-EPMC4010954 | biostudies-literature
| S-EPMC8304726 | biostudies-literature
| S-EPMC6164097 | biostudies-other
| S-EPMC6030378 | biostudies-literature
| S-EPMC10705291 | biostudies-literature
| S-EPMC5753241 | biostudies-literature
| S-EPMC9512797 | biostudies-literature