Unknown

Dataset Information

0

Quality control for terms and definitions in ontologies and taxonomies.


ABSTRACT:

Background

Ontologies and taxonomies are among the most important computational resources for molecular biology and bioinformatics. A series of recent papers has shown that the Gene Ontology (GO), the most prominent taxonomic resource in these fields, is marked by flaws of certain characteristic types, which flow from a failure to address basic ontological principles. As yet, no methods have been proposed which would allow ontology curators to pinpoint flawed terms or definitions in ontologies in a systematic way.

Results

We present computational methods that automatically identify terms and definitions which are defined in a circular or unintelligible way. We further demonstrate the potential of these methods by applying them to isolate a subset of 6001 problematic GO terms. By automatically aligning GO with other ontologies and taxonomies we were able to propose alternative synonyms and definitions for some of these problematic terms. This allows us to demonstrate that these other resources do not contain definitions superior to those supplied by GO.

Conclusion

Our methods provide reliable indications of the quality of terms and definitions in ontologies and taxonomies. Further, they are well suited to assist ontology curators in drawing their attention to those terms that are ill-defined. We have further shown the limitations of ontology mapping and alignment in assisting ontology curators in rectifying problems, thus pointing to the need for manual curation.

SUBMITTER: Kohler J 

PROVIDER: S-EPMC1482721 | biostudies-literature | 2006 Apr

REPOSITORIES: biostudies-literature

altmetric image

Publications

Quality control for terms and definitions in ontologies and taxonomies.

Köhler Jacob J   Munn Katherine K   Rüegg Alexander A   Skusa Andre A   Smith Barry B  

BMC bioinformatics 20060419


<h4>Background</h4>Ontologies and taxonomies are among the most important computational resources for molecular biology and bioinformatics. A series of recent papers has shown that the Gene Ontology (GO), the most prominent taxonomic resource in these fields, is marked by flaws of certain characteristic types, which flow from a failure to address basic ontological principles. As yet, no methods have been proposed which would allow ontology curators to pinpoint flawed terms or definitions in onto  ...[more]

Similar Datasets

| S-EPMC3589681 | biostudies-other
| S-EPMC5157009 | biostudies-literature
| S-EPMC5660792 | biostudies-literature
| S-EPMC4662321 | biostudies-literature
| S-EPMC5470020 | biostudies-literature
| S-EPMC7251821 | biostudies-literature
| S-EPMC6350067 | biostudies-literature
| S-EPMC6387518 | biostudies-literature
| S-EPMC6624059 | biostudies-literature
| S-EPMC7875264 | biostudies-literature