Unknown

Dataset Information

0

SSIF: Subsumption-based Sub-term Inference Framework to audit Gene Ontology.


ABSTRACT: MOTIVATION:The Gene Ontology (GO) is the unifying biological vocabulary for codifying, managing and sharing biological knowledge. Quality issues in GO, if not addressed, can cause misleading results or missed biological discoveries. Manual identification of potential quality issues in GO is a challenging and arduous task, given its growing size. We introduce an automated auditing approach for suggesting potentially missing is-a relations, which may further reveal erroneous is-a relations. RESULTS:We developed a Subsumption-based Sub-term Inference Framework (SSIF) by leveraging a novel term-algebra on top of a sequence-based representation of GO concepts along with three conditional rules (monotonicity, intersection and sub-concept rules). Applying SSIF to the October 3, 2018 release of GO suggested 1938 unique potentially missing is-a relations. Domain experts evaluated a random sample of 210 potentially missing is-a relations. The results showed SSIF achieved a precision of 60.61, 60.49 and 46.03% for the monotonicity, intersection and sub-concept rules, respectively. AVAILABILITY AND IMPLEMENTATION:SSIF is implemented in Java. The source code is available at https://github.com/rashmie/SSIF. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.

SUBMITTER: Abeysinghe R 

PROVIDER: S-EPMC7214018 | biostudies-literature | 2020 May

REPOSITORIES: biostudies-literature

altmetric image

Publications

SSIF: Subsumption-based Sub-term Inference Framework to audit Gene Ontology.

Abeysinghe Rashmie R   Hinderer Eugene W EW   Moseley Hunter N B HNB   Cui Licong L  

Bioinformatics (Oxford, England) 20200501 10


<h4>Motivation</h4>The Gene Ontology (GO) is the unifying biological vocabulary for codifying, managing and sharing biological knowledge. Quality issues in GO, if not addressed, can cause misleading results or missed biological discoveries. Manual identification of potential quality issues in GO is a challenging and arduous task, given its growing size. We introduce an automated auditing approach for suggesting potentially missing is-a relations, which may further reveal erroneous is-a relations  ...[more]

Similar Datasets

| S-EPMC7536087 | biostudies-literature
| S-EPMC2703932 | biostudies-literature
| S-EPMC3775452 | biostudies-literature
| S-EPMC5602685 | biostudies-literature
| S-EPMC3361142 | biostudies-literature
| S-EPMC3774028 | biostudies-literature
| S-EPMC5537389 | biostudies-other
| S-EPMC6602258 | biostudies-literature
| S-EPMC5751813 | biostudies-literature
| S-EPMC3849277 | biostudies-literature