Unknown

Dataset Information

0

A high-precision rule-based extraction system for expanding geospatial metadata in GenBank records.


ABSTRACT: The metadata reflecting the location of the infected host (LOIH) of virus sequences in GenBank often lacks specificity. This work seeks to enhance this metadata by extracting more specific geographic information from related full-text articles and mapping them to their latitude/longitudes using knowledge derived from external geographical databases.We developed a rule-based information extraction framework for linking GenBank records to the latitude/longitudes of the LOIH. Our system first extracts existing geospatial metadata from GenBank records and attempts to improve it by seeking additional, relevant geographic information from text and tables in related full-text PubMed Central articles. The final extracted locations of the records, based on data assimilated from these sources, are then disambiguated and mapped to their respective geo-coordinates. We evaluated our approach on a manually annotated dataset comprising of 5728 GenBank records for the influenza A virus.We found the precision, recall, and f-measure of our system for linking GenBank records to the latitude/longitudes of their LOIH to be 0.832, 0.967, and 0.894, respectively.Our system had a high level of accuracy for linking GenBank records to the geo-coordinates of the LOIH. However, it can be further improved by expanding our database of geospatial data, incorporating spell correction, and enhancing the rules used for extraction.Our system performs reasonably well for linking GenBank records for the influenza A virus to the geo-coordinates of their LOIH based on record metadata and information extracted from related full-text articles.

SUBMITTER: Tahsin T 

PROVIDER: S-EPMC4997033 | biostudies-literature | 2016 Sep

REPOSITORIES: biostudies-literature

altmetric image

Publications

A high-precision rule-based extraction system for expanding geospatial metadata in GenBank records.

Tahsin Tasnia T   Weissenbacher Davy D   Rivera Robert R   Beard Rachel R   Firago Mari M   Wallstrom Garrick G   Scotch Matthew M   Gonzalez Graciela G  

Journal of the American Medical Informatics Association : JAMIA 20160117 5


<h4>Objective</h4>The metadata reflecting the location of the infected host (LOIH) of virus sequences in GenBank often lacks specificity. This work seeks to enhance this metadata by extracting more specific geographic information from related full-text articles and mapping them to their latitude/longitudes using knowledge derived from external geographical databases.<h4>Materials and methods</h4>We developed a rule-based information extraction framework for linking GenBank records to the latitud  ...[more]

Similar Datasets

| S-EPMC5925778 | biostudies-literature
| S-EPMC6225896 | biostudies-literature
| S-EPMC2995682 | biostudies-literature
| S-EPMC3303717 | biostudies-literature
| S-EPMC7755405 | biostudies-literature
| S-EPMC2808878 | biostudies-literature
| S-EPMC2275786 | biostudies-literature
| S-EPMC5751806 | biostudies-literature
| S-EPMC6518432 | biostudies-literature
| S-EPMC6994485 | biostudies-literature