Unknown

Dataset Information

0

Exploring species-based strategies for gene normalization.


ABSTRACT: We introduce a system developed for the BioCreative II.5 community evaluation of information extraction of proteins and protein interactions. The paper focuses primarily on the gene normalization task of recognizing protein mentions in text and mapping them to the appropriate database identifiers based on contextual clues. We outline a ""fuzzy" dictionary lookup approach to protein mention detection that matches regularized text to similarly regularized dictionary entries. We describe several different strategies for gene normalization that focus on species or organism mentions in the text, both globally throughout the document and locally in the immediate vicinity of a protein mention, and present the results of experimentation with a series of system variations that explore the effectiveness of the various normalization strategies, as well as the role of external knowledge sources. While our system was neither the best nor the worst performing system in the evaluation, the gene normalization strategies show promise and the system affords the opportunity to explore some of the variables affecting performance on the BCII.5 tasks.

SUBMITTER: Verspoor K 

PROVIDER: S-EPMC2929766 | biostudies-literature | 2010 Jul-Sep

REPOSITORIES: biostudies-literature

altmetric image

Publications

Exploring species-based strategies for gene normalization.

Verspoor Karin K   Roeder Christophe C   Johnson Helen L HL   Cohen K Bretonnel KB   Baumgartner William A WA   Hunter Lawrence E LE  

IEEE/ACM transactions on computational biology and bioinformatics 20100701 3


We introduce a system developed for the BioCreative II.5 community evaluation of information extraction of proteins and protein interactions. The paper focuses primarily on the gene normalization task of recognizing protein mentions in text and mapping them to the appropriate database identifiers based on contextual clues. We outline a ""fuzzy" dictionary lookup approach to protein mention detection that matches regularized text to similarly regularized dictionary entries. We describe several di  ...[more]

Similar Datasets

| S-EPMC105386 | biostudies-literature
| S-EPMC4890630 | biostudies-literature
| S-EPMC7107212 | biostudies-literature
2010-07-18 | E-GEOD-22067 | biostudies-arrayexpress
2010-07-19 | GSE22067 | GEO
| S-EPMC7035612 | biostudies-literature
| S-EPMC3223928 | biostudies-literature
| S-EPMC2680405 | biostudies-literature
| S-EPMC6489209 | biostudies-literature
| S-EPMC6889127 | biostudies-literature