Unknown

Dataset Information

0

SciNER: Extracting Named Entities from Scientific Literature


ABSTRACT: The automated extraction of claims from scientific papers via computer is difficult due to the ambiguity and variability inherent in natural language. Even apparently simple tasks, such as isolating reported values for physical quantities (e.g., “the melting point of X is Y”) can be complicated by such factors as domain-specific conventions about how named entities (the X in the example) are referenced. Although there are domain-specific toolkits that can handle such complications in certain areas, a generalizable, adaptable model for scientific texts is still lacking. As a first step towards automating this process, we present a generalizable neural network model, SciNER, for recognizing scientific entities in free text. Based on bidirectional LSTM networks, our model combines word embeddings, subword embeddings, and external knowledge (from DBpedia) to boost its accuracy. Experiments show that our model outperforms a leading domain-specific extraction toolkit by up to 50%, as measured by F1 score, while also being easily adapted to new domains.

SUBMITTER: Krzhizhanovskaya V 

PROVIDER: S-EPMC7302801 | biostudies-literature | 2020 Jun

REPOSITORIES: biostudies-literature

Similar Datasets

| S-EPMC6351503 | biostudies-literature
| S-EPMC8772125 | biostudies-literature
| S-EPMC4022577 | biostudies-literature
| S-EPMC5156472 | biostudies-literature
| S-EPMC8070292 | biostudies-literature
| S-EPMC7647140 | biostudies-literature
| S-EPMC7787447 | biostudies-literature
| S-EPMC6341248 | biostudies-literature
| S-EPMC8568976 | biostudies-literature
| S-EPMC8293528 | biostudies-literature