Taxamatch, an algorithm for near ('fuzzy') matching of scientific names in taxonomic databases.
Ontology highlight
ABSTRACT: Misspellings of organism scientific names create barriers to optimal storage and organization of biological data, reconciliation of data stored under different spelling variants of the same name, and appropriate responses from user queries to taxonomic data systems. This study presents an analysis of the nature of the problem from first principles, reviews some available algorithmic approaches, and describes Taxamatch, an improved name matching solution for this information domain. Taxamatch employs a custom Modified Damerau-Levenshtein Distance algorithm in tandem with a phonetic algorithm, together with a rule-based approach incorporating a suite of heuristic filters, to produce improved levels of recall, precision and execution time over the existing dynamic programming algorithms n-gra
SUBMITTER: Rees T
PROVIDER: S-EPMC4172526 | biostudies-literature | 2014
REPOSITORIES: biostudies-literature
ACCESS DATA