Unknown

Dataset Information

Taxamatch, an algorithm for near ('fuzzy') matching of scientific names in taxonomic databases.


ABSTRACT: Misspellings of organism scientific names create barriers to optimal storage and organization of biological data, reconciliation of data stored under different spelling variants of the same name, and appropriate responses from user queries to taxonomic data systems. This study presents an analysis of the nature of the problem from first principles, reviews some available algorithmic approaches, and describes Taxamatch, an improved name matching solution for this information domain. Taxamatch employs a custom Modified Damerau-Levenshtein Distance algorithm in tandem with a phonetic algorithm, together with a rule-based approach incorporating a suite of heuristic filters, to produce improved levels of recall, precision and execution time over the existing dynamic programming algorithms n-gra

SUBMITTER: Rees T 

PROVIDER: S-EPMC4172526 | biostudies-literature | 2014

REPOSITORIES: biostudies-literature

altmetric image

Publications

Sorry, this publication's infomation has not been loaded in the Indexer, please go directly to PUBMED or Altmetric.

Similar Datasets