Unknown

Dataset Information

0

Medical Concept Representation Learning from Multi-source Data.


ABSTRACT: Representing words as low dimensional vectors is very useful in many natural language processing tasks. This idea has been extended to medical domain where medical codes listed in medical claims are represented as vectors to facilitate exploratory analysis and predictive modeling. However, depending on a type of a medical provider, medical claims can use medical codes from different ontologies or from a combination of ontologies, which complicates learning of the representations. To be able to properly utilize such multi-source medical claim data, we propose an approach that represents medical codes from different ontologies in the same vector space. We first modify the Pointwise Mutual Information (PMI) measure of similarity between the codes. We then develop a new negative sampling method for word2vec model that implicitly factorizes the modified PMI matrix. The new approach was evaluated on the code cross-reference problem, which aims at identifying similar codes across different ontologies. In our experiments, we evaluated cross-referencing between ICD-9 and CPT medical code ontologies. Our results indicate that vector representations of codes learned by the proposed approach provide superior cross-referencing when compared to several existing approaches.

SUBMITTER: Bai T 

PROVIDER: S-EPMC7047512 | biostudies-literature | 2019 Jul

REPOSITORIES: biostudies-literature

altmetric image

Publications

Medical Concept Representation Learning from Multi-source Data.

Bai Tian T   Egleston Brian L BL   Bleicher Richard R   Vucetic Slobodan S  

IJCAI : proceedings of the conference 20190701


Representing words as low dimensional vectors is very useful in many natural language processing tasks. This idea has been extended to medical domain where medical codes listed in medical claims are represented as vectors to facilitate exploratory analysis and predictive modeling. However, depending on a type of a medical provider, medical claims can use medical codes from different ontologies or from a combination of ontologies, which complicates learning of the representations. To be able to p  ...[more]

Similar Datasets

| S-EPMC8570806 | biostudies-literature
| S-EPMC7235813 | biostudies-literature
| S-EPMC11367907 | biostudies-literature
| S-EPMC10484909 | biostudies-literature
| S-EPMC10407241 | biostudies-literature
| S-EPMC8034601 | biostudies-literature
| S-EPMC5463197 | biostudies-other
| S-EPMC8922327 | biostudies-literature
| S-EPMC9044235 | biostudies-literature
| S-EPMC6487914 | biostudies-literature