Unknown

Dataset Information

0

ClotCatcher: a novel natural language model to accurately adjudicate venous thromboembolism from radiology reports.


ABSTRACT:

Introduction

Accurate identification of venous thromboembolism (VTE) is critical to develop replicable epidemiological studies and rigorous predictions models. Traditionally, VTE studies have relied on international classification of diseases (ICD) codes which are inaccurate - leading to misclassification bias. Here, we developed ClotCatcher, a novel deep learning model that uses natural language processing to detect VTE from radiology reports.

Methods

Radiology reports to detect VTE were obtained from patients admitted to Emory University Hospital (EUH) and Grady Memorial Hospital (GMH). Data augmentation was performed using the Google PEGASUS paraphraser. This data was then used to fine-tune ClotCatcher, a novel deep learning model. ClotCatcher was validated on both the EUH dataset alone and GMH dataset alone.

Results

The dataset contained 1358 studies from EUH and 915 studies from GMH (n = 2273). The dataset contained 1506 ultrasound studies with 528 (35.1%) studies positive for VTE, and 767 CT studies with 91 (11.9%) positive for VTE. When validated on the EUH dataset, ClotCatcher performed best (AUC = 0.980) when trained on both EUH and GMH dataset without paraphrasing. When validated on the GMH dataset, ClotCatcher performed best (AUC = 0.995) when trained on both EUH and GMH dataset with paraphrasing.

Conclusion

ClotCatcher, a novel deep learning model with data augmentation rapidly and accurately adjudicated the presence of VTE from radiology reports. Applying ClotCatcher to large databases would allow for rapid and accurate adjudication of incident VTE. This would reduce misclassification bias and form the foundation for future studies to estimate individual risk for patient to develop incident VTE.

SUBMITTER: Wang J 

PROVIDER: S-EPMC10652606 | biostudies-literature | 2023 Nov

REPOSITORIES: biostudies-literature

altmetric image

Publications

ClotCatcher: a novel natural language model to accurately adjudicate venous thromboembolism from radiology reports.

Wang Jeffrey J   de Vale Joao Souza JS   Gupta Saransh S   Upadhyaya Pulakesh P   Lisboa Felipe A FA   Schobel Seth A SA   Elster Eric A EA   Dente Christopher J CJ   Buchman Timothy G TG   Kamaleswaran Rishikesan R  

BMC medical informatics and decision making 20231116 1


<h4>Introduction</h4>Accurate identification of venous thromboembolism (VTE) is critical to develop replicable epidemiological studies and rigorous predictions models. Traditionally, VTE studies have relied on international classification of diseases (ICD) codes which are inaccurate - leading to misclassification bias. Here, we developed ClotCatcher, a novel deep learning model that uses natural language processing to detect VTE from radiology reports.<h4>Methods</h4>Radiology reports to detect  ...[more]

Similar Datasets

| S-EPMC3811072 | biostudies-literature
| S-EPMC8176715 | biostudies-literature
| S-EPMC9986939 | biostudies-literature
| S-EPMC7686874 | biostudies-literature
| S-EPMC8490271 | biostudies-literature
| S-EPMC9406429 | biostudies-literature
| S-EPMC10130381 | biostudies-literature
| S-EPMC6659158 | biostudies-literature
| S-EPMC4586346 | biostudies-literature
| S-EPMC9391313 | biostudies-literature