Dataset Information

Prediction of future gastric cancer risk using a machine learning algorithm and comprehensive medical check-up data: A case-control study.

ABSTRACT: A comprehensive screening method using machine learning and many factors (biological characteristics, Helicobacter pylori infection status, endoscopic findings and blood test results), accumulated daily as data in hospitals, could improve the accuracy of screening to classify patients at high or low risk of developing gastric cancer. We used XGBoost, a classification method known for achieving numerous winning solutions in data analysis competitions, to capture nonlinear relations among many input variables and outcomes using the boosting approach to machine learning. Longitudinal and comprehensive medical check-up data were collected from 25,942 participants who underwent multiple endoscopies from 2006 to 2017 at a single facility in Japan. The participants were classified into a case group (y?=?1) or a control group (y?=?0) if gastric cancer was or was not detected, respectively, during a 122-month period. Among 1,431 total participants (89 cases and 1,342 controls), 1,144 (80%) were randomly selected for use in training 10 classification models; the remaining 287 (20%) were used to evaluate the models. The results showed that XGBoost outperformed logistic regression and showed the highest area under the curve value (0.899). Accumulating more data in the facility and performing further analyses including other input variables may help expand the clinical utility.

SUBMITTER: Taninaga J

PROVIDER: S-EPMC6712020 | biostudies-literature | 2019 Aug

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Prediction of future gastric cancer risk using a machine learning algorithm and comprehensive medical check-up data: A case-control study.

Taninaga Junichi J Nishiyama Yu Y Fujibayashi Kazutoshi K Gunji Toshiaki T Sasabe Noriko N Iijima Kimiko K Naito Toshio T

Scientific reports 20190827 1

A comprehensive screening method using machine learning and many factors (biological characteristics, Helicobacter pylori infection status, endoscopic findings and blood test results), accumulated daily as data in hospitals, could improve the accuracy of screening to classify patients at high or low risk of developing gastric cancer. We used XGBoost, a classification method known for achieving numerous winning solutions in data analysis competitions, to capture nonlinear relations among many inp ...[more]

PMID: 31455831

Similar Datasets

Project description:Tumor markers are used to screen tens of millions of individuals worldwide at annual health check-ups, especially in East Asia. Machine learning (ML)-based algorithms that improve the diagnostic accuracy and clinical utility of these tests can have substantial impact leading to the early diagnosis of cancer. ML-based algorithms, including a cancer screening algorithm and a secondary organ of origin algorithm, were developed and validated using a large real world dataset (RWD) from asymptomatic individuals undergoing routine cancer screening at a Taiwanese medical center between May 2001 and April 2015. External validation was performed using data from the same period from a separate medical center. The data set included tumor marker values, age, and gender from 27,938 individuals, including 342 subsequently confirmed cancer cases. Separate gender-specific cancer screening algorithms were developed. For men, a logistic regression-based algorithm outperformed single-marker and other ML-based algorithms, with a mean area under the receiver operating characteristic curve (AUROC) of 0.7654 in internal and 0.8736 in external cross validation. For women, a random forest-based algorithm attained a mean AUROC of 0.6665 in internal and 0.6938 in external cross validation. The median time to cancer diagnosis (TTD) in men was 451.5, 204.5, and 28 days for the mild, moderate, and high-risk groups, respectively; for women, the median TTD was 229, 132, and 125 days for the mild, moderate, and high-risk groups. A second algorithm was developed to predict the most likely affected organ systems for at-risk individuals. The algorithm yielded 0.8120 sensitivity and 0.6490 specificity for men, and 0.8170 sensitivity and 0.6750 specificity for women. ML-derived algorithms, trained and validated by using a RWD, can significantly improve tumor marker-based screening for multiple types of early stage cancers, suggest the tissue of origin, and provide guidance for patient follow-up.

Dataset Information

Prediction of future gastric cancer risk using a machine learning algorithm and comprehensive medical check-up data: A case-control study.

Publications

Prediction of future gastric cancer risk using a machine learning algorithm and comprehensive medical check-up data: A case-control study.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets