Dataset Information

Information Retrieval Using Machine Learning for Biomarker Curation in the Exposome-Explorer.

ABSTRACT: Objective: In 2016, the International Agency for Research on Cancer, part of the World Health Organization, released the Exposome-Explorer, the first database dedicated to biomarkers of exposure for environmental risk factors for diseases. The database contents resulted from a manual literature search that yielded over 8,500 citations, but only a small fraction of these publications were used in the final database. Manually curating a database is time-consuming and requires domain expertise to gather relevant data scattered throughout millions of articles. This work proposes a supervised machine learning pipeline to assist the manual literature retrieval process. Methods: The manually retrieved corpus of scientific publications used in the Exposome-Explorer was used as training and testing sets for the machine learning models (classifiers). Several parameters and algorithms were evaluated to predict an article's relevance based on different datasets made of titles, abstracts and metadata. Results: The top performance classifier was built with the Logistic Regression algorithm using the title and abstract set, achieving an F2-score of 70.1%. Furthermore, we extracted 1,143 entities from these articles with a classifier trained for biomarker entity recognition. Of these, we manually validated 45 new candidate entries to the database. Conclusion: Our methodology reduced the number of articles to be manually screened by the database curators by nearly 90%, while only misclassifying 22.1% of the relevant articles. We expect that this methodology can also be applied to similar biomarkers datasets or be adapted to assist the manual curation process of similar chemical or disease databases.

SUBMITTER: Lamurias A

PROVIDER: S-EPMC8417071 | biostudies-literature |

REPOSITORIES: biostudies-literature

ACCESS DATA

Similar Datasets

Project description:Metabolites produced by the gut microbiota play an important role in the cross-talk with the human host. Many microbial metabolites are biologically active and can pass the gut barrier and make it into the systemic circulation, where they form the gut microbial exposome, i.e. the totality of gut microbial metabolites in body fluids or tissues of the host. A major difficulty faced when studying the microbial exposome and its role in health and diseases is to differentiate metabolites solely or partially derived from microbial metabolism from those produced by the host or coming from the diet. Our objective was to collect data from the scientific literature and build a database on gut microbial metabolites and on evidence of their microbial origin. Three types of evidence on the microbial origin of the gut microbial exposome were defined: (1) metabolites are produced in vitro by human faecal bacteria; (2) metabolites show reduced concentrations in humans or experimental animals upon treatment with antibiotics; (3) metabolites show reduced concentrations in germ-free animals when compared with conventional animals. Data was manually collected from peer-reviewed publications and inserted in the Exposome-Explorer database. Furthermore, to explore the chemical space of the microbial exposome and predict metabolites uniquely formed by the microbiota, genome-scale metabolic models (GSMMs) of gut bacterial strains and humans were compared. A total of 1848 records on one or more types of evidence on the gut microbial origin of 457 metabolites was collected in Exposome-Explorer. Data on their known precursors and concentrations in human blood, urine and faeces was also collected. About 66% of the predicted gut microbial metabolites (n = 1543) were found to be unique microbial metabolites not found in the human GSMM, neither in the list of 457 metabolites curated in Exposome-Explorer, and can be targets for new experimental studies. This new data on the gut microbial exposome, freely available in Exposome-Explorer ( http://exposome-explorer.iarc.fr/ ), will help researchers to identify poorly studied microbial metabolites to be considered in future studies on the gut microbiota, and study their functionalities and role in health and diseases.

Dataset Information

Information Retrieval Using Machine Learning for Biomarker Curation in the Exposome-Explorer.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets