Dataset Information

Pneumonia identification using statistical feature selection.

ABSTRACT: This paper describes a natural language processing system for the task of pneumonia identification. Based on the information extracted from the narrative reports associated with a patient, the task is to identify whether or not the patient is positive for pneumonia.A binary classifier was employed to identify pneumonia from a dataset of multiple types of clinical notes created for 426 patients during their stay in the intensive care unit. For this purpose, three types of features were considered: (1) word n-grams, (2) Unified Medical Language System (UMLS) concepts, and (3) assertion values associated with pneumonia expressions. System performance was greatly increased by a feature selection approach which uses statistical significance testing to rank features based on their association with the two categories of pneumonia identification.Besides testing our system on the entire cohort of 426 patients (unrestricted dataset), we also used a smaller subset of 236 patients (restricted dataset). The performance of the system was compared with the results of a baseline previously proposed for these two datasets. The best results achieved by the system (85.71 and 81.67 F1-measure) are significantly better than the baseline results (50.70 and 49.10 F1-measure) on the restricted and unrestricted datasets, respectively.Using a statistical feature selection approach that allows the feature extractor to consider only the most informative features from the feature space significantly improves the performance over a baseline that uses all the features from the same feature space. Extracting the assertion value for pneumonia expressions further improves the system performance.

SUBMITTER: Bejan CA

PROVIDER: S-EPMC3422830 | biostudies-other | 2012 Sep-Oct

REPOSITORIES: biostudies-other

ACCESS DATA

Publications

Pneumonia identification using statistical feature selection.

Bejan Cosmin Adrian CA Xia Fei F Vanderwende Lucy L Wurfel Mark M MM Yetisgen-Yildiz Meliha M

Journal of the American Medical Informatics Association : JAMIA 20120426 5

<h4>Objective</h4>This paper describes a natural language processing system for the task of pneumonia identification. Based on the information extracted from the narrative reports associated with a patient, the task is to identify whether or not the patient is positive for pneumonia.<h4>Design</h4>A binary classifier was employed to identify pneumonia from a dataset of multiple types of clinical notes created for 426 patients during their stay in the intensive care unit. For this purpose, three ...[more]

PMID: 22539080

Similar Datasets

Project description:There are many toxic chemicals to contaminate the world and cause harm to human and other organisms. How to quickly discriminate these compounds and characterize their potential molecular mechanism and toxicity is essential. High through put transcriptomics profiles such as microarray have been proven useful to identify biomarkers for different classification and toxicity prediction purposes. Here we aim to investigate how to use microarray to predict chemical contaminants and their possible mechanisms. In this study, we divided 105 compounds plus vehicle control into 14 compound classes. On the basis of gene expression profiles of in vitro primary cultured hepatocytes, we comprehensively compared various normalization, feature selection and classification algorithms for the classification of these 14 class compounds. We found that normalization had little effect on the averaged classification accuracy. Two support vector machine methods LibSVM and SMO had better classification performance. When feature sizes were smaller, LibSVM outperformed other classification methods. Simple logistic algorithm also performed well. At the training stage, usually the feature selection method SVM-RFE performed the best, and PCA was the poorest feature selection algorithm. But overall, SVM-RFE had the highest overfitting rate when an independent dataset used for a prediction in this case. Therefore, we developed a new feature selection algorithm called gradient method which had a pretty high training classification as well as prediction accuracy with the lowest over-fitting rate. Through the analysis of biomarkers that distinguished 14 class compounds, we found a goup of genes that mainly invovled in cell cylce were significanly downregulated by the metal and inflammatory compounds, but were induced by anti-microbial, cancer related drugs, pesticides, and PXR mediators. For in vitro experiment, primary cultured rat hepatocytes were treated one of 105 compounds with relative controls. At least three biological replicates were used for each unique condition. In total 531 arrays were used.

Dataset Information

Pneumonia identification using statistical feature selection.

Publications

Pneumonia identification using statistical feature selection.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets