Unknown

Dataset Information

0

Age-dependent topic modeling of comorbidities in UK Biobank identifies disease subtypes with differential genetic risk.


ABSTRACT: The analysis of longitudinal data from electronic health records (EHRs) has the potential to improve clinical diagnoses and enable personalized medicine, motivating efforts to identify disease subtypes from patient comorbidity information. Here we introduce an age-dependent topic modeling (ATM) method that provides a low-rank representation of longitudinal records of hundreds of distinct diseases in large EHR datasets. We applied ATM to 282,957 UK Biobank samples, identifying 52 diseases with heterogeneous comorbidity profiles; analyses of 211,908 All of Us samples produced concordant results. We defined subtypes of the 52 heterogeneous diseases based on their comorbidity profiles and compared genetic risk across disease subtypes using polygenic risk scores (PRSs), identifying 18 disease subtypes whose PRS differed significantly from other subtypes of the same disease. We further identified specific genetic variants with subtype-dependent effects on disease risk. In conclusion, ATM identifies disease subtypes with differential genome-wide and locus-specific genetic risk profiles.

SUBMITTER: Jiang X 

PROVIDER: S-EPMC10632146 | biostudies-literature | 2023 Nov

REPOSITORIES: biostudies-literature

altmetric image

Publications

Age-dependent topic modeling of comorbidities in UK Biobank identifies disease subtypes with differential genetic risk.

Jiang Xilin X   Zhang Martin Jinye MJ   Zhang Yidong Y   Durvasula Arun A   Inouye Michael M   Holmes Chris C   Price Alkes L AL   McVean Gil G  

Nature genetics 20231009 11


The analysis of longitudinal data from electronic health records (EHRs) has the potential to improve clinical diagnoses and enable personalized medicine, motivating efforts to identify disease subtypes from patient comorbidity information. Here we introduce an age-dependent topic modeling (ATM) method that provides a low-rank representation of longitudinal records of hundreds of distinct diseases in large EHR datasets. We applied ATM to 282,957 UK Biobank samples, identifying 52 diseases with he  ...[more]

Similar Datasets

| S-EPMC9272510 | biostudies-literature
| S-EPMC9142639 | biostudies-literature
| S-EPMC6422276 | biostudies-literature
| S-EPMC7454409 | biostudies-literature
| S-EPMC8806365 | biostudies-literature
| S-EPMC7996196 | biostudies-literature
| S-EPMC5681697 | biostudies-literature
| S-EPMC7010560 | biostudies-literature
| S-EPMC10436456 | biostudies-literature
| S-EPMC10990010 | biostudies-literature