Dataset Information

A general framework for two-stage analysis of genome-wide association studies and its application to case-control studies.

ABSTRACT: Two-stage analyses of genome-wide association studies have been proposed as a means to improving power for designs including family-based association and gene-environment interaction testing. In these analyses, all markers are first screened via a statistic that may not be robust to an underlying assumption, and the markers thus selected are then analyzed in a second stage with a test that is independent from the first stage and is robust to the assumption in question. We give a general formulation of two-stage designs and show how one can use this formulation both to derive existing methods and to improve upon them, opening up a range of possible further applications. We show how using simple regression models in conjunction with external data such as average trait values can improve the power of genome-wide association studies. We focus on case-control studies and show how it is possible to use allele frequencies derived from an external reference to derive a powerful two-stage analysis. An illustration involving the Wellcome Trust Case-Control Consortium data shows several genome-wide-significant associations, subsequently validated, that were not significant in the standard analysis. We give some analytic properties of the methods and discuss some underlying principles.

SUBMITTER: Wason JM

PROVIDER: S-EPMC3376500 | biostudies-literature | 2012 May

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

A general framework for two-stage analysis of genome-wide association studies and its application to case-control studies.

Wason James M S JM Dudbridge Frank F

American journal of human genetics 20120501 5

Two-stage analyses of genome-wide association studies have been proposed as a means to improving power for designs including family-based association and gene-environment interaction testing. In these analyses, all markers are first screened via a statistic that may not be robust to an underlying assumption, and the markers thus selected are then analyzed in a second stage with a test that is independent from the first stage and is robust to the assumption in question. We give a general formulat ...[more]

PMID: 22560088

Similar Datasets

Project description:BackgroundIn microbiome studies, it is important to detect taxa which are associated with pathological outcomes at the lowest definable taxonomic rank, such as genus or species. Traditionally, taxa at the target rank are tested for individual association, followed by the Benjamini-Hochberg (BH) procedure to control for false discovery rate (FDR). However, this approach neglects the dependence structure among taxa and may lead to conservative results. The taxonomic tree of microbiome data represents alignment from phylum to species rank and characterizes evolutionary relationships across microbial taxa. Taxa that are closer on the tree usually have similar responses to the exposure (environment). The statistical power in microbial association tests can be enhanced by efficiently employing the prior evolutionary information via the taxonomic tree.MethodsWe propose a two-stage microbial association mapping framework (massMap) which uses grouping information from the taxonomic tree to strengthen statistical power in association tests at the target rank. massMap first screens the association of taxonomic groups at a pre-selected higher taxonomic rank using a powerful microbial group test OMiAT. The method then proceeds to test the association for each candidate taxon at the target rank within the significant taxonomic groups identified in the first stage. Hierarchical BH (HBH) and selected subset testing (SST) procedures are evaluated to control the FDR for the two-stage structured tests.ResultsOur simulations show that massMap incorporating OMiAT and the advanced FDR controlling methodologies largely alleviates the multiplicity issue. It is statistically more powerful than the traditional association mapping directly at the target rank while controlling the FDR at desired levels under most scenarios. In our real data analyses, massMap detects more or the same amount of associated species with smaller adjusted p values compared to the traditional method, which further illustrates the efficiency of the proposed framework. The R package of massMap is publicly available at https://sites.google.com/site/huilinli09/software and https://github.com/JiyuanHu/ .ConclusionsmassMap is a novel microbial association mapping framework and achieves additional efficiency by utilizing the intrinsic taxonomic structure of microbiome data.

Project description:As genome-wide association studies expand beyond populations of European ancestry, the role of admixture will become increasingly important in the continued discovery and fine-mapping of variation influencing complex traits. Although admixture is commonly viewed as a confounding influence in association studies, approaches such as admixture mapping have demonstrated its ability to highlight disease susceptibility regions of the genome. In this study, we illustrate a powerful two-stage testing strategy designed to uncover trait-associated single nucleotide polymorphisms in the presence of ancestral allele frequency differentiation. In the first stage, we conduct an association scan by using predicted genotypic values based on regional admixture estimates. We then select a subset of promising markers for inclusion in a second-stage analysis, where association is tested between the observed genotype and the phenotype conditional on the predicted genotype. We prove that, under the null hypothesis, the test statistics used in each stage are orthogonal and asymptotically independent. Using simulated data designed to mimic African-American populations in the case of a quantitative trait, we show that our two-stage procedure maintains appropriate control of the family wise error rate and has higher power under realistic effect sizes than the one-stage testing procedure in which all markers are tested for association simultaneously with control of admixture. We apply the proposed procedure to a study of height in 201 African-Americans genotyped at 108 ancestry informative markers. The two-stage procedure identified two statistically significant markers rs1985080 (PTHB1/BBS9) and rs952718 (ABCA12). PTHB1/BBS9 is downregulated by parathyroid hormone in osteoblastic cells and is thought to be involved in parathyroid hormone action in bones and may play a role in height. ABCA12 is a member of the superfamily of ATP-binding cassette transporters and its potential involvement in height is unclear.

Dataset Information

A general framework for two-stage analysis of genome-wide association studies and its application to case-control studies.

Publications

A general framework for two-stage analysis of genome-wide association studies and its application to case-control studies.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets