bioRxiv ScienceSearch

Biology subjects

Hailiang Huang

Publications and source records attributed to Hailiang Huang.

5 recordsLinked to original sources

Insights into the genetic epidemiology of Crohn’s and rare diseases in the Ashkenazi Jewish population

As part of a broader collaborative network of exome sequencing studies, we developed a jointly called data set of 5,685 Ashkenazi Jewish exomes. We make publicly available a resource of site and allele frequencies, which should serve as a reference for medical genetics in the Ashkenazim. We estimate that 30% of protein-coding alleles present in the Ashkenazi Jewish population at frequencies greater than 0.2% are significantly more frequent (mean 7.6-fold) than their maximum frequency observed in other reference populations. Arising via a well-described founder effect, this catalog of enriched alleles can contribute to differences in genetic risk and overall prevalence of diseases between populations. As validation we document 151 AJ enriched protein-altering alleles that overlap with \"pathogenic\" ClinVar alleles, including those that account for 10-100 fold differences in prevalence between AJ and non-AJ populations of some rare diseases including Gaucher disease (GBA, p.Asn409Ser, 8-fold enrichment); Canavan disease (ASPA, p.Glu285Ala, 12-fold enrichment); and Tay-Sachs disease (HEXA, c.1421+1G>C, 27-fold enrichment; p.Tyr427IlefsTer5, 12-fold enrichment). We next sought to use this catalog, of well-established relevance to Mendelian disease, to explore Crohns disease, a common disease with an estimated two to four-fold excess prevalence in AJ. We specifically evaluate whether strong acting rare alleles, enriched by the same founder-effect, contribute excess genetic risk to Crohns disease in AJ, and find that ten rare genetic risk factors in NOD2 and LRRK2 are strongly enriched in AJ, including several novel contributing alleles, show evidence of association to CD. Independently, we find that genomewide common variant risk defined by GWAS shows a strong difference between AJ and non-AJ European control population samples (0.97 s.d. higher, p<10-16). Taken together, the results suggest coordinated selection in AJ population for higher CD risk alleles in general. The results and approach illustrate the value of exome sequencing data in case-control studies along with reference data sets like ExAC to pinpoint genetic variation that contributes to variable disease predisposition across populations.

Genetics

Bootstrat: Population Informed Bootstrapping for Rare Variant Tests

Recent advances in genotyping and sequencing technologies have made detecting rare variants in large cohorts possible. Various analytic methods for associating disease to rare variants have been proposed, including burden tests, C-alpha and SKAT. Most of these methods, however, assume that samples come from a homogeneous population, which is not realistic for analyses of large samples. Not correcting for population stratification causes inflated p-values and false-positive associations. Here we propose a population-informed bootstrap resampling method that controls for population stratification (Bootstrat) in rare variant tests. In essence, the Bootstrat procedure uses genetic distance to create a phenotype probability for each sample. We show that this empirical approach can effectively correct for population stratification while maintaining statistical power comparable to established methods of controlling for population stratification. The Bootstrat scheme can be easily applied to existing rare variant testing methods with reasonable computational complexity.\n\nAuthor SummaryRecent technology advances have enabled large-scale analysis of rare variants, but properly testing rare variants remains a significant challenge as most rare variant testing methods assume a sample of homogenous ethnicity, an assumption often not true for large cohorts. Failure to account for this heterogeneity increases the type I error rate. Here we propose a bootstrap scheme applicable to most existing rare variant testing methods to control for population heterogeneity. This scheme uses a randomization layer to establish a null distribution of the test statistics while preserving the sample genetic relationships. The null distribution is then used to calculate an empirical p-value that accounts for population heterogeneity. We demonstrate how this scheme successfully controls the type I error rate without loss of statistical power.

Genomics

Integrative genetic and epigenetic analysis uncovers regulatory mechanisms of autoimmune disease

Genome-wide association studies in autoimmune and inflammatory diseases (AID) have uncovered hundreds of loci mediating risk1,2. These associations are preferentially located in non-coding DNA regions3,4 and in particular to tissue-specific Dnase I hypersensitivity sites (DHS)5,6. Whilst these analyses clearly demonstrate the overall enrichment of disease risk alleles on gene regulatory regions, they are not designed to identify individual regulatory regions mediating risk or the genes under their control, and thus uncover the specific molecular events driving disease risk. To do so we have departed from standard practice by identifying regulatory regions which replicate across samples, and connect them to the genes they control through robust re-analysis of public data. We find substantial evidence of regulatory potential in 132/301 (44%) risk loci across nine autoimmune and inflammatory diseases, and are able to prioritize a single gene in 104/132 (79%) of these. Thus, we are able to generate testable mechanistic hypotheses of the molecular changes that drive disease risk.

Genetics

The Cancer Epitope Trees of 23 Early Cervical Cancers in Chinese Women

Emerging evidences suggest the heterogeneity of cancers limits the efficacy of immunotherapy. To search for optimal therapeutic targets, we used whole-exome sequencing data from 23 early cervical tumors from Chinese women to investigate the hierarchical structure of the somatic mutations and the predicted neo-epitopes based on their strong binding with major histocompatibility complex class I molecules. We found each tumor carried 117 mutations and 61 neo-epitopes in average and displayed a unique phylogenic tree and \"cancer neo-epitope tree\" comprising different compositions of mutations or neo-epitopes. Conceivably, the neo-epitopes at the top of the tree shared by all cancer cells are the optimal therapeutic targets that might lead to a cure. Human papillomavirus can be used as therapeutic target in only a proportion of cases where the integrated genome exits without active infection. Therefore, the \"cancer neo-epitope tree\" will serve as an important source to determine of the optimal immunotherapeutic target.

Cancer Biology

Association mapping of inflammatory bowel disease loci to single variant resolution

Inflammatory bowel disease (IBD) is a chronic gastrointestinal inflammatory disorder that affects millions worldwide. Genome-wide association studies (GWAS) have identified 200 IBD-associated loci, but few have been conclusively resolved to specific functional variants. Here we report fine-mapping of 94 IBD loci using high-density genotyping in 67,852 individuals. Of the 139 independent associations identified in these regions, 18 were pinpointed to a single causal variant with >95% certainty, and an additional 27 associations to a single variant with >50% certainty. These 45 variants are significantly enriched for protein-coding changes (n=13), direct disruption of transcription factor binding sites (n=3) and tissue specific epigenetic marks (n=10), with the latter category showing enrichment in specific immune cells among associations stronger in CD and gut mucosa among associations stronger in UC. The results of this study suggest that high-resolution, fine-mapping in large samples can convert many GWAS discoveries into statistically convincing causal variants, providing a powerful substrate for experimental elucidation of disease mechanisms.

Genetics