bioRxiv ScienceSearch

Biology subjects

Wallace, C.

Publications and source records attributed to Wallace, C..

11 recordsLinked to original sources

Accurate error control in high dimensional association testing using conditional false discovery rates

High-dimensional hypothesis testing is ubiquitous in the biomedical sciences, and informative covariates may be employed to improve power. The conditional false discovery rate (cFDR) is widely-used approach suited to the setting where the covariate is a set of p-values for the equivalent hypotheses for a second trait. Although related to the Benjamini-Hochberg procedure, it does not permit any easy control of type-1 error rate, and existing methods are over-conservative. We propose a new method for type-1 error rate control based on identifying mappings from the unit square to the unit interval defined by the estimated cFDR, and splitting observations so that each map is independent of the observations it is used to test. We also propose an adjustment to the existing cFDR estimator which further improves power. We show by simulation that the new method more than doubles potential improvement in power over unconditional analyses compared to existing methods. We demonstrate our method on transcriptome-wide association studies, and show that the method can be used in an iterative way, enabling the use of multiple covariates successively. Our methods substantially improve the power and applicability of cFDR analysis.

genomics

Improved consistency in estimates of conditional false discovery rates increases power relative to both existing methods and parametric estimators

A common aim in high-dimensional association studies is the identification of the subset of investigated variables associated with a trait of interest. Using association statistics on the same variables for a second related trait can improve power. An important quantity in such analyses is the conditional false-discovery rate (cFDR), the probability of non-association with the trait of interest given p-value thresholds for both traits. The cFDR can be used for hypothesis testing and as a posterior probability in its own right. In this paper, we propose new estimators for the cFDR based on kernel density estimates and mixture-Gaussian models of effect sizes, the latter also allowing estimation of a local form of cFDR (cfdr). We also propose a general non-parametric improvement to existing estimators based on estimating a posterior probability previously estimated at 1. We find that new estimators have the desirable property of smooth rejection regions, but, unexpectedly, do not improve the power of the method, even when distributional assumptions are true. Furthermore, we find that although the local cfdr represents a theoretically optimal decision boundary, noisiness in its estimation means it is less powerful than corresponding cFDR estimates. We find, however, that the non-parametric adjustment increases power for every estimator. We demonstrate the best method on transcriptome-wide association study datasets for breast and ovarian cancers. The findings from this analysis are of both theoretical and pragmatic interest, giving insight into the nature of cFDR and the behaviour of false-discovery rates in a two-dimensional setting. Our methods allow improved control over the behaviour of the cFDR estimator and improved power in high-dimensional hypothesis testing.

genomics

G-quadruplex dynamics contribute to epigenetic regulation of mitochondrial function

Single-stranded DNA or RNA sequences rich in guanine (G) can adopt non-canonical structures known as G-quadruplexes (G4). Predicted G4-forming sequences in the mitochondrial genome are enriched on the heavy-strand and have been associated with formation of deletion breakpoints that cause mitochondrial disorders. However, the functional roles of G4 structures in regulating mitochondrial respiration in non-cancerous cells remain unclear. Here, we demonstrate that RHPS4, previously thought to be a nuclear G4-ligand, localizes primarily to mitochondria in live cells by mechanisms involving mitochondrial membrane potential. We find that RHPS4 exposure causes an acute inhibition of mitochondrial transcript elongation, leading to respiratory complex depletion. At higher ligand doses, RHPS4 causes mitochondrial DNA (mtDNA) replication pausing and genome depletion. Using these different levels of RHPS4 exposure, we describe discrete nuclear gene expression responses associated with mitochondrial transcription inhibition or with mtDNA depletion. Importantly, a mtDNA variant with increased anti-parallel G4-forming characteristic shows a stronger respiratory defect in response to RHPS4, supporting the conclusion that mitochondrial sensitivity to RHPS4 is G4-structure mediated. Thus, we demonstrate a direct role for G4 perturbation in mitochondrial genome replication, transcription processivity, and respiratory function in normal cells and describe the first molecule that differentially recognizes G4 structures in mtDNA.

molecular biology

simGWAS: a fast method for simulation of large scale case-control GWAS summary statistics

MotivationMethods for analysis of GWAS summary statistics have encouraged data sharing and democratised the analysis of different diseases. Ideal validation for such methods is application to simulated data, where some \"truth\" is known. As GWAS increase in size, so does the computational complexity of such evaluations; standard practice repeatedly simulates and analyses genotype data for all individuals in an example study.\n\nResultsWe have developed a novel method based on an alternative approach, directly simulating GWAS summary data, without individual data as an intermediate step. We mathematically derive the expected statistics for any set of causal variants and their effect sizes, conditional upon control haplotype frequencies (available from public reference datasets). Simulation of GWAS summary output can be conducted independently of sample size by simulating random variates about these expected values. Across a range of scenarios, our method, produces very similar output to that from simulating individual genotypes with a substantial gain in speed even for modest sample sizes. Fast simulation of GWAS summary statistics will enable more complete and rapid evaluation of summary statistic methods as well as opening new potential avenues of research in fine mapping and gene set enrichment analysis.\n\nAvailability and ImplementationOur method is available under a GPL license as an R package from http://github.com/chr1swallace/simGWAS\n\nContactcew54@cam.ac.uk\n\nSupplementary InformationSupplementary Information is appended.

genomics

Lysosome enlargement during inhibition of the lipid kinase PIKfyve proceeds through lysosome coalescence

Lysosomes receive and degrade cargo from endocytosis, phagocytosis and autophagy. They also play an important role in sensing and instructing cells on their metabolic state. The lipid kinase PIKfyve generates phosphatidylinositol-3,5-bisphosphate to modulate lysosome function. PIKfyve inhibition leads to impaired degradative capacity, ion dysregulation, abated autophagic flux, and a massive enlargement of lysosomes. Collectively, this leads to various physiological defects including embryonic lethality, neurodegeneration and overt inflammation. While being the most dramatic phenotype, the reasons for lysosome enlargement remain unclear. Here, we examined whether biosynthesis and/or fusion-fission dynamics contribute to swelling. First, we show that PIKfyve inhibition activates TFEB, TFE3 and MITF enhancing lysosome gene expression. However, this did not augment lysosomal protein levels during acute PIKfyve inhibition and deletion of TFEB and/or related proteins did not impair lysosome swelling. Instead, PIKfyve inhibition led to fewer but enlarged lysosomes, suggesting that an imbalance favouring lysosome fusion over fission causes lysosome enlargement. Indeed, conditions that abated fusion curtailed lysosome swelling in PIKfyve-inhibited cells.\n\nSummary statementPIKfyve inhibition causes lysosomes to coalesce, resulting in fewer, enlarged lysosomes. We also show that TFEB-mediated lysosome biosynthesis does not contribute to swelling.

cell biology

DNA methylation oscillation defines classes of enhancers

Understanding the regulatory landscape of human cells requires the integration of genomic and epigenomic maps, capturing combinatorial levels of cell type-specific and invariant activity states.\n\nHere, we segmented whole-genome bisulfite sequencing-derived methylomes into consecutive blocks of co-methylation (COMETs) to obtain spatial variation patterns of DNA methylation (DNAm oscillations) integrated with histone modifications and promoter-enhancer interactions derived from promoter capture Hi-C (PCHi-C) sequencing of the same purified blood cells.\n\nMapping DNAm oscillations onto regulatory genome annotation revealed that enhancers are enriched for DNAm hyper-oscillations (>30-fold), where multiple machine learning models support DNAm as predictive of enhancer location. Based on this analysis, we report overall predictive power of 99% for DNAm oscillations, 77.3% for DNaseI, 41% for CGIs, 20% for UMRs and 0% for LMRs, demonstrating the power of DNAm oscillations over other methods for enhancer prediction. Methylomes of activated and non-activated CD4+ T cells indicate that DNAm oscillations exist in both states irrespective of activation; hence they can be used to determine the location of latent enhancers.\n\nOur approach advances the identification of tissue-specific regulatory elements and outperforms previous approaches defining enhancer classes based on DNA methylation.

genomics

Fine mapping chromatin contacts in capture Hi-C data

Hi-C and capture Hi-C (CHi-C) are used to map physical contacts between chromatin regions in cell nuclei using high-throughput sequencing. Analysis typically proceeds considering the evidence for contacts between each possible pair of fragments independent from other pairs. This can produce long runs of fragments which appear to all make contact with the same baited fragment of interest. We hypothesised that these long runs could result from a smaller subset of direct contacts and propose a new method, based on a Bayesian sparse variable selection approach, which attempts to fine map these direct contacts.\n\nOur model is conceptually novel, exploiting the spatial pattern of counts in CHi-C data, and prioritises fragments with biological properties that would be expected of true contacts. For bait fragments corresponding to gene promoters, we identify contact fragments with active chromatin and contacts that correspond to edges found in previously defined enhancer-target networks; conversely, for intergenic bait fragments, we identify contact fragments corresponding to promoters for genes expressed in that cell type. We show that long runs of apparently co-contacting fragments can typically be explained using a subset of direct contacts consisting of < 10% of the number in the full run, suggesting that greater resolution can be extracted from existing datasets. Our results appear largely complementary to the those from a per-fragment analytical approach, suggesting that they provide an additional level of interpretation that may be used to increase resolution for mapping direct contacts in CHi-C experiments.

genomics

Cells With Treg-Specific FOXP3 Demethylation But Low CD25 Are Prevalent In Autoimmunity

Identification of alterations in the cellular composition of the human immune system is key to understanding the autoimmune process. Recently, a subset of FOXP3+ cells with low CD25 expression was found to be increased in peripheral blood from systemic lupus erythematosus (SLE) patients, although its functional significance remains controversial. Here we find in comparisons with healthy donors that the frequency of FOXP3+ cells within CD127lowCD25low CD4+ T cells (here defined as CD25lowFOXP3+ T cells) is increased in patients affected by autoimmune disease of varying severity, from combined immunodeficiency with active autoimmunity, SLE to type 1 diabetes. We show that CD25lowFOXP3+ T cells share phenotypic features resembling conventional CD127lowCD25highFOXP3+ Tregs, including demethylation of the Treg-specific epigenetic control region in FOXP3 that is highly enriched in HELIOS+ cells, and lack of IL-2 production. As compared to conventional Tregs, more CD25lowFOXP3+HELIOS+ T cells are in cell cycle (33.0% vs 20.7% Ki-67+; P = 1.3 x 10-9) and express the late-stage inhibitory receptor PD-1 (67.2% vs 35.5%; P = 4.0 x 10-18), while having reduced expression of the early-stage inhibitory receptor CTLA-4, as well as other Treg markers, such as FOXP3 and CD15s. The number of CD25lowFOXP3+ T cells are highly correlated (P = 1.2 x 10-19) with the proportion of CD25highFOXP3+ T cells in cell cycle (Ki-67+). These findings suggest that CD25lowFOXP3+ T cells represent a subset of Tregs that are derived from CD25highFOXP3+ T cells, and are a peripheral marker of recent Treg expansion in response to an autoimmune reaction in tissues.\n\nHighlights- FOXP3+ compartment within CD127lowCD25low T cells is expanded in autoimmune patients.\n\n- Increased numbers of CD25lowFOXP3+ T cells are a circulating marker of autoimmunity.\n\n- CD25lowFOXP3+ HELIOS+ T cells are fully demethylated at the FOXP3 TSDR.\n\n- CD25lowFOXP3+ T cells could represent a terminal stage of regulatory T cells.

immunology

Type 1 diabetes genome-wide association analysis with imputation identifies five new risk regions

Type 1 diabetes genotype datasets have undergone several well powered genome wide analysis studies (GWAS), identifying 57 associated regions at the time of analysis. There are still many regions of smaller effect size or low frequency left to discover, and better exploitation of existing type 1 diabetes cohorts with meta analysis and imputation can precede the acquisition of new or larger cohorts. An existing dataset of 5,913 case and 8,829 control samples was analysed using genome-wide microarrays (Affymetrix GeneChip 500K and Illumina Infinium 550K) with imputation via IMPUTE2 with the 1000 Genomes Project (phase 3) reference panel. Genotyping coverage was doubled in known association regions, and increased by four fold in other regions compared to previous studies. Our analysis resulted in new index variants for 17/57 regions, an expanded set of plausible candidate SNPs for 17 regions, and five novel type 1 diabetes association regions at 1p31.3, 1q24.3, 1q31.2, 2q11.2 and 11q12.2. Candidate genes for the new loci included ITGB3BP, FASLG, RGS1, AFF3 and CD5/CD6. Further prioritisation of causal genes and causal variants will require detailed RNA and protein expression studies, in conjunction with genome annotation studies including analysis of physical promoter-enhancer interactions.

genomics

A rare IL2RA haplotype identifies SNP rs61839660 as causal for autoimmunity

IL2RA is associated with multiple autoimmune diseases including type 1 diabetes (T1D). Higher expression of IL2RA mRNA and its protein product CD25 in T lymphocytes is associated with a T1D-protective haplotype. Here we show that a rare variation of this haplotype that loses the protective allele at a single SNP, rs61839660, reduces IL2RA expression and T1D protection, identifying it as the causal factor in disease.

genetics

Chromosome contacts in activated T cells identify autoimmune disease-candidate genes

BackgroundAutoimmune disease-associated variants are preferentially found in regulatory regions in immune cells, particularly CD4+ T cells. Linking such regulatory regions to gene promoters in disease-relevant cell contexts facilitates identification of candidate disease genes.\n\nResultsWithin four hours, activation of CD4+ T cells invokes changes in histone modifications and enhancer RNA transcription that correspond to altered expression of the interacting genes identified by promoter capture Hi-C. By integrating promoter capture Hi-C data with genetic associations for five autoimmune diseases we prioritised 245 candidate genes with a median distance from peak signal to prioritised gene of 153 kb. Just under half (108/245) prioritised genes related to activation-sensitive interactions. This included IL2RA, where allele-specific expression analyses were consistent with its interaction-mediated regulation, illustrating the utility of the approach.\n\nConclusionsOur systematic experimental framework offers an alternative approach to candidate causal gene identification for variants with cell state-specific functional effects, with achievable sample sizes.

genomics