bioRxiv ScienceSearch

Biology subjects

David K. Gifford

Publications and source records attributed to David K. Gifford.

4 recordsLinked to original sources

Discovering DNA motifs and genomic variants associated with DNA methylation

DNA methylation plays a crucial role in establishing tissue-specific gene expression. However, our incomplete understanding of the cis elements that regulate DNA methylation prevents us from interpreting the functional effects of non-coding variants. We present CpGenie (http://cpgenie.csail.mit.edu), a deep convolutional neural network that learns a regulatory sequence code of DNA methylation and enables allele-specific DNA methylation prediction with single-nucleotide sensitivity. Variant annotations from CpGenie accurately identify methylation quantitative trait loci (meQTL) and contribute to the prioritization of functional non-coding variants including expression quantitative trait loci (eQTL) and disease-associated mutations.

Genomics

Accurate eQTL prioritization with an ensemble-based framework

Expression quantitative trait loci (eQTL) analysis links sequence variants with gene expression change and serves as a successful approach to fine-map variants causal for complex traits and understand their pathogenesis. In this work, we present an ensemble-based computational framework, EnsembleExpr, for eQTL prioritization. When trained on data from massively parallel reporter assays (MPRA), EnsembleExpr accurately predicts reporter expression levels from DNA sequence and identifies sequence variants that exhibit significant allele-specific reporter expression. This framework achieved the best performance in the \"eQTL-causal SNPs\" open challenge in the Fourth Critical Assessment of Genome Interpretation (CAGI 4). We envision EnsembleExpr to be a powerful resource for interpreting non-coding regulatory variants and prioritizing disease-associated mutations for downstream validation.

Genomics

Modular Combinatorial Binding among Human Trans-acting Factors Reveals Direct and Indirect Factor Binding

The combinatorial binding of trans-acting factors (TFs) to regulatory genomic regions is an important basis for the spatial and temporal specificity of gene regulation. We present a new computational approach that reveals how TFs are organized into combinatorial regulatory programs. We define a regulatory program to be a set of TFs that bind together at a regulatory region. Unlike other approaches to characterizing TF binding, we permit a regulatory region to be bound by one or more regulatory programs. We have developed a method called regulatory program discovery (RPD) that produces compact and coherent regulatory programs from in vivo binding data using a topic model. Using RPD we find that the binding of 115 TFs in K562 cells can be organized into 49 interpretable regulatory programs that bind ~140,000 distinct regulatory regions in a modular manner. The discovered regulatory programs recapitulate many published protein-protein physical interactions and have consistent functional annotations of chromatin states. We found that, for certain TFs, direct (motif present) and indirect (motif absent) binding is characterized by distinct sets of binding partners and that the binding of other TFs can predict whether the TF binds directly or indirectly with high accuracy. Joint analysis across two cell types reveals both cell-type-specific and shared regulatory programs and that thousands of regulatory regions use different programs in different cell types. Overall, our results provide comprehensive cell-type-specific combinatorial binding maps and suggest a modular organization of binding programs in regulatory regions.

Genomics

Whole Genome Regulatory Variant Evaluation for Transcription Factor Binding

The majority of disease-associated variants identified in genome-wide association studies (GWAS) reside in noncoding regions of the genome with regulatory roles. Thus being able to interpret the functional consequence of a variant is essential for identifying causal variants in the analysis of GWAS studies. We present GERV (Generative Evaluation of Regulatory Variants), a novel computational method for predicting regulatory variants that affect transcription factor binding. GERV learns a k-mer based generative model of transcription factor binding from ChIP-seq and DNase-seq data, and scores variants by computing the change of predicted ChIP-seq reads between the reference and alternate allele. The k-mers learned by GERV capture more sequence determinants of transcription factor binding than a motif-based approach alone, including both a transcription factors canonical motif as well as associated co-factor motifs. We show that GERV outperforms existing methods in predicting SNPs associated with allele-specific binding. GERV correctly predicts a validated causal variant among linked SNPs, and prioritizes the variants previously reported to modulate the binding of FOXA1 in breast cancer cell lines. Thus, GERV provides a powerful approach for functionally an-notating and prioritizing causal variants for experimental follow-up analysis. The implementation of GERV and related data are available at http://gerv.csail.mit.edu/

Bioinformatics