bioRxiv Science⌕ Search

Biology subjects

Theesfeld, C. L.

Publications and source records attributed to Theesfeld, C. L..

3 recordsLinked to original sources

An automated framework for efficiently designing deep convolutional neural networks in genomics

Convolutional neural networks (CNNs) have become a standard for analysis of biological sequences. Tuning of network architectures is essential for CNNs performance, yet it requires substantial knowledge of machine learning and commitment of time and effort. This process thus imposes a major barrier to broad and effective application of modern deep learning in genomics. Here, we present AMBER, a fully automated framework to efficiently design and apply CNNs for genomic sequences. AMBER designs optimal models for user-specified biological questions through the state-of-the-art Neural Architecture Search (NAS). We applied AMBER to the task of modelling genomic regulatory features and demonstrated that the predictions of the AMBER-designed model are significantly more accurate than the equivalent baseline non-NAS models and match or even exceed published expert-designed models. Interpretation of AMBER architecture search revealed its design principles of utilizing the full space of computational operations for accurately modelling genomic sequences. Furthermore, we illustrated the use of AMBER to accurately discover functional genomic variants in allele-specific binding and disease heritability enrichment. AMBER provides an efficient automated method for designing accurate deep learning models in genomics.

bioinformatics↗

Genome-wide landscape of RNA-binding protein dysregulation reveals a major impact on psychiatric disorder risk

Despite the strong genetic basis of psychiatric disorders, the molecular origins of these diseases are still largely unmapped. RNA-binding proteins (RBPs) are responsible for most post-transcriptional regulation, from splicing to translational to localization. RBPs thus act as key gatekeepers of cellular homeostasis, especially in the brain. Here, we leverage a deep learning approach to interrogate variant effects genome-wide, and discover that the dysregulation of RBP target sites is a principal contributor to psychiatric disorder risk. We show that specific modes of RBP regulation are genetically linked to the heritability of psychiatric disorders, and demonstrate that diverse RBP regulatory functions are reflected in distinct genome-wide negative selection signatures. Notably, RBP dysregulation has a stronger impact on psychiatric disorders than common coding region variants and explains heritability not currently captured by large-scale molecular QTL studies (expression QTLs and splicing QTLs). We share genome-wide profiles of RBP target site dysregulation, which we used to identify DDHD2 as a candidate schizophrenia risk gene, in a public web server. This resource provides a novel analytical framework to connect the full range of RNA regulation to complex disease.

genetics↗

DeepArk: modeling cis-regulatory codes of model species with deep learning

To enable large-scale analyses of regulatory logic in model species, we developed DeepArk (https://DeepArk.princeton.edu), a set of deep learning models of the cis-regulatory codes of four widely-studied species: Caenorhabditis elegans, Danio rerio, Drosophila melanogaster, and Mus musculus. DeepArk accurately predicts the presence of thousands of different context-specific regulatory features, including chromatin states, histone marks, and transcription factors. In vivo studies show that DeepArk can predict the regulatory impact of any genomic variant (including rare or not previously observed), and enables the regulatory annotation of understudied model species.

bioinformatics↗