bioRxiv Science⌕ Search

Biology subjects

Horlacher, M.

Publications and source records attributed to Horlacher, M..

3 recordsLinked to original sources

Towards In-Silico CLIP-seq: Predicting Protein-RNA Interaction via Sequence-to-Signal Learning

AO_SCPLOWBSTRACTC_SCPLOWUnraveling sequence determinants which drive protein-RNA interaction is crucial for studying binding mechanisms and the impact of genomic variants. While CLIP-seq allows for transcriptome-wide profiling of in vivo protein-RNA interactions, it is limited to expressed transcripts, requiring computational imputation of missing binding information. Existing classification-based methods predict binding with low resolution and depend on prior labeling of transcriptome regions for training. We present RBPNet, a novel deep learning method, which predicts CLIP crosslink count distribution from RNA sequence at single-nucleotide resolution. By training on up to a million regions, RBPNet achieves high generalization on eCLIP, iCLIP and miCLIP assays, outperforming state-of-the-art classifiers. CLIP-seq suffers from various technical biases, complicating downstream interpretation. RBPNet performs bias correction by modeling the raw signal as a mixture of the protein-specific and background signal. Through model interrogation via Integrated Gradients, RBPNet identifies predictive sub-sequences corresponding to known binding motifs and enables variant-impact scoring via in silico mutagenesis. Together, RBPNet improves inference of protein-RNA interaction, as well as mechanistic interpretation of predictions.

bioinformatics↗

Transfer learning reveals sequence determinants of regulatory element accessibility

Dysfunction of regulatory elements through genetic variants is a central mechanism in the pathogenesis of disease. To better understand disease etiology, there is consequently a need to understand how DNA encodes regulatory activity. Deep learning methods show great promise for modeling of biomolecular data from DNA sequence but are limited to large input data for training. Here, we develop ChromTransfer, a transfer learning method that uses a pre-trained, cell-type agnostic model of open chromatin regions as a basis for fine-tuning on regulatory sequences. We demonstrate superior performances with ChromTransfer for learning cell-type specific chromatin accessibility from sequence compared to models not informed by a pre-trained model. Importantly, ChromTransfer enables fine-tuning on small input data with minimal decrease in accuracy. We show that ChromTransfer uses sequence features matching binding site sequences of key transcription factors for prediction. Together, these results demonstrate ChromTransfer as a promising tool for learning the regulatory code.

bioinformatics↗

Computational Mapping of the Human-SARS-CoV-2 Protein-RNA Interactome

Strong evidence suggests that human human RNA-binding proteins (RBPs) are critical factors for viral infection, yet there is no feasible experimental approach to map exact binding sites of RBPs across the SARS-CoV-2 genome systematically at a large scale. We investigated the role of RBPs in the context of SARS-CoV-2 by constructing the first in silico map of human RBP / viral RNA interactions at nucleotide-resolution using two deep learning methods (pysster and DeepRiPe) trained on data from CLIP-seq experiments. We evaluated conservation of RBP binding between 6 other human pathogenic coronaviruses and identified sites of conserved and differential binding in the UTRs of SARS-CoV-1, SARS-CoV-2 and MERS. We scored the impact of variants from 11 viral strains on protein-RNA interaction, identifying a set of gain-and loss of binding events. Lastly, we linked RBPs to functional data and OMICs from other studies, and identified MBNL1, FTO and FXR2 as potential clinical biomarkers. Our results contribute towards a deeper understanding of how viruses hijack host cellular pathways and are available through a comprehensive online resource (https://sc2rbpmap.helmholtz-muenchen.de).

bioinformatics↗