bioRxiv Science⌕ Search

Biology subjects

Ferrando, A. A.

Publications and source records attributed to Ferrando, A. A..

2 recordsLinked to original sources

GET: a foundation model of transcription across human cell types

Transcriptional regulation, involving the complex interplay between regulatory sequences and proteins, directs all biological processes. Computational models of transcription lack generalizability to accurately extrapolate in unseen cell types and conditions. Here, we introduce GET, an interpretable foundation model designed to uncover regulatory grammars across 213 human fetal and adult cell types. Relying exclusively on chromatin accessibility data and sequence information, GET achieves experimental-level accuracy in predicting gene expression even in previously unseen cell types. GET showcases remarkable adaptability across new sequencing platforms and assays, enabling regulatory inference across a broad range of cell types and conditions, and uncovering universal and cell type specific transcription factor interaction networks. We evaluated its performance on prediction of regulatory activity, inference of regulatory elements and regulators, and identification of physical interactions between transcription factors. Specifically, we show GET outperforms current models in predicting lentivirus-based massive parallel reporter assay readout with reduced input data. In fetal erythroblasts, we identify distal (>1Mbp) regulatory regions that were missed by previous models. In B cells, we identified a lymphocyte-specific transcription factor-transcription factor interaction that explains the functional significance of a leukemia-risk predisposing germline mutation. In sum, we provide a generalizable and accurate model for transcription together with catalogs of gene regulation and transcription factor interactions, all with cell type specificity.

bioinformatics↗

Computational structure prediction methods enable the systematic identification of oncogenic mutations

Oncogenic mutations are associated with the activation of key pathways necessary for the initiation, progression and treatment-evasion of tumors. While large genomic studies provide the opportunity of identifying these mutations, the vast majority of variants have unclear functional roles presenting a challenge for the use of genomic studies in the clinical/therapeutic setting. Recent developments in predicting protein structures enable the systematic large-scale characterization of structures providing a link from genomic data to functional impact. Here, we observed that most oncogenic mutations tend to occur in protein regions that undergo conformation changes in the presence of the activating mutation or when interacting with a protein partner. By combining evolutionary information and protein structure prediction, we introduce the Evolutionary and Structure (ES) score, a computational approach that enables the systematic identification of hotspot somatic mutations in cancer. The predicted sites tend to occur in Short Linear Motifs and protein-protein interfaces. We test the use of ES-scores in genomic studies in pediatric leukemias that easily recapitulates the main mechanisms of resistance to targeted and chemotherapy drugs. To experimentally test the functional role of the predictions, we performed saturated mutagenesis in NT5C2, a protein commonly mutated in relapsed pediatric lymphocytic leukemias. The approach was able to capture both commonly mutated sites and identify previously uncharacterized functionally relevant regions that are not frequently mutated in these cancers. This work shows that the characterization of protein structures provides a link between large genomic studies, with mostly variants of unknown significance, to functional systematic characterization, prioritizing variants of interest in the therapeutic setting and informing on their possible mechanisms of action.

cancer biology↗