bioRxiv Science⌕ Search

Biology subjects

Jagota, M.

Publications and source records attributed to Jagota, M..

5 recordsLinked to original sources

Learning antibody sequence constraints from allelic inclusion

Antibodies and B-cell receptors (BCRs) are produced by B cells, and are built of a heavy chain and a light chain. Although each B cell could express two different heavy chains and four different light chains, usually only a unique pair of heavy chain and light chain is expressed--a phenomenon known as allelic exclusion. However, a small fraction of naive-B cells violate allelic exclusion by expressing two productive light chains, one of which has impaired function; this has been called allelic inclusion. We demonstrate that these B cells can be used to learn constraints on antibody sequence. Using large-scale single-cell sequencing data from humans, we find examples of light chain allelic inclusion in thousands of naive-B cells, which is an order of magnitude larger than existing datasets. We train machine learning models to identify the abnormal sequences in these cells. The resulting models correlate with antibody properties that they were not trained on, including polyreactivity, surface expression, and mutation usage in affinity maturation. These correlations are larger than what is achieved by existing antibody modeling approaches, indicating that allelic inclusion data contains useful new information. We also investigate the impact of similar selection forces on the heavy chain in mouse, and observe that pairing with the surrogate light chain significantly restricts heavy chain diversity.

immunology↗

Dissecting the oncogenic mechanisms of POT1 cancer mutations through deep scanning mutagenesis

Mutations in the shelterin protein POT1 are associated with diverse cancers, but their role in cancer progression remains unclear. To resolve this, we performed deep scanning mutagenesis in POT1 locally haploid human stem cells to assess the impact of POT1 variants on cellular viability and cancer-associated telomeric phenotypes. Though POT1 is essential, frame-shift mutants are rescued by chemical ATR inhibition, indicating that POT1 is not required for telomere replication or lagging strand synthesis. In contrast, a substantial fraction of clinically-validated pathogenic mutations support normal cellular proliferation, but still drive ATR-dependent telomeric DNA damage signaling and ATR-independent telomere elongation. Moreover, this class of cancer-associated POT1 variants elongates telomeres more rapidly than POT1 frame-shifts, indicating they actively drive oncogenesis and are not simple loss-of-function mutations.

cancer biology↗

Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects

Genetic variation in the human genome is a major determinant of individual disease risk, but the vast majority of missense variants have unknown etiological effects. Here, we present a robust learning framework for leveraging saturation mutagenesis experiments to construct accurate computational predictors of proteome-wide missense variant pathogenicity. We train cross-protein transfer (CPT) models using deep mutational scanning data from only five proteins and achieve state-of-the-art performance on clinical variant interpretation for unseen proteins across the human proteome. High sensitivity is crucial for clinical applications and our model CPT-1 particularly excels in this regime. For instance, at 95% sensitivity of detecting human disease variants annotated in ClinVar, CPT-1 improves specificity to 68%, from 27% for ESM-1v and 55% for EVE. Furthermore, for genes not used to train REVEL, a supervised method widely used by clinicians, we show that CPT-1 compares favorably with REVEL. Our framework combines predictive features derived from general protein sequence models, vertebrate sequence alignments, and AlphaFold2 structures, and it is adaptable to the future inclusion of other sources of information. We find that vertebrate alignments, albeit rather shallow with only 100 genomes, provide a strong signal for variant pathogenicity prediction that is complementary to recent deep learning-based models trained on massive amounts of protein sequence data. We release predictions for all possible missense variants in 90% of human genes. Our results demonstrate the utility of mutational scanning data for learning properties of variants that transfer to unseen proteins.

genetics↗

The Impact of Protein Dynamics on Residue-Residue Coevolution and Contact Prediction

The need to maintain protein structure constrains evolution at the sequence level, and patterns of coevolution in homologous protein sequences can be used to predict their 3D structures with high accuracy. Our understanding of the relationship between protein structure and evolution has traditionally been benchmarked by computational models ability to predict contacts from a single representative, experimentally determined structure per protein family. However, proteins in vivo are highly dynamic and can adopt multiple functionally relevant conformations. Here we demonstrate that interactions that stabilize alternate conformations, as well those that mediate conformational changes, impose an underappreciated but significant set of evolutionary constraints. We analyze the extent of these constraints over 56 paralogous G protein coupled receptors (GPCRs), {beta}-arrestin and the human SARS-CoV2 receptor ACE2. Specifically, we observe that contacts uniquely found in molecular dynamics (MD) simulation data and alternate-conformation crystal structures are successfully predicted by unsupervised language models. In GPCRs, adding these contacts as positives increases the percentage of top contacts classified as true positives, as predicted by a state-of-the-art language model, from 69% to 87%. Our results show that protein dynamics impose constraints on molecular evolution and demonstrate the ability of unsupervised language models to measure these constraints.

bioinformatics↗

Transferability of Geometric Patterns from Protein Self-Interactions to Protein-Ligand Interactions

There is significant interest in developing machine learning methods to model protein-ligand interactions but a scarcity of experimentally resolved protein-ligand structures to learn from. Protein self-contacts are a much larger source of structural data that could be leveraged, but currently it is not well understood how this data source differs from the target domain. Here, we characterize the 3D geometric patterns of protein self-contacts as probability distributions. We then present a flexible statistical framework to assess the transferability of these patterns to protein-ligand contacts. We observe that the level of transferability from protein self-contacts to protein-ligand contacts depends on contact type, with many contact types exhibiting high transferability. We then demonstrate the potential of leveraging information from these geometric patterns to aid in ligand pose-selection problems in protein-ligand docking. We publicly release our extracted data on geometric interaction patterns to enable further exploration of this problem.

biophysics↗