bioRxiv Science⌕ Search

Biology subjects

Beguir, K.

Publications and source records attributed to Beguir, K..

5 recordsLinked to original sources

The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics

Closing the gap between measurable genetic information and observable traits is a longstanding challenge in genomics. Yet, the prediction of molecular phenotypes from DNA sequences alone remains limited and inaccurate, often driven by the scarcity of annotated data and the inability to transfer learning between prediction tasks. Here, we present an extensive study of foundation models pre-trained on DNA sequences, named the Nucleotide Transformer, ranging from 50M up to 2.5B parameters and integrating information from 3,202 diverse human genomes, as well as 850 genomes selected across diverse phyla, including both model and non-model organisms. These transformer models yield transferable, context-specific representations of nucleotide sequences, which allow for accurate molecular phenotype prediction even in low-data settings. We show that the developed models can be fine-tuned at low cost and despite low available data regime to solve a variety of genomics applications. Despite no supervision, the transformer models learned to focus attention on key genomic elements, including those that regulate gene expression, such as enhancers. Lastly, we demonstrate that utilizing model representations can improve the prioritization of functional genetic variants. The training and application of foundational models in genomics explored in this study provide a widely applicable stepping stone to bridge the gap of accurate molecular phenotype prediction from DNA sequence. Code and weights available on GitHub in Jax and HuggingFace in Pytorch. Example notebooks to apply these models to any downstream task are available on HuggingFace.

genomics↗

Progressive loss of conserved spike protein neutralizing antibody sites in Omicron sublineages is balanced by preserved T-cell recognition epitopes

The continued evolution of the SARS-CoV-2 Omicron variant has led to the emergence of numerous sublineages with different patterns of evasion from neutralizing antibodies. We investigated neutralizing activity in immune sera from individuals vaccinated with SARS-CoV-2 wild-type spike (S) glycoprotein-based COVID-19 mRNA vaccines after subsequent breakthrough infection with Omicron BA.1, BA.2, or BA.4/BA.5 to study antibody responses against sublineages of high relevance. We report that exposure of vaccinated individuals to infections with Omicron sublineages, and especially with BA.4/BA.5, results in a boost of Omicron BA.4.6, BF.7, BQ.1.1, and BA.2.75 neutralization, but does not efficiently boost neutralization of sublineages BA.2.75.2 and XBB. Accordingly, we found in in silico analyses that with occurrence of the Omicron lineage a large portion of neutralizing B-cell epitopes were lost, and that in Omicron BA.2.75.2 and XBB less than 12% of the wild-type strain epitopes are conserved. In contrast, HLA class I and class II presented T-cell epitopes in the S glycoprotein were highly conserved across the entire evolution of SARS-CoV-2 including Alpha, Beta, and Delta and Omicron sublineages, suggesting that CD8+ and CD4+ T-cell recognition of Omicron BQ.1.1, BA.2.75.2, and XBB may be largely intact. Our study suggests that while some Omicron sublineages effectively evade B-cell immunity by altering neutralizing antibody epitopes, S protein-specific T-cell immunity, due to the very nature of the polymorphic cell-mediated immune, response is likely to remain unimpacted and may continue to contribute to prevention or limitation of severe COVID-19 manifestation.

immunology↗

Peptide-MHC Structure Prediction With Mixed Residue and Atom Graph Neural Network

Neoantigen-targeting vaccines have achieved breakthrough success in cancer immunotherapy by eliciting immune responses against neoantigens, which are proteins uniquely produced by cancer cells. During the immune response, the interactions between peptides and major histocompatibility complexes (MHC) play an important role as peptides must be bound and presented by MHC to be recognised by the immune system. However, only limited experimentally determined peptide-MHC (pMHC) structures are available, and in-silico structure modelling is therefore used for studying their interactions. Current approaches mainly use Monte Carlo sampling and energy minimisation, and are often computationally expensive. On the other hand, the advent of large high-quality proteomic data sets has led to an unprecedented opportunity for deep learning-based methods with pMHC structure prediction becoming feasible with these trained protein folding models. In this work, we present a graph neural network-based model for pMHC structure prediction, which takes an amino acid-level pMHC graph and an atomic-level peptide graph as inputs and predicts the peptide backbone conformation. With a novel weighted reconstruction loss, the trained model achieved a similar accuracy to AlphaFold 2, requiring only 1.7M learnable parameters compared to 93M, representing a more than 98% reduction in the number of required parameters.

bioinformatics↗

So ManyFolds, So Little Time: Efficient Protein Structure Prediction With pLMs and MSAs

In recent years, machine learning approaches for de novo protein structure prediction have made significant progress, culminating in AlphaFold which approaches experimental accuracies in certain settings and heralds the possibility of rapid in silico protein modelling and design. However, such applications can be challenging in practice due to the significant compute required for training and inference of such models, and their strong reliance on the evolutionary information contained in multiple sequence alignments (MSAs), which may not be available for certain targets of interest. Here, we first present a streamlined AlphaFold architecture and training pipeline that still provides good performance with significantly reduced computational burden. Aligned with recent approaches such as OmegaFold and ESMFold, our model is initially trained to predict structure from sequences alone by leveraging embeddings from the pretrained ESM-2 protein language model (pLM). We then compare this approach to an equivalent model trained on MSA-profile information only, and find that the latter still provides a performance boost - suggesting that even state-of-the-art pLMs cannot yet easily replace the evolutionary information of homologous sequences. Finally, we train a model that can make predictions from either the combination, or only one, of pLM and MSA inputs. Ultimately, we obtain accuracies in any of these three input modes similar to models trained uniquely in that setting, whilst also demonstrating that these modalities are complimentary, each regularly outperforming the other.

bioinformatics↗

Early Computational Detection of Potential High Risk SARS-CoV-2 Variants

The ongoing COVID-19 pandemic is leading to the discovery of hundreds of novel SARS-CoV-2 variants on a daily basis. While most variants do not impact the course of the pandemic, some variants pose a significantly increased risk when the acquired mutations allow better evasion of antibody neutralisation in previously infected or vaccinated subjects or increased transmissibility. Early detection of such high risk variants (HRVs) is paramount for the proper management of the pandemic. However, experimental assays to determine immune evasion and transmissibility characteristics of new variants are resource-intensive and time-consuming, potentially leading to delays in appropriate responses by decision makers. Here we present a novel in silico approach combining spike (S) protein structure modelling and large protein transformer language models on S protein sequences to accurately rank SARS-CoV-2 variants for immune escape and fitness potential. These metrics can be combined into an automated Early Warning System (EWS) capable of evaluating new variants in minutes and risk-monitoring variant lineages in near real-time. The system accurately pinpoints the putatively dangerous variants by selecting on average less than 0.3% of the novel variants each week. With only the S protein nucleotide sequence as input, the EWS detects HRVs earlier and with better precision than baseline metrics such as the growth metric (which requires real-world observations) or random sampling. Notably, Omicron BA.1 was flagged by the EWS on the day its sequence was made available. Additionally, our immune escape and fitness metrics were experimentally validated using in vitro pseudovirus-based virus neutralisation test (pVNT) assays and binding assays. The EWS flagged as potentially dangerous all 16 variants (Alpha-Omicron BA.1/2/4/5) designated by the World Health Organisation (WHO) with an average lead time of more than one and a half months ahead of them being designated as such. One-Sentence SummaryA COVID-19 Early Warning System combining structural modelling with machine learning to detect and monitor high risk SARS-CoV-2 variants, identifying all 16 WHO designated variants on average more than one and a half months in advance by selecting on average less than 0.3% of the weekly novel variants.

bioinformatics↗