bioRxiv Science⌕ Search

Biology subjects

Rancati, S.

Publications and source records attributed to Rancati, S..

2 recordsLinked to original sources

SARITA: A Large Language Model for Generating the S1 Subunit of the SARS-CoV-2 Spike Protein

The COVID-19 pandemic has profoundly impacted global health, economics, and daily life, with over 776 million cases and 7 million deaths from December 2019 to November 2024. Since the original SARS-CoV-2 Wuhan strain emerged, the virus has evolved into variants such as Alpha, Beta, Gamma, Delta, and Omicron, all characterized by mutations in the Spike glycoprotein, critical for viral entry into human cells via its S1 and S2 subunits. The S1 subunit, binding to the ACE2 receptor and mutating frequently, affects infectivity and immune evasion; the more conserved S2, on the other hand, facilitates membrane fusion. Predicting future mutations is crucial for developing vaccines and treatments adaptable to emerging strains, enhancing preparedness and intervention design. Generative Large Language Models (LLMs) are becoming increasingly common in the field of genomics, given their ability to generate realistic synthetic biological sequences, including applications in protein design and engineering. Here we present SARITA, an LLM with up to 1.2 billion parameters, based on GPT-3 architecture, designed to generate high-quality synthetic SARS-CoV-2 Spike S1 sequences. SARITA is trained via continuous learning on the pre-existing protein model RITA. When trained on Alpha, Beta, and Gamma variants (data up to February 2021 included), SARITA correctly predicts the evolution of future S1 mutations, including characterized mutations of Delta, Omicron and Iota variants. Furthermore, we show how SARITA outperforms alternative approaches, including other LLMs, in terms of sequence quality, realism, and similarity with real-world S1 sequences. These results indicate the potential of SARITA to predict future SARS-CoV-2 S1 evolution, potentially aiding in the development of adaptable vaccines and treatments.

bioinformatics↗

Forecasting dominance of SARS-CoV-2 lineages by anomaly detection using deep AutoEncoders

The coronavirus disease of 2019 (COVID-19) pandemic is characterized by sequential emergence of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) variants, lineages, and sublineages, outcompeting previously circulating ones because of, among other factors, increased transmissibility and immune escape. We propose DeepAutoCoV, an unsupervised deep learning anomaly detection system to predict future dominant lineages (FDLs). We define FDLs as viral (sub)lineages that will constitute more than 10% of all the viral sequences added to the GISAID database on a given week. DeepAutoCoV is trained and validated by assembling global and country-specific data sets from over 16 million Spike protein sequences sampled over a period of about 4 years. DeepAutoCoV successfully flags FDLs at very low frequencies (0.01% - 3%), with median lead times of 4-17 weeks, and predicts FDLs [~]5 and [~]25 times better than a baseline approach For example, the B.1.617.2 vaccine reference strain was flagged as FDL when its frequency was only 0.01%, more than a year before it was considered for an updated COVID-19 vaccine. Furthermore, DeepAutoCoV outputs interpretable results by pinpointing specific mutations potentially linked to increased fitness, and may provide significant insights for the optimization of public health pre-emptive intervention strategies. Key PointsO_LIIntroduction of DeepAutoCoV: The article introduces DeepAutoCoV, an unsupervised deep learning anomaly detection system designed to predict future dominant lineages (FDLs) of SARS-CoV-2. FDLs are defined as viral (sub)lineages that will constitute more than 10% of all viral sequences added to the GISAID database in a given week; C_LIO_LIPerformance and Predictive Capability: DeepAutoCoV successfully flags FDLs at very low frequencies (0.01% to 3%), with median lead times of 4 to 17 weeks before they become dominant. It predicts FDLs approximately 5 to 25 times better than baseline approaches. For instance, the B.1.617.2 vaccine reference strain was identified when its frequency was only 0.01%, over a year before it was considered for vaccine updates; C_LIO_LIInterpretable Results and Mutation Identification: The system provides interpretable results by pinpointing specific mutations that may be linked to increased fitness, offering insights that can optimize public health interventions. Key FDL mutations, such as those found in Delta and Omicron variants, are identified and analyzed for their potential impact on viral spread and immune escape; C_LIO_LIAdvantages and Applications: DeepAutoCoV is advantageous because it does not require prior assumptions about which protein sites are more likely to mutate. Its application in genomic surveillance systems could significantly reduce the time needed for public health responses to emerging variants, enabling early interventions such as vaccine updates; C_LIO_LIEvaluation and Comparisons: The performance of DeepAutoCoV was tested over four years of global and national surveillance data, demonstrating superior predictive power compared to other supervised and unsupervised methods. The system is periodically updated to adapt to the evolving viral landscape, making it a robust tool for ongoing surveillance efforts. C_LI

bioinformatics↗