bioRxiv Science⌕ Search

bioRxiv · 10.1101/2022.04.28.489618

A Pan-Coronavirus Vaccine Candidate: Nine Amino Acid Substitutions in the ORF1ab Gene Attenuate 99% of 365 Unique Coronaviruses: A Comparative Effectiveness Research Study

Abstract

BackgroundThe COVID-19 pandemic has been a watershed event. Industry and governments have reacted, investing over US$105 billion in vaccine research.1 The Holy Grail is a universal, pan-coronavirus, vaccine to protect humankind from future SARS-CoV-2 variants and the thousands of similar coronaviruses with pandemic potential.2 This paper proposes a new vaccine candidate that appears to attenuate the SARS-Cov-2 coronavirus variants to render it safe to use as a vaccine. Moreover, these results indicate it may be efficacious against 99% of 365 coronaviruses. This research model is wet-dry-wet; it originated in genomic sequencing laboratories, evolved to computational modeling, and the candidate result now require validation back in a wet lab. ObjectivesThis studys purpose was to test the hypothesis that machine learning applied to sequenced coronaviruses genomes could identify which amino acid substitutions likely attenuate the viruses to produce a safe and effective pan-coronavirus vaccine candidate. This candidate is now eligible to be pre-clinically then clinically tested and proven. If validated, it would constitute a traditional attenuated virus vaccine to protect against hundreds of coronaviruses, including the many future variants of SARS-CoV-2 predicted from continuously recombining in unvaccinated populations and spreading by modern mass travel. MethodsUsing machine learning, this was an in silico comparative effectiveness research study on trinucleotide functions in nonstructural proteins of 365 novel coronavirus genomes. Sequences of 7,097 codons in the ORF1ab gene were collected from 65 global locations infecting 68 species and reported to the US National Institute of Health. The data were proprietarily transformed twice to enable machine learning ingestion, mapping, and interpretation. The set of 2,590,405 data points was randomly divided into three cohorts: 255 (70%) observations for training; and two cohorts of 55 (15%) observations each for testing. Machine learning models were trained in the statistical programming language R and compared to identify which mixture of the 7.097 x 1023 possible amino-acid-location combinations would attenuate SARS-CoV-2 and other coronaviruses that have infected humans. ResultsContests of machine-learning algorithms identified nine amino-acid point substitutions in the ORF1ab gene that likely attenuate 98.98% of 365 (361) novel coronaviruses. Notably, seven substitutions are for the amino acid alanine. Most of the locations (5 of 9) are in nonstructural proteins (NSPs) 2 and 3. The substitutions are alanine to (1) valine at codon 4273; (2) leucine at codon 5077; (3) phenylalanine at codon 2001; (4) leucine at codon 372; (5) proline at codon 354; (6) phenylalanine at codon 2811; (7) phenylalanine at codon 4703; (8) leucine to serine at codon 2333; and, (9) threonine to alanine at codon 5131. ConclusionsThe primary outcome is a new, highly promising, pan-coronavirus vaccine candidate based on nine amino-acid substitutions in the ORF1ab gene. The secondary outcome was evidence that sequences of wet-dry lab collaborations - here machine learning analysis of viral genomes informing codon functions -- may discover new broader and more stable vaccines candidates more quickly and inexpensively than traditional methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Luellen, E.. 2022-04-28. A Pan-Coronavirus Vaccine Candidate: Nine Amino Acid Substitutions in the ORF1ab Gene Attenuate 99% of 365 Unique Coronaviruses: A Comparative Effectiveness Research Study. https://doi.org/10.1101/2022.04.28.489618

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Living electronic transistors with tunable conductivity

Electroactive bacteria, like Shewanella oneidensis, can couple the oxidation of organic electron donors to the reduction of external conductive surfaces, such as minerals and electrodes, by utilizing multiheme cytochromes to carry charge from within the cell to external surfaces. Additionally, multiheme cytochromes facilitate gateable, long-distance (micrometer-scale) redox conduction along the outer membrane and across multiple cells bridging electrodes. While electroactive microbes are being used to develop bioelectrochemical devices, there have been limited efforts to use synthetic biology to exert additional control over microbes serving as device components. Thus, this work implements an optogenetic biofilm patterning gene circuit and a small molecule sensor in S. oneidensis to simultaneously control cell deposition and cytochrome expression. This allows for photolithographic patterning of biofilms possessing tunable electrical properties controlled with small molecules. This system demonstrates tunable electrochemical activity, redox conduction, intrinsic biofilm conductivity, and negative differential transconductance as a function of cytochrome expression. Additionally, temperature-dependent measurements of this tunable biofilm conduction reveal changes in activation energy as a function of cytochrome expression. Through this combination of synthetic biology and electrochemistry, simultaneous control over biofilm geometry and conductivity sheds light on fundamental microbial electron transport processes, and it enables the construction of living electronic devices.

synthetic biology↗

Evolutionary stabilisation of stressful metabolism via integrated biocomputing and essential-gene metabolic locking circuits

Synthetic genetic circuits enable microbial differentiation from growth to production, yet metabolic burden, imbalance and toxicity frequently drive strain degeneration. Yeast strains engineered to produce different terpene products exhibited divergent genetic responses to metabolic stresses, but commonly underwent progressive loss of induction of synthetic GAL regulatory circuits, either across the entire population or within subpopulations. Using di- and tri-input biocomputing circuits, the essential glutamine synthetase gene GLN1 was coupled to GAL induction, thereby enabling stabilisation and evolutionary adaptation of the synthetic genetic circuits and stressful heterologous terpene synthetic pathways. The integrated biocomputing and metabolic coupling circuit systems not only prevent strain degeneration but also enable interrogation of non-degenerative evolutionary shifts, providing a platform for metabolic engineering optimisation.

synthetic biology↗

Unbiased and scalable reduction of diverse bacterial genomes

The genome is a complex, integrated system where the functions and regulatory interactions of its many components remain poorly understood. Genome minimization aims to reduce genomic complexity by removing non-essential elements to reveal the fundamental building blocks of cellular life. However, current minimization strategies are often slow and species-specific due to a reliance on prior information, and limited to producing single, isolated strains, which obscures the diverse ways a genome can adapt to large-scale DNA removal. Here we show the development and application of Stochastic Lineage-based Iterative Minimization (SLIM) a modular, high-throughput platform for unbiased genome reduction across phylogenetically diverse bacteria. We apply SLIM to generate a library of genome-reduced Escherichia coli lineages. We then interrogate the lineages, identifying both universal and lineage-specific transcriptional and translational reprogramming in response to deletions. We demonstrate that these expression dynamics drive environment-dependent fitness, allowing us to pinpoint a single gene deletion in one genome-reduced lineage as the driver of a measurable environmental growth defect. Beyond E. coli, we successfully deploy SLIM in phylogenetically distinct bacterial taxa to rapidly reduce the genomes of Shigella flexneri and Pseudomonas putida, distinct genus and order respectively from E. coli, without species-specific optimization. Our results establish a scalable, generalizable framework for navigating the vast landscape of minimized genomes, providing a powerful new tool for functional discovery and the rational design of synthetic genomic chassis.

synthetic biology↗