bioRxiv Science⌕ Search

Biology subjects

Spence, M. A.

Publications and source records attributed to Spence, M. A..

5 recordsLinked to original sources

Leveraging ancestral sequence reconstruction for protein representation learning

Protein language models (PLMs) convert amino acid sequences into the numerical representations required to train machine learning (ML) models. Many PLMs are large (>600 M parameters) and trained on a broad span of protein sequence space. However, these models have limitations in terms of predictive accuracy and computational cost. Here, we use multiplexed Ancestral Sequence Reconstruction (mASR) to generate small but focused functional protein sequence datasets for PLM training. Compared to large PLMs, this local ancestral sequence embedding (LASE) produces representations 10-fold faster and with higher predictive accuracy. We show that due to the evolutionary nature of the ASR data, LASE produces smoother fitness landscapes in which protein variants that are closer in fitness value become numerically closer in representation space. This work contributes to the implementation of ML-based protein design in real-world settings, where data is sparse and computational resources are limited.

bioinformatics↗

A broad-spectrum macrocyclic peptide inhibitor of the SARS-CoV-2 spike protein

The ongoing COVID-19 pandemic has had great societal and health consequences. Despite the availability of vaccines, infection rates remain high due to immune evasive Omicron sublineages. Broad-spectrum antivirals are needed to safeguard against emerging variants and future pandemics. We used mRNA display under a reprogrammed genetic code to find a spike-targeting macrocyclic peptide that inhibits SARS-CoV-2 Wuhan strain infection and pseudoviruses containing spike proteins of SARS-CoV-2 variants or related sarbecoviruses. Structural and bioinformatic analyses reveal a conserved binding pocket between the receptor binding domain, N-terminal domain and S2 region, distal to the ACE2 receptor-interaction site. Our data reveal a hitherto unexplored site of vulnerability in sarbecoviruses that peptides and potentially other drug-like molecules can target. Significance statementThis study reports on the discovery of a macrocyclic peptide that is able to inhibit SARS-CoV-2 infection by exploiting a new vulnerable site in the spike glycoprotein. This region is highly conserved across SARS-CoV-2 variants and the subgenus sarbecovirus. Due to the inaccessability and mutational contraint of this site, it is anticipated to be resistant to the development of resistance through antibody selective pressure. In addition to the discovery of a new molecule for development of potential new peptide or biomolecule therapeutics, the discovery of this broadly active conserved site can also stimulate a new direction of drug development, which together may prevent future outbreaks of related viruses.

biochemistry↗

The rugged DNA-binding sequence-fitness landscape of the LacI/GalR Family is a product of asymmetry in the operator:repressor complex

How a proteins function influences the shape of its fitness landscape, smooth or rugged, is a fundamental question in evolutionary biochemistry. Smooth landscapes arise when incremental mutational steps lead to a progressive change in function, as commonly seen in enzymes and binding proteins. On the other hand, rugged landscapes are poorly understood because of the inherent unpredictability of how sequence changes affect function. Here, we experimentally characterize the entire sequence phylogeny, comprising 1158 extant and ancestral sequences, of the DNA-binding domain (DBD) of the LacI/GalR transcriptional repressor family. Our analysis revealed an extremely rugged landscape with rapid switching of specificity even between adjacent nodes. Further, the ruggedness arises due to the necessity of the repressor to simultaneously evolve specificity for asymmetric operators and disfavors potentially adverse regulatory crosstalk. Our study provides fundamental insight into evolutionary, molecular, and biophysical rules of genetic regulation through the lens of fitness landscapes.

biochemistry↗

Comprehensive phylogenetic analysis of the ribonucleotide reductase family reveals an ancestral clade and the role of insertions and extensions in diversification

Ribonucleotide reductases (RNRs) are used by all organisms and many viruses to catalyze an essential step in the de novo biosynthesis of DNA precursors. RNRs are remarkably diverse by primary sequence and cofactor requirement, while sharing a conserved fold and radical-based mechanism for nucleotide reduction. Here, we structurally aligned the diverse RNR family by the conserved catalytic barrel to reconstruct the first large-scale phylogeny consisting of 6,779 sequences that unites all extant classes of the RNR family and performed evo-velocity analysis to independently validate our evolutionary model. With a robust phylogeny in-hand, we uncovered a novel, phylogenetically distinct clade that is placed as ancestral to the classes I and II RNRs, which we have termed clade O. We employed small-angle X-ray scattering (SAXS), cryogenic-electron microscopy (cryo-EM), and AlphaFold2 to investigate a member of this clade from Synechococcus phage S-CBP4 and report the most minimal RNR architecture to-date. Using the catalytic barrel as a starting point for diversification, we traced the evolutionarily relatedness of insertions and extensions that confer the diversity observed in the RNR family. Based on our analyses, we propose an evolutionary model of diversification in the RNR family and delineate how our phylogeny can be used as a roadmap for targeted future study.

bioinformatics↗

A comprehensive phylogenetic analysis of the serpin superfamily

Serine protease inhibitors (serpins) are found in all kingdoms of life and play essential roles in multiple physiological processes. Owing to the diversity of the superfamily, phylogenetic analysis is challenging and prokaryotic serpins have been speculated to have been acquired from Metazoa through horizontal gene transfer (HGT) due to their unexpectedly high homology. Here we have leveraged a structural alignment of diverse serpins to generate a comprehensive 6000-sequence phylogeny that encompasses serpins from all kingdoms of life. We show that in addition to a central "hub" of highly conserved serpins, there has been extensive diversification of the superfamily into many novel functional clades. Our analysis indicates that the hub proteins are ancient and are similar because of convergent evolution, rather than the alternative hypothesis of HGT. This work clarifies longstanding questions in the evolution of serpins and provides new directions for research in the field of serpin biology.

bioinformatics↗