bioRxiv Science⌕ Search

Biology subjects

Lamb, K. D.

Publications and source records attributed to Lamb, K. D..

3 recordsLinked to original sources

PLM-interact: extending protein language models to predict protein-protein interactions

Computational prediction of protein structure from amino acid sequences alone has been achieved with unprecedented accuracy, yet the prediction of protein-protein interactions (PPIs) remains an outstanding challenge. Here we assess the ability of protein language models (PLMs), routinely applied to protein folding, to be retrained for PPI prediction. Existing PPI prediction models that exploit PLMs use a pre-trained PLM feature set, ignoring that the proteins are physically interacting. Our novel method, PLM-interact, goes beyond a single protein, jointly encoding protein pairs to learn their relationships, analogous to the next-sentence prediction task from natural language processing. This approach provides a significant improvement in performance: Trained on human-human PPIs, PLM-interact predicts mouse, fly, worm, E. coli and yeast PPIs, with 16-28% improvements in AUPR compared with state-of-the-art PPI models. Additionally, it can detect changes that disrupt or cause PPIs and be applied to virus-host PPI prediction. Our work demonstrates that large language models can be extended to learn the intricate relationships among biomolecules from their sequences alone.

bioinformatics↗

From a single sequence to evolutionary trajectories: protein language models capture the evolutionary potential of SARS-CoV-2 protein sequences

Protein language models (PLMs) capture features of protein three-dimensional structure from amino acid sequences alone, without requiring multiple sequence alignments (MSA). The concepts of grammar and semantics from natural language have the potential to capture functional properties of proteins. Here, we investigate how these representations enable assessment of variation due to mutation. Applied to SARS-CoV-2s spike protein using in silico deep mutational scanning (DMS) we demonstrate the PLM, ESM-2, has learned the sequence context within which variation occurs, capturing evolutionary constraint. This recapitulates what conventionally requires MSA data to predict. Unlike other state-of-the-art methods which require protein structures or multiple sequences for training, we show what can be accomplished using an unmodified pretrained PLM. We demonstrate that the grammaticality and semantic scores represent novel metrics. Applied to SARS-CoV-2 variants across the pandemic we show that ESM-2 representations encode the evolutionary history between variants, as well as the distinct nature of variants of concern upon their emergence, associated with shifts in receptor binding and antigenicity. PLM likelihoods can also identify epistatic interactions among sites in the protein. Our results here affirm that PLMs are broadly useful for variant-effect prediction, including unobserved changes, and can be applied to understand novel viral pathogens with the potential to be applied to any protein sequence, pathogen or otherwise.

bioinformatics↗

SARS-CoV-2's evolutionary capacity is mostly driven by host antiviral molecules

The COVID-19 pandemic has been characterised by sequential variant-specific waves shaped by viral, individual human and population factors. SARS-CoV-2 variants are defined by their unique combinations of mutations and there has been a clear adaptation to human infection since its emergence in 2019. Here we use machine learning models to identify shared signatures, i.e., common underlying mutational processes, and link these to the subset of mutations that define the variants of concern (VOCs). First, we examined the global SARS-CoV-2 genomes and associated metadata to determine how viral properties and public health measures have influenced the magnitude of waves, as measured by the number of infection cases, in different geographic locations using regression models. This analysis showed that, as expected, both public health measures and not virus properties alone are associated with the rise and fall of regional SARS-CoV-2 reported infection numbers. This impact varies geographically. We attribute this to intrinsic differences such as vaccine coverage, testing and sequencing capacity, and the effectiveness of government stringency. In terms of underlying evolutionary change, we used non-negative matrix factorisation to observe three distinct mutational signatures, unique in their substitution patterns and exposures from the SARS-CoV-2 genomes. Signatures 0, 1 and 3 were biased to C[->]T, T[->]C/A[->]G and G[->]T point mutations as would be expected of host antiviral molecules APOBEC, ADAR and ROS effects, respectively. We also observe a shift amidst the pandemic in relative mutational signature activity from predominantly APOBEC-like changes to an increasingly high proportion of changes consistent with ADAR editing. This could represent changes in how the virus and the host immune response interact, and indicates how SARS-CoV-2 may continue to accumulate mutations in the future. Linkage of the detected mutational signatures to the VOC defining amino acids substitutions indicates the majority of SARS-CoV-2s evolutionary capacity is likely to be associated with the action of host antiviral molecules rather than virus replication errors.

genomics↗