bioRxiv Science⌕ Search

Biology subjects

Goverde, C. A.

Publications and source records attributed to Goverde, C. A..

5 recordsLinked to original sources

Structures of the Foamy virus fusion protein reveal an unexpected link with the F protein of paramyxo- and pneumoviruses

Foamy viruses (FVs) constitute a subfamily of retroviruses. Their envelope glycoprotein (Env) drives the merger of viral and cellular membranes during entry into cells. The only available structures of retroviral Envs are those from human and simian immunodeficiency viruses from the subfamily of orthoretroviruses, which are only distantly related to the FVs. We report here the cryo-EM structures of the FV Env ectodomain in the pre- and post-fusion states, which demonstrate structural similarity with the fusion protein (F) of paramyxo- and pneumoviruses, implying an evolutionary link between the two viral fusogens. Based on the structural information on the FV Env in two states, we propose a mechanistic model for its conformational change, highlighting how the interplay of its structural elements could drive the structural rearrangement. The structural knowledge on the FV Env now provides a framework for functional investigations such as the FV cell tropism and molecular features controlling the Env fusogenicity, which can benefit the design of FV Env variants with improved features for use as gene therapy vectors.

microbiology↗

AF2BIND: Predicting ligand-binding sites using the pair representation of AlphaFold2

Predicting ligand-binding sites, particularly in the absence of previously resolved homologous structures, presents a significant challenge in structural biology. Here, we leverage the internal pairwise representation of AlphaFold2 (AF2) to train a model, AF2BIND, to accurately predict small-molecule-binding residues given only a target protein. AF2BIND uses 20 "bait" amino acids to optimally extract the binding signal in the absence of a small-molecule ligand. We find that the AF2 pair representation outperforms other neural-network representations for binding-site prediction. Moreover, unique combinations of the 20 bait amino acids are correlated with chemical properties of the ligand.

bioinformatics↗

Exploring "dark matter" protein folds using deep learning

De novo protein design aims to explore uncharted sequence-and structure areas to generate novel proteins that have not been sampled by evolution. One of the main challenges in de novo design involves crafting "designable" structural templates that can guide the sequence search towards adopting the target structures. Here, we present an approach to learn patterns of protein structure based on a convolutional variational autoencoder, dubbed Genesis. We coupled Genesis with trRosetta to design sequences for a set of protein folds and found that Genesis is capable of reconstructing native-like distance-and angle distributions for five native folds and three novel, so-called "dark-matter" folds as a demonstration of generalizability. We used a high-throughput assay to characterize protease resistance of the designs, obtaining encouraging success rates for folded proteins and further biochemically characterized folded designs. The Genesis framework enables the exploration of the protein sequence and fold space within minutes and is not bound to specific protein topologies. Our approach addresses the backbone designability problem, showing that structural patterns in proteins can be efficiently learned by small neural networks and could ultimately contribute to the de novo design of proteins with new functions.

bioinformatics↗

An atlas of protein homo-oligomerization across domains of life

Protein structures are essential to understand cellular processes in molecular detail. While advances in AI revealed the tertiary structure of proteins at scale, their quaternary structure remains mostly unknown. Here, we describe a scalable strategy based on AlphaFold2 to predict homo-oligomeric assemblies across four proteomes spanning the tree of life. We find that 50% of archaeal, 45% of bacterial, and 20% of eukaryotic proteomes form homomers. Our predictions accurately capture protein homo-oligomerization, recapitulate megadalton complexes, and unveil hundreds of novel homo-oligomer types. Analyzing these datasets reveals coiled-coil regions as major enablers of quaternary structure evolution in Eukaryotes. Integrating these structures with omics data shows that a majority of known protein complexes are symmetric. Finally, these datasets provide a structural context for interpreting disease mutations, which we find enriched at interfaces. Our strategy is applicable to any organism and provides a comprehensive view of homo-oligomerization in proteomes, protein networks, and disease. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=193 SRC="FIGDIR/small/544317v1_ufig1.gif" ALT="Figure 1"> View larger version (79K): org.highwire.dtl.DTLVardef@1507a12org.highwire.dtl.DTLVardef@7e522aorg.highwire.dtl.DTLVardef@1445410org.highwire.dtl.DTLVardef@eb09f7_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Computational design of soluble analogues of integral membrane protein structures

De novo design of complex protein folds using solely computational means remains a significant challenge. Here, we use a robust deep learning pipeline to design complex folds and soluble analogues of integral membrane proteins. Unique membrane topologies, such as those from GPCRs, are not found in the soluble proteome and we demonstrate that their structural features can be recapitulated in solution. Biophysical analyses reveal high thermal stability of the designs and experimental structures show remarkable design accuracy. The soluble analogues were functionalized with native structural motifs, standing as a proof-of-concept for bringing membrane protein functions to the soluble proteome, potentially enabling new approaches in drug discovery. In summary, we designed complex protein topologies and enriched them with functionalities from membrane proteins, with high experimental success rates, leading to a de facto expansion of the functional soluble fold space.

bioinformatics↗