bioRxiv Science⌕ Search

Biology subjects

Ljubetic, A.

Publications and source records attributed to Ljubetic, A..

4 recordsLinked to original sources

Multivalent designed proteins protect against SARS-CoV-2 variants of concern

Escape variants of SARS-CoV-2 are threatening to prolong the COVID-19 pandemic. To address this challenge, we developed multivalent protein-based minibinders as potential prophylactic and therapeutic agents. Homotrimers of single minibinders and fusions of three distinct minibinders were designed to geometrically match the SARS-CoV-2 spike (S) trimer architecture and were optimized by cell-free expression and found to exhibit virtually no measurable dissociation upon binding. Cryo-electron microscopy (cryoEM) showed that these trivalent minibinders engage all three receptor binding domains on a single S trimer. The top candidates neutralize SARS-CoV-2 variants of concern with IC50 values in the low pM range, resist viral escape, and provide protection in highly vulnerable human ACE2-expressing transgenic mice, both prophylactically and therapeutically. Our integrated workflow promises to accelerate the design of mutationally resilient therapeutics for pandemic preparedness. One-Sentence SummaryWe designed, developed, and characterized potent, trivalent miniprotein binders that provide prophylactic and therapeutic protection against emerging SARS-CoV-2 variants of concern.

synthetic biology↗

Interpreting Neural Networks for Biological Sequences by Learning Stochastic Masks

Sequence-based neural networks can learn to make accurate predictions from large biological datasets, but model interpretation remains challenging. Many existing feature attribution methods are optimized for continuous rather than discrete input patterns and assess individual feature importance in isolation, making them ill-suited for interpreting non-linear interactions in molecular sequences. Building on work in computer vision and natural language processing, we developed an approach based on deep generative modeling - Scrambler networks - wherein the most salient sequence positions are identified with learned input masks. Scramblers learn to generate Position-Specific Scoring Matrices (PSSMs) where unimportant nucleotides or residues are scrambled by raising their entropy. We apply Scramblers to interpret the effects of genetic variants, uncover non-linear interactions between cis-regulatory elements, explain binding specificity for protein-protein interactions, and identify structural determinants of de novo designed proteins. We show that interpretation based on a generative model allows for efficient attribution across large datasets and results in high-quality explanations, often outperforming state-of-the-art methods.

genomics↗

Ensuring scientific reproducibility in bio-macromolecular modeling via extensive, automated benchmarks

Each year vast international resources are wasted on irreproducible research. The scientific community has been slow to adopt standard software engineering practices, despite the increases in high-dimensional data, complexities of workflows, and computational environments. Here we show how scientific software applications can be created in a reproducible manner when simple design goals for reproducibility are met. We describe the implementation of a test server framework and 40 scientific benchmarks, covering numerous applications in Rosetta bio-macromolecular modeling. High performance computing cluster integration allows these benchmarks to run continuously and automatically. Detailed protocol captures are useful for developers and users of Rosetta and other macromolecular modeling tools. The framework and design concepts presented here are valuable for developers and users of any type of scientific software and for the scientific community to create reproducible methods. Specific examples highlight the utility of this framework and the comprehensive documentation illustrates the ease of adding new tests in a matter of hours.

bioinformatics↗

A Multiplexed Bacterial Two-Hybrid for Rapid Characterization of Protein-Protein Interactions and Iterative Protein Design

Myriad biological functions require protein-protein interactions (PPIs), and engineered PPIs are crucial for applications ranging from drug design to synthetic cell circuits. Understanding and engineering specificity in PPIs is particularly challenging as subtle sequence changes can drastically alter specificity. Coiled-coils are small protein domains that have long served as a simple model for studying the sequence-determinants of specificity and have been used as modular building blocks to build large protein nanostructures and synthetic circuits. Despite their simple rules and long-time use, building large sets of well-behaved orthogonal pairs that can be used together is still challenging because predictions are often inaccurate, and, as the library size increases, it becomes difficult to test predictions at scale. To address these problems, we first developed a method called the Next-Generation Bacterial Two-Hybrid (NGB2H), which combines gene synthesis, a bacterial two-hybrid assay, and a high-throughput next-generation sequencing readout, allowing rapid exploration of interactions of programmed protein libraries in a quantitative and scalable way. After validating the NGB2H system on previously characterized libraries, we designed, built, and tested large sets of orthogonal synthetic coiled-coils. In an iterative set of experiments, we assayed more than 8,000 PPIs, used the dataset to train a novel linear model-based coiled-coil scoring algorithm, and then characterized nearly 18,000 interactions to identify the largest set of orthogonal PPIs to date with twenty-two on-target interactions.

molecular biology↗