bioRxiv Science⌕ Search

Biology subjects

Dudzic, P.

Publications and source records attributed to Dudzic, P..

10 recordsLinked to original sources

Protein Design Viz (PDV): lightweight protein structure visualization studio with validated quantitative analytics and antibody-specific toolkit

Motivation: Molecular visualization is dominated by two families of software. Desktop programs such as PyMOL, VMD and UCSF ChimeraX are powerful but heavy to install and operate. Web viewers such as Jmol, 3Dmol.js, NGL, Mol* and iCn3D are light and installation-free, but typically stop at rendering, lacking deeper functionality needed for even basic protein sequence, structural analyses and design tasks. Tools building upon these often lack comprehensive sequence/structure manipulation, surface, interface and antibody-specific analytics functionality or require an external backend. There is a need for a tool that is as frictionless as a web viewer yet carries the analytical depth normally reserved for the desktop or the command line software. Results: We present Protein Design Viz (PDV), a self-contained molecular visualization studio delivered as a single offline HTML file that runs entirely in the browser. Built as an extensive modification of 3Dmol.js, PDV combines a full visualization workflow: multi-object scenes, representations, coloring palette, linked sequence track, publication-quality outline rendering and portable sessions. Additionally we re-implemented four commonly used macromolecular analyses measures from scratch in client-side JavaScript: a Shrake-Rupley solvent-accessible surface area (SASA) engine, an antibody numbering and germline-assignment engine, a non-covalent interaction detector, and a developability-liability scanner. Each engine is validated against its established reference. PDV's SASA reproduces FreeSASA at Pearson r {approx} 0.997-0.998 across 2,582 structures spanning proteins, nucleic acids and ligands; its numbering reproduces RIOT for over 99.88% of 1.3 million residue positions across 16,996 sequences and four schemes; its interaction detector reproduces PLIP at macro-F1 0.82, matching Arpeggio as closely as PLIP itself does. PDV brings validated, quantitative structural analysis into a no-install, simple to use tool. Availability and implementation: PDV is a single HTML file, free for noncommercial use under the PolyForm Noncommercial License 1.0.0, available from pdv.naturalantibody.com. It requires only a WebGL-capable browser and runs fully offline.

bioinformatics↗

ASD: Antigen-Specific Antibody Database

The development of computational models addressing therapeutic antibodies faces significant challenges due to the scarcity of data. A critical data element is the set of antibody-antigen interaction pairs associated with sequences. To address this issue, we developed the Antigen Specific Antibody Database (ASD, https://naturalantibody.com/agab/), a database aggregating antibody-antigen interaction data from multiple studies with standardized formatting and annotations. Our dataset compilation strategy resulted in data from 15 distinct sources, resulting in 1,097,946 unique antibody-antigen interactions (with 9,575 unique antigens). The ASD database captures diverse affinity measures and qualitative binding assessment, along with metadata including UniProt and PDB identifiers, target protein names, confidence levels, and experimental conditions. Through this integration drive, we make available an ample resource of interaction data gathered from the public domain to act as a foundation for model development and further data generation.

bioinformatics↗

NAStructuralDB : Structural database to facilitate computational studies of molecular modeling and recognition of proteins with special focus on antibody-antigen interactions.

Studying the interactions between antibodies and antigens is fundamental to the development of novel therapeutic biologics. Predictions of such interactions start with data collection. Though there exist reliable resources to identify antibody structures in the Protein Data Bank (PDB), such data still requires substantial processing to be usable in predictive tasks. Redundancy in sequences needs to be removed to avoid data leakages between train, test and validation sets. Descriptors such as surface accessibility, secondary structure and antibody region information need to be additionally annotated. Information on inter- and intra-molecular contacts, which is crucial to studying paratope/epitope information, needs to be collected. The specialized immunoglobulin format of Nanobodies(R) requires a separate dataset mirroring that of antibodies, given that their structure contains only a single VHH chain. Because antibody-antigen structures account for a small amount of all protein-protein contacts, having a molecular contact reference from other proteins is also desired. To address these issues, we introduce NAStructuralDB (https://naturalantibody.com/na-structural/), a dataset of processed structures of antibodies, Nanobodies(R), proteins and their complexes with molecular contact information and associated annotations. We use the opportunity of having collected the contact data to provide a reference of binding propensities of different residues across distinct contact types. We anticipate that this dataset will accelerate a broad range of predictive tasks by standardizing common, time-consuming data preparation steps in antibody and protein design.

bioinformatics↗

Herding cats: predicting immunogenicity from heterogeneous clinical trials data

Antibodies represent the largest and fastest growing class of biologic therapeutics, yet forecasting their clinical performance, particularly immunogenicity, remains a major hurdle in drug development. Despite hundreds of antibody-based drugs progressing through clinical pipelines, systematic integration of their clinical outcomes has been limited by fragmented and heterogeneous data. Here, we present the Therapeutic Antibody Database, a comprehensive and curated resource that links therapeutic antibodies to clinical trial outcomes, with a dedicated focus on immunogenicity. Our dataset is sourced from approximately 11,500 anti-drug antibody (ADA) measurements across diverse molecules and indications, offering an unprecedented view into the clinical manifestation of immune responses to biologics. In order to evaluate the main drivers of ADA, we evaluate gathered immunogenicity incidence and prevalence data against various therapeutic descriptors which includes sequence, structure and contextual features related to therapeutics. We find that most tools have very poor performance, and we pinpoint the causes of it, demonstrating the need for systems immunology approaches incorporating clinical metadata beyond biochemical properties of the molecules alone.

bioinformatics↗

Benchmarking antigen-aware inverse folding methods for antibody design.

Computational antibody design has seen many recent advances pioneered via the use of language models and advanced structure prediction tools. Developing a de novo antibody against a specific antigen requires structural awareness that most language models lack. A prominent class of machine learning methods combining the best of language model and structural worlds is inverse folding. This approach aims to predict a sequence that would fit a given structure. Such methods are now increasingly used to predict alternate sequences given a structure of a binder. It is known that, just like language models, such methods have certain predictive power in identifying binders. Here we performed a set of tests to reveal where, if at all, such methods provide value in the realistic setting of antibody discovery.

bioinformatics↗

AbDesign: Database of point mutants of antibodies with associated structures reveals poor generalization of binding predictions from machine learning models.

Antibodies are naturally evolved molecular recognition scaffolds that can bind a variety of surfaces. Their designability is crucial to the development of biologics, with computational methods holding promise in accelerating the delivery of medicines to the clinic. Modeling antibody-antigen recognition is prohibitively difficult, with data paucity being one of the biggest hurdles. Current affinity datasets comprise a small number of experimental measurements, which are often not standardized between molecules. Here, we address these issues by creating a dataset of seven antigens with two antibodies each, for which we introduce a heterogeneous set of mutations to the CDR-H3 measured by ELISA. Each of the parental complexes has a known crystal structure. We perform benchmarking of state-of-the-art affinity prediction algorithms to gauge their effectiveness. Current computational methods exhibit significant limitations in accurately predicting the effects of single-point mutations. In contrast, the older empirical, physics-based method FoldX, performs well in identifying mutants that retain binding. These findings highlight the need for more resources like the one presented here -- large, molecularly diverse, and experimentally consistent datasets.

bioinformatics↗

nanoFOLD : sequence design of nanobodies via inverse folding

Antibodies devoid of light chains are a promising class of biotherapeutics. Computational methods that address these molecules are crucially needed to accelerate the traditional, long and expensive experimental process of their discovery. Inverse folding, wherein one is tasked to predict a sequence given molecular coordinates, is an established method in scaffold-based protein design. Here we develop an inverse folding method speci[fi]c to nanobodies. We demonstrate its application in nanobody-engineering scenarios of enriching binders from next-generation sequencing experiments and novel binder design.

bioinformatics↗

Conserved heavy/light contacts and germline preferences revealed by a large-scale analysis of natively paired human antibody sequences and structural data.

Antibody next-generation sequencing (NGS) datasets have become crucial to develop computational models addressing this successful class of therapeutics. Although antibodies are composed of both heavy and light chains, most NGS sequencing depositions provide them in unpaired form, reducing their utility. Here we introduce PairedAbNGS, a novel database with paired heavy/light antibody chains. To the best of our knowledge, this is the largest resource for paired natural antibody sequences with 58 bioprojects and over 14 million assembled productive sequences. We make the database accessible at http://naturalantibody.com/paired-ngs as a valuable tool for biological and machine-learning applications. Using this dataset, we investigated heavy and light chain variable (V) gene pairing preferences and found significant biases beyond gene usage frequencies, possibly due to receptor editing favoring less autoreactive combinations. Analyzing the available antibody structures from the Protein Data Bank, we studied conserved contact residues between heavy and light chains, particularly interactions between the CDR3 region of one chain and the FWR2 region of the opposite chain. Examination of amino acid pairs at key contact sites revealed significant deviations of amino acids distributions compared to random pairings, in the heavy chains CDR3 region contacting the opposite chain, indicating specific interactions might be crucial for proper chain pairing. This observation is further reinforced by preferential IGHV-IGLJ and IGLV-IGHJ pairing preferences. We hope that both our resources and the findings would contribute to improving the engineering of biological drugs.

immunology↗

RIOT - Rapid Immunoglobulin Overview Tool - annotation of nucleotide and amino acid immunoglobulin sequences using an open germline database.

Antibodies are a cornerstone of the immune system, playing a pivotal role in identifying and neutralizing infections caused by bacteria, viruses, and other pathogens. Understanding their structure, and function, can provide insights into both the bodys natural defenses and the principles behind many therapeutic interventions, including vaccines and antibody-based drugs. The analysis and annotation of antibody sequences, including the identification of variable, diversity, joining, and constant genes, as well as the delineation of framework regions and complementarity-determining regions, is essential for understanding their structure and function. Currently analyzing large volumes of antibody sequences is routine in antibody discovery, requiring fast and accurate tools. While there are existing tools designed for the annotation and numbering of antibody sequences, they often have limitations such as being restricted to either nucleotide or amino acid sequences, reliance on non-uniform germline databases, or slow execution times. Here we present Rapid Immunoglobulin Overview Tool (RIOT), a novel open-source solution for antibody numbering that addresses these shortcomings. RIOT handles nucleotide and amino acid sequence processing, comes with a free germline database, and is computationally efficient. We hope the tool will facilitate rapid annotation of antibody sequencing outputs for the benefit of understanding antibody biology and discovering novel therapeutics. AvailabilityRIOT is available at https://github.com/NaturalAntibody/riot_na.

bioinformatics↗

Structural pre-training improves physical accuracy of antibody structure prediction using deep learning.

Protein folding problem obtained a practical solution recently, owing to advances in deep learning. There are classes of proteins though, such as antibodies, that are structurally unique, where the general solution still lacks. In particular, the prediction of the CDR-H3 loop, which is an instrumental part of an antibody in its antigen recognition abilities, remains a challenge. Antibody-specific deep learning frameworks were proposed to tackle this problem noting great progress, both on accuracy and speed fronts. Oftentimes though, the original networks produce physically implausible bond geometries that then need to undergo a time-consuming energy minimization process. Here we hypothesized that pre-training the network on a large, augmented set of models with correct physical geometries, rather than a small set of real antibody X-ray structures, would allow the network to learn better bond geometries. We show that fine-tuning such a pre-trained network on a task of shape prediction on real X-ray structures improves the number of correct peptide bond distances. We further demonstrate that pre-training allows the network to produce physically plausible shapes on an artificial set of CDR-H3s, showing the ability to generalize to the vast antibody sequence space. We hope that our strategy will benefit the development of deep learning antibody models that rapidly generate physically plausible geometries, without the burden of time-consuming energy minimization.

bioinformatics↗