bioRxiv Science⌕ Search

Biology subjects

Croasdale-Wood, R.

Publications and source records attributed to Croasdale-Wood, R..

4 recordsLinked to original sources

Generative Language Modeling for Antibody CDR Grafting and Alignment-driven De Novo Design

Antibodies recognise their targets through hypervariable complementarity-determining regions (CDRs), which are interleaved with conserved frameworks in sequence space, making de novo CDR design an infilling problem. Autoregressive models generate residues left-to-right, which precludes full framework context during CDR generation and conflates framework and CDR likelihoods, leaving no natural prompt-response interface for feedback to steer generation. We present GenCDR, a family of LLaMa-based autoregressive language models that read all frameworks as a conditioning prompt and generate all CDRs jointly as a variable-length response, making CDR likelihoods a clean, separable target for reward attribution. The family comprises IgGenCDR, p-IgGenCDR, and NanoGenCDR, trained on unpaired, paired, and nanobody chains, respectively. GenCDR achieves the highest CDR recovery among autoregressive models and produces natural, diverse, human-like CDRs whose likelihoods correlate with fitness and developability assays. The prompt-response boundary also enables principled alignment: reward signals for binding affinity, expression, or developability can be composed to steer CDR generation. Over four rounds of alignment against antibody-antigen co-folding and developability objectives, we find that NanoGenCDR, which uses no explicit antigen encoding, can reach in silico structural interface metrics competitive with those of a structure-conditioned diffusion pipeline at roughly half the sampling budget, with more natural, developable designs. The same interface can be extended to integrate experimental feedback, opening a path to closed-loop antibody de novo design.

synthetic biology↗

Accelerating Antibody Development: Sequence and Structure-Based Models for Predicting Developability Properties through Size Exclusion Chromatography

Experimental screening for biopharmaceutical developability properties typically relies on resource-intensive, and time-consuming assays such as size exclusion chromatography (SEC). This study highlights the potential of in silico models to accelerate the screening process by exploring sequence and structure-based machine learning techniques. Specifically, we compared surrogate models based on pre-computed features extracted from sequence and predicted structure with sequence-based approaches using protein language models (PLMs) like ESM-2. In addition to different end-to-end fine-tuning strategies for PLM, we have also investigated the integration of the structural information of the antibodies into the prediction pipeline through graph neural networks (GNN). We applied these different methods for predicting protein aggregation propensity using a dataset of approximately 1200 Immunoglobulin G (IgG1) molecules. Through this empirical evaluation, our study identifies the most effective in silico approach for predicting developability properties for SEC assays, thereby adding insights to existing screening efforts for accelerating the antibody development process.

bioinformatics↗

p-IgGen: A Paired Antibody Generative Language Model

A key challenge in antibody drug discovery is designing novel sequences that are free from developability issues - such as aggregation, polyspecificity, poor expression, or low solubility. Here, we present p-IgGen, a protein language model for paired heavylight chain antibody generation. The model generates diverse, antibody-like sequences with pairing properties found in natural antibodies. We also create a finetuned version of p-IgGen that biases the model to generate antibodies with 3D biophysical properties that fall within distributions seen in clinical-stage therapeutic antibodies.

immunology↗

Enhancement of antibody thermostability and affinity by computational design in the absence of antigen

Over the last two decades, therapeutic antibodies have emerged as a rapidly expanding domain within the field biologics. In silico tools that can streamline the process of antibody discovery and optimization are critical to support a pipeline that is growing more numerous and complex every year. In this study, DeepAb, a deep learning model for predicting antibody Fv structure directly from sequence, was used to design 200 potentially stabilized variants of an anti-hen egg lysozyme (HEL) antibody. We sought to determine whether DeepAb can enhance the stability of these antibody variants without relying on or predicting the antibody-antigen interface, and whether this stabilization could increase antibody affinity without impacting their developability profile. The 200 variants were produced through a robust highthroughput method and tested for thermal and colloidal stability (Tonset, Tm, Tagg), affinity (KD) relative to the parental antibody, and for developability parameters (non-specific binding, aggregation propensity, self-association). In the designed clones, 91% and 94% exhibited increased thermal and colloidal stability and affinity, respectively. Of these, 10% showed a significantly increased affinity for HEL (5-to 21-fold increase), with most clones retaining the favorable developability profile of the parental antibody. These data open the possibility of in silico antibody stabilization and affinity maturation without the need to predict the antibody-antigen interface, which is notoriously difficult in the absence of crystal structures.

molecular biology↗