bioRxiv Science⌕ Search

Biology subjects

Santolla, N.

Publications and source records attributed to Santolla, N..

2 recordsLinked to original sources

Kukulu: Diffusion-Based Reconstruction of Antibody CDR Loops using a Structure-Aware Joint Embedding Predictive Architecture

Antibody complementarity-determining regions (CDRs), especially CDR-H3, are a dominant source of binding specificity but remain difficult to design due to coupled sequence-structure constraints and local geometric variability. Here we present Kukulu, a structure-aware Joint Embedding Predictive Architecture (JEPA) combined with conditional diffusion for CDR loop reconstruction in antibody-antigen complexes. Our pipeline prepares structures by chain-aware cleanup, Fv trimming, Chothia-indexed CDR identification, and in silico CDR masking, then trains on paired prepared/masked structures represented in an atom37 format. The model uses a context encoder over masked structures, a transformer predictor for latent CDR representations, and a diffusion head that reconstructs loop coordinates, atom presence, and residue identities under geometry-aware losses. During generation, Kukulu denoises only masked CDR residues while preserving frame-work context, then optionally rebuilds sidechains with local frame templates and performs post-generation structural relaxation. This manuscript provides a methods-focused overview of the models implementation details and an evaluation protocol based on structure quality and docking-oriented scoring for integration into existing antibody design workflows.

synthetic biology↗

peleke-1: A Suite of Protein Language Models Fine-Tuned for Targeted Antibody Sequence Generation

The discovery of therapeutic antibodies is a traditionally arduous process. Today, the lab-based process of antibody discovery consists of several time-consuming steps that involve live animal immunization, B-cell harvesting, hybridoma creation, and then downstream engineering and evaluation. However, the use of artificial intelligence in drug design has previously been shown effective in the rapid generation of proteinspecific binders, small molecules, and even antibody therapeutics, thereby replacing some of the primary steps of the drug discovery process. Here we present peleke-1, a suite of protein language models fine-tuned from state-of-the-art large language models using curated antibody-antigen complex data. These models generate targeted antibody Fv sequences for a given antigen sequence input at-scale. This suite of models provides a reliable, artificial intelligence-driven approach for in silico therapeutic antibody discovery along with an open-source framework for future antibody language model tuning.

immunology↗