bioRxiv Science⌕ Search

Biology subjects

Grandguillaume, I.

Publications and source records attributed to Grandguillaume, I..

2 recordsLinked to original sources

ARID-sf: A physics-informed Deep Learning scoring function to improve Antibody-Antigen docking model ranking

Accurate prediction of antibody-antigen (Ab-Ag) complexation is crucial for understanding immune responses, diagnostics, and the development of therapeutic antibodies. While molecular docking generates conformations, current scoring functions struggle to identify nearnative poses, particularly for Ab-Ag interactions. We present ARID-sf (Antibody-antigen Residue Interface Docking scoring function), which combines classical force field potentials with structural features and protein language model embeddings through a self-attention-based neural network architecture. ARID-sf was trained on >1.5 million docking models and evaluated across four independent test sets comprising 806 cases and 700,000+ docking models. ARID-sf consistently outperforms other functions on increasingly challenging docking scenarios. Critically, ARID-sf maintains performance across diverse sequence identities and increasing conformational change required to reach the bound state, starting with unbound components, demonstrating robust generalization. ARID-sf is parallelizable and can process thousands of docking models per minute, enabling practical application in computational Ab engineering pipelines. The code, trained network, and complete pipeline are freely available at https://github.com/DSIMB/ARID-sf.git.

bioinformatics↗

ANABAG: Annotated Antibody Antigen dataset with unique features for Antibody Engineering Applications

The analysis and prediction of antibody-antigen (Ab-Ag) interactions often overlook critical structural features such as glycosylation, physical chemical conditions like pH and salt concentration, as well as the lack of standardized criteria for selecting complexes based on structural properties and sequence identity. Common practices in dataset construction rely on removing redundancy using sequence identity thresholds, which can inadvertently exclude complexes with alternative binding modes that share identical sequences. To enable more precise Ab-Ag modeling and antibody engineering, it is essential to incorporate richer structural and physical information into both physics-based and machine learning models. To address these limitations, we present ANABAG, a new curated dataset of Ab-Ag complexes annotated at the residue level with UniProt sequence information and enriched with a wide range of structural and physicochemical features. The dataset allows flexible filtering of complexes using a variety of descriptors available at both the complex and residue levels. Selected features are ready to use in machine learning workflows, while the structural files are compatible with antibody design and docking pipelines like Rosetta or Haddock. The complete dataset is available on Zenodo, and all accompanying scripts and usage documentation can be accessed via GitHub at https://github.com/DSIMB/anabag-handler.git.

bioinformatics↗