bioRxiv Science⌕ Search

Biology subjects

Malherbe, C.

Publications and source records attributed to Malherbe, C..

4 recordsLinked to original sources

Benchmarking Generative Models for Antibody Design

Generative models trained on antibody sequences and structures have shown great potential in advancing machine learning-assisted antibody engineering and drug discovery. Current state-of-the-art models are primarily evaluated using two categories of in silico metrics: sequence-based metrics, such as amino acid recovery (AAR), and structure-based metrics, including root-mean-square deviation (RMSD), predicted alignment error (pAE), and interface predicted template modeling (ipTM). While metrics such as pAE and ipTM have been shown to be useful filters for experimental success, there is no evidence that they are suitable for ranking, particularly for antibody sequence designs. Furthermore, no reliable sequence-based metric for ranking has been established. In this work, using real-world experimental data from fourteen diverse datasets, we extensively benchmark a range of generative models, including LLM-style, diffusion-based, and graph-based models. We show that log-likelihood scores from these generative models have promising correlation with experimentally measured binding affinities, suggesting that log-likelihood can potentially serve as a reliable metric for ranking antibody sequence designs. Additionally, we scale up one of the diffusion-based models by training it on a large and diverse synthetic dataset, significantly enhancing its ability to rank antibodies based on their binding affinities. We also evaluate non-log-likelihood-based metrics on ten datasets and find that, while they are less consistent for ranking, they provide complementary information. Structure-, energy-, and sequence-based scores appear to be orthogonal and may be used together to increase the likelihood of experimental success. Our implementation is available at: https://github.com/AstraZeneca/DiffAbXL

bioinformatics↗

IgBlend: Unifying 3D Structures and Sequences in Antibody Language Models

Large language models (LLMs) trained on antibody sequences have shown significant potential in the rapidly advancing field of machine learning-assisted antibody engineering and drug discovery. However, current state-of-the-art antibody LLMs often overlook structural information, which could enable the model to more effectively learn the functional properties of antibodies by providing richer, more informative data. In response to this limitation, we introduce IgBlend, which integrates both the 3D coordinates of backbone atoms (C-alpha, N, and C) and antibody sequences. Our model is trained on a diverse dataset containing over 4 million unique structures and more than 200 million unique sequences, including heavy and light chains as well as nanobodies. We rigorously evaluate IgBlend using established benchmarks such as sequence recovery, complementarity-determining region (CDR) editing and inverse folding and demonstrate that IgBlend consistently outperforms current state-of-the-art models across all benchmarks. Furthermore, experimental validation shows that the models log probabilities correlate well with measured binding affinities.

bioinformatics↗

Ground-truth validation of uni- and multivariate lesion inference approaches

Lesion analysis aims to reveal causal contributions of brain regions to brain functions. Various strategies have been used for such lesion inferences. These approaches can be broadly categorized as univariate or multivariate methods. Here we analysed data from 581 patients with acute ischemic injury, parcellated into 41 Brodmann areas, and systematically investigated the inferences made by two univariate and two multivariate lesion analysis methods via ground-truth simulations, in which we defined a priori contributions of brain areas to assumed brain function. Particularly, we analysed single-region models, with only single areas presumed to contribute functionally, and multiple-region models, with two contributing regions that interacted in a synergistic, redundant or mutually inhibitory mode. The functional contributions could vary in proportion to the lesion damage or in a binary way. The analyses showed a considerably better performance of the tested multivariate than univariate methods in terms of accuracy and mis-inference error. Specifically, the univariate approaches of Lesion Symptom Mapping (LSM) as well as Lesion Symptom Correlation (LSC) mis-inferred substantial contributions from several areas even in the single-region models, and also after accounting for lesion size. By contrast, the multivariate approaches of Multi- Area Pattern Prediction (MAPP), which is based on machine learning, and Multi-perturbation Shapley value Analysis (MSA), based on coalitional game theory, delivered consistently higher accuracy and specificity. Our findings suggest that the tested multivariate approaches produce largely reliable lesion inferences, without requiring lesion size consideration, while the application of the univariate methods may yield substantial mis-localizations that limit the reliability of functional attributions.

neuroscience↗

Grapevine leaf MALDI-MS imaging reveals the localisation of a putatively identified sucrose metabolite associated to Plasmopara viticola development

Despite well-established pathways and metabolites involved in grapevine-Plasmopara viticola interaction, information on the molecules involved in the first moments of pathogen contact with the leaf surface and their specific location is still missing. To understand and localise these molecules, we analysed grapevine leaf discs infected with P. viticola with MSI. Plant material preparation was optimised, and different matrices and solvents were tested. Our data shows that trichomes hamper matrix deposition and the ion signal. Results show that putatively identified sucrose presents a higher accumulation and a non-homogeneous distribution in the infected leaf discs in comparison with the controls. This accumulation was mainly on the veins, leading to the hypothesis that sucrose metabolism is being manipulated by the development structures of P. viticola. Up to our knowledge this is the first time that the localisation of a putatively identified sucrose metabolite was shown to be associated to P. viticola infection sites.

plant biology↗