bioRxiv Science⌕ Search

Biology subjects

Comeau, S. R.

Publications and source records attributed to Comeau, S. R..

3 recordsLinked to original sources

Benchmarking antigen-aware inverse folding methods for antibody design.

Computational antibody design has seen many recent advances pioneered via the use of language models and advanced structure prediction tools. Developing a de novo antibody against a specific antigen requires structural awareness that most language models lack. A prominent class of machine learning methods combining the best of language model and structural worlds is inverse folding. This approach aims to predict a sequence that would fit a given structure. Such methods are now increasingly used to predict alternate sequences given a structure of a binder. It is known that, just like language models, such methods have certain predictive power in identifying binders. Here we performed a set of tests to reveal where, if at all, such methods provide value in the realistic setting of antibody discovery.

bioinformatics↗

nanoFOLD : sequence design of nanobodies via inverse folding

Antibodies devoid of light chains are a promising class of biotherapeutics. Computational methods that address these molecules are crucially needed to accelerate the traditional, long and expensive experimental process of their discovery. Inverse folding, wherein one is tasked to predict a sequence given molecular coordinates, is an established method in scaffold-based protein design. Here we develop an inverse folding method speci[fi]c to nanobodies. We demonstrate its application in nanobody-engineering scenarios of enriching binders from next-generation sequencing experiments and novel binder design.

bioinformatics↗

Predicting the Purity of Multispecific Antibodies From Sequence Using Machine Learning: Methods and Applications

Multispecific antibodies are prominent therapeutic agents, but many molecular formats and drug candidates that show promise during molecular discovery stages cannot be scaled up and developed into drugs due to inadequate developability. During the discovery stages, the selection of molecule format(s), molecule design, purity, and initial physiochemical stability testing criteria largely rely on scientists experience. Machine learning, however, can identify hidden trends in large datasets, aiding in the selection of drug candidates with improved developability. In this study, we present a machine learning approach to predict antibody purity, measured by the percentage of monomer after protein A purification. Using the amino acid sequences of variable regions, molecular formats, germlines and germline pairings, and calculated physiochemical properties as inputs, machine learning models were trained to predict the percentage of monomer for a given multispecific antibody (Figure 1). The dataset employed in this study consists of [~]500 multi-specific antibodies generated during BIs internal drug discovery programs. Our results indicate that machine learning, when applied to sequence, germline, and format data, can effectively predict antibody percentage of monomer. Incorporating this approach into high-throughput multispecific antibody screening processes can save time and resources by reducing the need to test a large subset of potentially unstable antibodies. While this study focused on percentage of monomer as a test case, similar approaches can be employed to predict other antibody properties, such as melting temperature (Tm), hydrophobicity (aHIC), and solution stability properties (AC-SINS). O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=47 SRC="FIGDIR/small/570217v1_fig1.gif" ALT="Figure 1"> View larger version (11K): org.highwire.dtl.DTLVardef@f23246org.highwire.dtl.DTLVardef@c2b81aorg.highwire.dtl.DTLVardef@1c4e59dorg.highwire.dtl.DTLVardef@1bed683_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1C_FLOATNO Overview of ML model for predicting multispecific antibody purity from sequence, germline and format information. C_FIG

molecular biology↗