bioRxiv Science⌕ Search

Biology subjects

Gordon, G. L.

Publications and source records attributed to Gordon, G. L..

3 recordsLinked to original sources

ANARCII: A Generalised Language Model for Antigen Receptor Numbering

Antigen receptor numbering allows the rapid delineation of the antigen-binding regions of antibody and T cell receptor (TCR) sequences, from sequence alone. It also allows the comparison of the vast diversity of antigen receptors in a consistent frame of reference. Numbering of antigen receptors is currently achieved by aligning sequences to a reference set. This approach may result in different numbering, depending on the reference set used or may fail to number query sequences derived from new species or rare sequence types. To address this problem, we have built a new numbering method (ANARCII) which requires no alignment step and is based on a Seq2Seq language model. Our results show that ANARCII can deal with the complexity that arises in experimentally collected sequencing data and generalise to sequences which are highly dissimilar to those in training. In test sets designed to contain challenging and ambiguous sequence patterns ANARCII numbering was identical to existing methods for over 99.99% of conserved residues and over 99.94% for complete CDR regions. The lightweight architecture allows numbering of over 90,000 sequences per minute on a single A100 GPU. Furthermore, the ANARCII package can be conditioned to fit rare sequence types and provide new training data for fine-tuning. We demonstrate that fine-tuned versions of ANARCII can correctly number other immunoglobulin domains such as TCRs and VNARs. Our model is freely available as a web tool (https://github.com/oxpig/ANARCII), as well as a package for high throughput numbering of next generation sequencing data (https://opig.stats.ox.ac.uk/webapps/sabdab-sabpred/sabpred/anarcii/).

bioinformatics↗

PLAbDab-nano: a database of camelid and shark nanobodies from patents and literature

Nanobodies are essential proteins of the adaptive immune systems of camelid and shark species, complementing conventional antibodies. Properties such as their relatively small size, solubility and high thermostability make VHH and VNAR modalities a promising therapeutic format and a valuable resource for a wide range of biological applications. The volume of academic literature and patents related to nanobodies has risen significantly over the past decade. Here, we present PLAbDab-nano, a nanobody complement to the Patent and Literature Antibody Database (PLAbDab). PLAbDab-nano is a selfupdating, searchable repository containing approximately 5000 annotated VHH and VNAR sequences. We describe the methods used to curate the entries in PLAbDab-nano, and highlight how PLAbDab-nano could be used to design diverse libraries, as well as find sequences similar to known patented or therapeutic entries. PLAbDab-nano is freely available as a searchable web server (opig.stats.ox.ac.uk/webapps/plabdab-nano/). Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=193 SRC="FIGDIR/small/604232v1_ufig1.gif" ALT="Figure 1"> View larger version (40K): org.highwire.dtl.DTLVardef@5d81d4org.highwire.dtl.DTLVardef@f69f21org.highwire.dtl.DTLVardef@1493c54org.highwire.dtl.DTLVardef@117e3c5_HPS_FORMAT_FIGEXP M_FIG C_FIG

immunology↗

A comparison of the binding sites of antibodies and single-domain antibodies

Antibodies are the largest class of biotherapeutics. However, in recent years, single-domain antibodies have gained traction due to their smaller size and comparable binding affinity. Antibodies (Abs) and single-domain antibodies (sdAbs) differ in the structures of their binding sites: most significantly, single-domain antibodies lack a light chain and so have just three CDR loops. Given this inherent structural difference, it is important to understand whether Abs and sdAbs are distinguishable in how they engage a binding partner and thus, whether they are suited to different types of epitopes. In this study, we use non-redundant sequence and structural datasets to compare the paratopes, epitopes and antigen interactions of Abs and sdAbs. We demonstrate that even though sd-Abs have smaller paratopes, they target epitopes of equal size to those targeted by Abs. To achieve this, the paratopes of sdAbs contribute more interactions per residue than the paratopes of Abs. Additionally, we find that conserved framework residues are of increased importance in the paratopes of sd-Abs, suggesting that they include non-specific interactions to achieve comparable affinity. Further-more, the epitopes of sdAbs and Abs cannot be distinguished by their shape. For our datasets, sd-Abs do not target more concave epitopes than Abs: we posit that this may be explained by differences in the orientation and compaction of sdAb and Ab CDR-H3 loops. Overall, our results have important implications for the engineering and humanization of sdAbs, as well as the selection of the best modality for targeting a particular epitope.

immunology↗