bioRxiv Science⌕ Search

Biology subjects

Teixeira, A. A. R.

Publications and source records attributed to Teixeira, A. A. R..

3 recordsLinked to original sources

An interpretable open platform for sequence-based antibody developability prediction

Antibody developability is increasingly predictable from sequence, yet software and trained models are rarely made available. We present DELPHI, open software for training developability predictors from labelled antibody assay data, together with ready-to-run, retrainable models. DELPHI compares 25 language-model and classifier combinations under CDR H3-cluster cross-validation that reduces sequence-similarity leakage, measures how performance changes with labelled training-set size, and reports residue-level model attributions. Applied to in-house polyreactivity and size-exclusion (SEC) data, it reaches mean AUC 0.959 and 0.933. Trained on those data alone, it transfers to a 246,293-antibody public library (AUC 0.950 with our deployed model) and ranks polyreactivity at a level similar to the best reported Ginkgo competition point estimate, without training on its data. Any laboratory can screen candidates before running assays, generate residue-level engineering hypotheses, and retrain DELPHI for a new assay.

bioinformatics↗

OpenAntigens: a structure-aware database for antigen construct design across the human cell-surface and secreted proteome

Choosing which region of a protein to express remains poorly standardized in antibody discovery, recombinant reagent generation, structural biology and computational binder design. For human cell-surface and secreted proteins, this requires reconciling topology, processing, predicted and experimental structure, modifications, interaction partners, orthologs, paralogs and cross-reactivity risk before ordering DNA. OpenAntigens is a free, no-login database of construct-design reports for 5328 human secreted, GPI-anchored, single-pass and multipass proteins. It integrates UniProt topology, AlphaFold pLDDT and PAE, PDB precedent, InterPro and Pfam domains, mouse and cynomolgus orthologs, paralog and family context, Open Targets disease associations, partner and assembly context, and BLAST searches. It provides 55 305 construct suggestions spanning full design regions, PDB-backed boundaries, annotated domains, pLDDT/PAE-derived regions and membrane-expression options, plus 148 722 sequence-similarity hits to help choose constructs and assess cross-reactivity. For targets with compatible AlphaFold models, the interactive designer links sequence, structure, pLDDT and PAE, allowing users to revise boundaries and export species-equivalent sequences with real-time cysteine and modification warnings. OpenAntigens places reproducible construct suggestions, comparative context and browser editing in one workflow, reducing manual reconciliation across resources. OpenAntigens is available at openantigens.org.

bioinformatics↗

High-Throughput Machine Learning-Aided Antibody Discovery for Cell Surface Antigens

Machine learning (ML) has the potential to revolutionize antibody design and selection, but its success depends on access to extensive, well-curated datasets of antibody-antigen interactions. To address this need, we developed a synthetic Fab yeast display library optimized for seamless ML integration, focusing on sequence diversity within the CDRH3 loop. The library incorporates key sequence features derived from human B cell repertoires essential for efficient antibody generation captured in a compact antigen recognition module (ARM) format. Built using the VH1-69 heavy chain and four light chains, the library was evaluated against ten human and murine cell surface antigens, including PD-L1, TIGIT, and ROBO1. This approach yielded hundreds of antibodies with robust biophysical properties, validated for functional performance in flow cytometry and immunohistochemistry. Furthermore, ML analysis identified additional antibodies for ROBO2 and PD-L2 from the aggregate sequencing data, demonstrating utility for hybrid in silico and experimental workflows. We provide a publicly accessible dataset comprising more than 68,000 Fab sequences and 486 characterized antibodies. This study establishes an ML-compatible framework designed to accelerate and streamline antibody discovery and development.

biophysics↗