bioRxiv Science⌕ Search

Biology subjects

Deibler, K.

Publications and source records attributed to Deibler, K..

5 recordsLinked to original sources

Latent generative search unlocks de novo design of untapped biomolecular interactions at scale

De novo protein design has advanced rapidly, yet designing binders to polar, solvent-exposed epitopes and small, flexible ligands remains challenging. Such hydrated surfaces and flexible molecules, including carbohydrates, provide few of the hydrophobic contacts favoured by current methods and have largely resisted de novo binders. To address this challenge, here we introduce latent generative search for binder design, a novel framework that uses reward-guided search at inference time to steer the Proteina-Complexa generative model. The model codesigns sequence and structure - generating them together in a continuous latent space - and thereby removes the inverse-folding step on which current methods rely. In a screen of more than one million designs by multiplexed phage display, latent generative search produced more validated binders than every other method tested, its codesigned sequences surpassing post hoc redesign. It delivered high-affinity binders across therapeutic receptors, a viral attachment protein and intracellular signalling targets. Our approach also accessed previously untapped biology, generating the first de novo proteins that bind a free carbohydrate, including one that discriminates between blood-group antigens - a polar, flexible target class beyond the reach of current design methods.

bioengineering↗

Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery

Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model's performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation), showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.

bioinformatics↗

PeptideMTR: Scaling SMILES-Based Language Models for Therapeutic Peptide Engineering

Therapeutic peptides occupy a unique middle ground in drug discovery, offering the high specificity of protein interactions with the chemical diversity of small molecules, yet they currently fall in a computational blind spot. Existing foundation models cannot handle them effectively: protein models are restricted to natural amino acids, while chemical models struggle to process large, polymer-like sequences. This disconnect has forced the field to rely on static chemical descriptors that fail to capture subtle chemical details or on complex multi-embedding pipelines that are custom tailored to specific datasets. To bridge this gap, we present PeptideCLM-2, a suite of chemical language models trained on over 100 million molecules to natively represent complex peptide chemistry. This modeling approach expands the available toolkit of machine learning models for therapeutic peptides. Benchmarking results show strong performance versus prior methods for predicting development endpoints including membrane diffusion, biological function, and half life.

bioinformatics↗

Predicting peptide aggregation with protein language model embeddings

Amyloid fibrils, a form of peptide aggregate, are associated with multiple diseases and hinder the development of therapeutics. The experimental characterization of aggregating peptides is resource-intensive and data are scarce, limiting the development of accurate models. We present a deep-learning model, PALM (Predicting Aggregation with Language Model embeddings), which uses transfer learning to predict aggregation from embeddings extracted from a pretrained protein language model (pLM). PALM is trained on the WaltzDB-2.0 dataset to classify peptides and identify aggregation-prone regions within a sequence at single-residue resolution. Compared to existing models, it exhibits competitive performance on diverse held-out experimental datasets. We find that PALM fails to identify single mutations that increase the rate of aggregation of amyloid beta peptide; however, training the PALM architecture on a larger dataset, CANYA NNK1-3, substantially improves performance in this task. These results show that transfer learning with pLM embeddings improves performance when training on small datasets, but highlight that challenging tasks, such as predicting the effect of single mutations, require more experimental data.

bioinformatics↗

De novo design of miniprotein agonists and antagonists targeting G protein-coupled receptors

G protein-coupled receptors (GPCRs) play key roles in physiology and are central targets for drug discovery and development, but the design of protein agonists and antagonists has been challenging as GPCRs are integral membrane proteins and conformationally dynamic. Here we describe computational de novo design methods and a high throughput "receptor diversion" microscopy-based screen for generating GPCR binding miniproteins with high affinity, potency and selectivity, and the use of these methods to generate agonists for MRGPRX1, NK1R and CCR5, as well as antagonists for CXCR4, CCR5, OXTR, GLP1R, GIPR, GCGR, PTH1R and CGRPR.. Cryo-electron microscopy data reveals atomic-level agreement between designed and experimentally determined structures for CGRPR- and CXCR4-bound antagonists and MRGPRX1-bound agonists. Our de novo design and screening approach opens new frontiers in GPCR drug discovery and development.

bioengineering↗