bioRxiv Science⌕ Search

Biology subjects

Wenckstern, J.

Publications and source records attributed to Wenckstern, J..

2 recordsLinked to original sources

The Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates

Affinity reagents such as antibodies are indispensable for interrogating proteins biological function. Yet they are costly and frequently unreliable, with unknown sequences, posing challenges to reproducible experimental research. Deep learning-based protein design can now in silico generate affinity reagents achieving reliable experimental success rates, but has remained largely confined to specialist laboratories. Here we present the Human Bindome, a proteome-scale atlas of high-confidence in silico protein binder candidates. By embedding the experimentally benchmarked BindCraft method in an accelerated, parallelized framework with automated domain-level target selection, we generated 306,146 binder candidates covering 8,296 human proteins (40.9% of the full proteome). Every candidate carries a defined sequence, a predicted binder-target structure model, and in silico confidence metrics. We characterize proteome-wide coverage and show that binder epitopes frequently overlap functional sites. This positions the Bindome as a resource of genetically encodable perturbagens for site-specific, modular control of protein function. The Bindome is freely available through a web interface (https://bindome.epfl.ch), with agentic, natural-language querying and as data splits for machine-learning model development. We anticipate that the Bindome will be valuable for the scientific community by providing affinity and perturbation reagents with broad applications in dissecting biological mechanisms as well as in drug and target discovery.

synthetic biology↗

A Diffusion-Based Autoencoder for Learning Patient-Level Representations from Single-Cell Data

Single-cell RNA sequencing (scRNA-seq) offers insights into cellular heterogeneity and tissue composition, yet leveraging this data for patient-level clinical predictions remains challenging due to the set-structured nature of single-cell data, as well as the scarcity of labeled samples. To address these challenges, we introduce scSet, a diffusion-based autoencoder that learns patient-level representations from sets of single-cell transcriptomes. Our method uses a transformer-based encoder to process variably sized and unordered cell inputs, coupled with a conditional diffusion decoder for self-supervised learning on unlabeled data. By pre-training on large-scale unlabeled datasets, scSet generates robust patient representations that can be fine-tuned for downstream clinical prediction tasks. We demonstrate the effectiveness of scSet patient embeddings for clinical prediction across multiple real-world datasets, where they outperform existing patient representations, even with limited labeled data. This work represents an important step toward bridging the gap between single-cell resolution and patient-level insights. Code is available at https://github.com/clinicalml/scset.

bioinformatics↗