bioRxiv Science⌕ Search

Biology subjects

Heyer, K.

Publications and source records attributed to Heyer, K..

2 recordsLinked to original sources

Germline-encoded V(D)J gene usage does not impose strict constraints on the epitope-specificity of T cell receptors

The theoretical diversity of T cell receptors (TCRs), generated through V(D)J recombination, is enormous, yet the diversity of TCRs capable of recognizing the same epitope remains unknown. Defining this TCR solution space is essential for uncovering basic principles that govern TCR specificity. Using single-cell RNA and TCR sequencing, we generated ultra-deep (more than 4000 unique TCRs per epitope) epitope-specific TCR libraries derived from 560 immunized C57BL/6 mice, identifying over 27,000 unique epitope-reactive TCRs across three distinct CD8+ T cell epitopes presented by two major histocompatibility complex (MHC) class I alleles. Saturation analyses indicated that the solution space for all studied epitopes comprises many tens of thousands of unique TCRs. Despite highly skewed and peptide-dependent VJ-usage patterns, nearly the entire set of functional germline V/ and J/ segments was detected at least once within each epitope-specific repertoire. Therefore, diversity of epitope-specific TCRs is not limited by distinct germline combinations but rather can emerge from a near-to-complete combinatorial space of - and -chain, V and J segments paired with compatible CDR3 sequences.

immunology↗

Efficient Search of Ultra-Large Synthesis On-Demand Libraries with Chemical Language Models

Ultra-large building block catalogs provide inexpensive access to billions of synthesis-on-demand molecules, but the combinatorial scale renders conventional virtual screening impractical. We present Vector Virtual Screen (VVS), a score-function-agnostic machine learning framework for efficient navigation of combinatorial libraries and rapid identification of promising molecules for experimental validation. VVS comprises four key innovations: (i) the Embedding Decomposer, which factors molecules into building blocks in latent space; (ii) ChemRank, a correlation-based loss that improves retrieval precision; (iii) BBKNN, an algorithm for nearest-neighbor search directly in building block space; and (iv) a multi-scale hill-climbing algorithm for gradient-based navigation of molecular embedding vector databases. Across diverse scoring functions, VVS consistently outperforms existing methods in retrieving high-scoring molecules while evaluating only a fraction of the library, achieving orders-of-magnitude runtime improvements. By turning ultra-large libraries into tractable search spaces, VVS enables virtual screening to keep pace with the rapid expansion of chemical space and adapt seamlessly to future advances in scoring functions.

bioinformatics↗