bioRxiv Science⌕ Search

Biology subjects

Diamant, N.

Publications and source records attributed to Diamant, N..

8 recordsLinked to original sources

TxConformal: Controlling False Discoveries in AI-Driven Therapeutic Discovery

Artificial Intelligence (AI) is transforming therapeutic discovery by scoring a large set of promising candidates and prioritizing a shortlist for further investigation. Quantifying the reliability of AI scores and preventing false positives among selected candidates is key to the efficiency of the discovery process. Conformal prediction (CP) has emerged as a popular tool for guiding such prioritization, especially via the conformal selection framework to control false discovery rates (FDR) in selecting top-ranked candidates under distributional shift1, 2. However, deploying these advances in real-world therapeutic discovery remains challenging: distribution shifts are difficult to quantify and correct in high-dimensional biomedical data, and practical workflows often require flexible error metrics. Here, we present TO_SCPLOWXC_SCPLOWCO_SCPLOWONFORMALC_SCPLOW, a general framework for trustworthy decision making when building shortlists using AI scores. TO_SCPLOWXC_SCPLOWCO_SCPLOWONFORMALC_SCPLOW adjusts for distribution shift by balancing the hidden representations in AI models and then provides confidence measures for true discoveries of target biological properties. These confidence measures, interpretable as p-values, can be used in conjunction with statistical multiple testing procedures to derive selection decisions with limited false positives or to estimate the errors in given selection decisions. TO_SCPLOWXC_SCPLOWCO_SCPLOWONFORMALC_SCPLOW controls the false positive rate in six real-world tasks spanning various therapeutic discovery stages, modalities, and AI models with realistic data splits. When selecting promising combinatorial genetic perturbations, TO_SCPLOWXC_SCPLOWCO_SCPLOWONFORMALC_SCPLOW nearly halves false-positive selections compared to baseline methods, substantially reducing unnecessary experimental costs by tens of thousands of dollars. When selecting stable protein structures under mutant shifts, TO_SCPLOWXC_SCPLOWCO_SCPLOWONFORMALC_SCPLOW identifies about 10 times more proteins than baseline methods at stringent thresholds when running at a target FDR level of 10%, recovering over 90% of valuable candidates that baseline methods miss due to unaccounted distribution shifts. Furthermore, we demonstrate that TO_SCPLOWXC_SCPLOWCO_SCPLOWONFORMALC_SCPLOW robustly supports various alternative error metrics suitable for resource-constrained settings. Finally, in a prospective fixed-budget virtual screening campaign for novel antibiotic discovery, TO_SCPLOWXC_SCPLOWCO_SCPLOWONFORMALC_SCPLOW predicted false positives in close agreement with experimental outcomes, with substantial improvements over simple baselines.

bioinformatics↗

Nona: A unifying multimodal masking framework for functional genomics

The non-coding genome encodes complex regulatory logic that orchestrates gene expression and cell identity. While machine learning models for functional genomics have advanced our understanding of the cis-regulatory code, sequence-to-function models, DNA language models, and generative models have evolved as separate paradigms despite probing the same underlying regulatory biology. We introduce Nona, a multimodal masked modeling framework that unifies these paradigms by learning jointly from DNA sequence and base-resolution functional genomics data. Beyond unifying existing modeling paradigms, Nona enables entirely new modeling objectives. We demonstrate its versatility through three applications: (1) a context-aware sequence-to-function model that improves local predictions by up to 13% by correcting systematic errors in sequence-to-function predictions; (2) a functional language model that integrates functional data into language modeling, learns relevant regulatory sequence motifs, and enables regulatory element design through masked discrete diffusion; (3) functional genotyping, which reveals an unrecognized privacy vulnerability in processed ATAC-seq data and re-identifies individuals from genetic databases with perfect accuracy. Together, these results establish masking as a universal interface for integrated modeling of functional genomics data, unifying disparate approaches while opening new directions for understanding and engineering the regulatory genome.

genomics↗

Deep Evolutionary Fitness Inference for Variant Nomination from Directed Evolution

Iterative screening techniques, such as directed evolution, enable high-throughput affinity maturation to optimize binders to molecular interfaces. However, the decision problem of selecting variants from rich, evolved populations to enter low-throughput follow-up methods remains a significant bottleneck. Here, we present evolutionary fitness inference (EVFI) and DeepEVFI, two machine learning methods that model directed evolution from time-series sequencing data, and infer fitness, a variants ability to enrich under selection pressure. Our methods flexibly handle mutation mechanisms and starting populations that may be partially unknown - settings relevant to drug discovery - and achieve strong performance on a diverse set of experimental data. We conducted two experimental directed evolution campaigns, using antibodies and macrocyclic peptides libraries to identify and optimize binders to therapeutically relevant targets. EVFI and DeepEVFI identified tighter binders that were missed by human experts using conventional frequency-based approaches, including "rising stars" with low frequency. Beyond initial hit discovery, EVFI and Deep-EVFI enables labeling large-scale sequence-fitness datasets and identifying variants of initial binders with diverse properties.

bioinformatics↗

Foundation Model Attributions Reveal Shared Inflammatory Program Across Diseases

Determining a genes functional significance within a cellular context has long been a challenge, as absolute expression level is an unreliable indicator. We introduce SIGnature, a framework for scoring gene importance by leveraging attributions derived from single-cell RNA-sequencing (scRNA-seq) foundation models. Attribution scores reduce technical noise, emphasize regulatory genes, and facilitate cross-dataset comparison - a core challenge for scRNA-seq analyses. We developed the SIGnature package as a tool for generating and querying attributions, enabling rapid gene set searches across massive scRNA-seq atlases. We demonstrated its utility using the MS1 monocyte signature, a poorly understood gene program activated in severe COVID-19 and sepsis. Searching 400 studies revealed novel associations between the MS1 signature and multiple hyperinflammatory conditions, including Kawasaki disease. Experimental validation confirmed Kawasaki disease patient serum induces the MS1 phenotype. These findings highlight that SIGnature can uncover shared mechanisms across conditions, demonstrating its power for large-scale signature scoring and cross-disease analysis.

bioinformatics↗

Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states

Single-cell RNA-seq (scRNA-seq) has become a prominent tool for studying human biology and disease. The availability of massive scRNA-seq datasets and advanced machine learning techniques has recently driven the development of single-cell foundation models that provide informative and versatile cell representations based on expression profiles. However, to understand disease states, we need to consider entire tissue ecosystems, simultaneously considering many different interacting cells. Here, we tackle this challenge by generating patient-level representations derived from multi-cellular expression context measured with scRNA-seq of tissues. We develop PaSCient, a novel model that employs a multi-level representation learning paradigm and provides importance scores at the individual cell and gene levels for fine-grained analysis across multiple cell types and gene programs characteristic of a given disease. We apply PaSCient to learn a disease model across a large-scale scRNA-seq atlas of 24.3 million cells from over 5,000 patients. Comprehensive and rigorous benchmarking demonstrates the superiority of PaSCient in disease classification and its multiple downstream applications, including dimensionality reduction, gene/cell type prioritization, and patient subgroup discovery.

bioinformatics↗

Decoding sequence determinants of gene expression in diverse cellular and disease states

Sequence-to-function models that predict gene expression from genomic DNA sequence have proven valuable for many biological tasks, including understanding cis-regulatory syntax and interpreting non-coding genetic variants. However, current state-of-the-art models have been trained largely on bulk expression profiles from healthy tissues or cell lines, and have not learned the properties of precise cell types and states that are captured in large-scale single-cell transcriptomic datasets. Thus, they lack the ability to perform these tasks at the resolution of specific cell types or states across diverse tissue and disease contexts. To address this gap, we present Decima, a model that predicts the cell type- and condition- specific expression of a gene from its surrounding DNA sequence. Decima is trained on single-cell or single-nucleus RNA sequencing data from over 22 million cells, and successfully predicts the cell type-specific expression of unseen genes based on their sequence alone. Here, we demonstrate Decimas ability to reveal the cis-regulatory mechanisms driving cell type-specific gene expression and its changes in disease, to predict non-coding variant effects at cell type resolution, and to design regulatory DNA elements with precisely tuned, context-specific functions.

genomics↗

A high-throughput phenotypic screen combined with an ultra-large-scale deep learning-based virtual screening reveals novel scaffolds of antibacterial compounds

The proliferation of multi-drug-resistant bacteria underscores an urgent need for novel antibiotics. Traditional discovery methods face challenges due to limited chemical diversity, high costs, and difficulties in identifying structurally novel compounds. Here, we explore the integration of small molecule high-throughput screening with a deep learning-based virtual screening approach to uncover new antibacterial compounds. Leveraging a diverse library of nearly 2 million small molecules, we conducted comprehensive phenotypic screening against a sensitized Escherichia coli strain that, at a low hit rate, yielded thousands of hits. We trained a deep learning model, GNEprop, to predict antibacterial activity, ensuring robustness through out-of-distribution generalization techniques. Virtual screening of over 1.4 billion compounds identified potential candidates, of which 82 exhibited antibacterial activity, illustrating a 90X improved hit rate over the high-throughput screening experiment GNEprop was trained on. Importantly, a significant portion of these newly identified compounds exhibited high dissimilarity to known antibiotics, indicating promising avenues for further exploration in antibiotic discovery.

bioinformatics↗

Scalable querying of human cell atlases via a foundational model reveals commonalities across fibrosis-associated macrophages

Single-cell RNA-seq (scRNA-seq) studies have profiled over 100 million human cells across diseases, developmental stages, and perturbations to date. A singular view of this vast and growing expression landscape could help reveal novel associations between cell states and diseases, discover cell states in unexpected tissue contexts, and relate in vivo cells to in vitro models. However, these require a common, scalable representation of cell profiles from across the body, a general measure of their similarity, and an efficient way to query these data. Here, we present SCimilarity, a metric learning framework to learn and search a unified and interpretable representation that annotates cell types and instantaneously queries for a cell state across tens of millions of profiles. We demonstrate SCimilarity on a 22.7 million cell corpus assembled across 399 published scRNA-seq studies, showing accurate integration, annotation and querying. We experimentally validated SCimilarity by querying across tissues for a macrophage subset originally identified in interstitial lung disease, and showing that cells with similar profiles are found in other fibrotic diseases, tissues, and a 3D hydrogel system, which we then repurposed to yield this cell state in vitro. SCimilarity serves as a foundational model for single cell gene expression data and enables researchers to query for similar cellular states across the entire human body, providing a powerful tool for generating novel biological insights from the growing Human Cell Atlas.

bioinformatics↗