bioRxiv Science⌕ Search

Biology subjects

Haber, E.

Publications and source records attributed to Haber, E..

5 recordsLinked to original sources

Uncertainty-aware synthetic lethality prediction with pretrained foundation models

Synthetic lethality (SL) offers a promising paradigm for targeted cancer therapy, yet experimental identification of SL gene pairs remains costly, context-dependent, and biased toward well-studied genes. Existing computational approaches often rely on curated protein-protein interaction (PPI) networks and Gene Ontology (GO) annotations, which limit their ability to generalize to novel genes. Here we introduce CO_SCPLOWILANTROC_SCPLOWO_SCPCAP-C_SCPCAPO_SCPLOWSLC_SCPLOW, a two-stage, graph-free framework that leverages pretrained biological foundation models to predict SL pairs with calibrated uncertainty. In Stage 1, we apply a pretrained single-cell foundation model to bulk RNA-seq profiles of cancer cell lines to obtain context-aware embeddings and perform in silico gene knockouts to generate delta embeddings. These perturbation signals are further conditioned on a data-driven gene prior and supervised with CRISPR viability readouts to learn knockout-aware viability embeddings. In Stage 2, we derive pairwise features from these embeddings and train a lightweight classifier to distinguish SL from non-SL pairs. To enable reliable experimental prioritization, CO_SCPLOWILANTROC_SCPLOWO_SCPCAP-C_SCPCAPO_SCPLOWSLC_SCPLOW incorporates conformal prediction, producing calibrated and interpretable prediction sets that highlight high-confidence SL candidates. Across two evaluation settings, including zero-shot generalization to unseen gene pairs and to unseen genes, ablation analyses show that viability pretraining and the gene prior substantially improve performance while avoiding reliance on PPI and GO features. CO_SCPLOWILANTROC_SCPLOWO_SCPCAP-C_SCPCAPO_SCPLOWSLC_SCPLOW therefore transforms pretrained biological representations into practical, uncertainty-aware hypotheses that support robust and scalable discovery of therapeutic targets.

bioinformatics↗

HEIMDALL: A Modular Framework for Tokenization in Single-Cell Foundation Models

Foundation models for single-cell RNA-sequencing (scRNA-seq) data are emerging as powerful tools for single-cell analysis, yet their performance depends critically on how cells are tokenized into model inputs. Single-cell data lack a canonical tokenization scheme, and many design choices in current single-cell foundation models (scFMs) remain heuristic, entangled, and difficult to evaluate. Here, we introduce HO_SCPLOWEIMDALLC_SCPLOW, a unified framework for dissecting and redesigning tokenizers in scFMs. By decomposing existing tokenization strategies into individual design choices, HO_SCPLOWEIMDALLC_SCPLOW enables attribution of the components that underlie robust generalization, allowing more principled design of improved tokenizers. Combining HO_SCPLOWEIMDALLC_SCPLOW with a minimal transformer backbone, we find that tokenizer design is instrumental for generalization in challenging distribution-shift settings such as cross-tissue, cross-species, and cross-gene-panel cell type classification, as well as reverse perturbation prediction. We show that, while tokenizer choice has little effect in scenarios with matched train and test data, it becomes imperative under distribution shift. Rather than identifying a single globally optimal tokenizer, HO_SCPLOWEIMDALLC_SCPLOW reveals that robust transfer depends on a small number of tokenization design axes - especially gene identity, expression encoding, and ordering - that expose different biological priors to the model. In this sense, universal transferability in scFMs still depends on a non-universal tokenizer interface. Together, these findings establish tokenization as a critical design axis in scFMs and provide design principles and reusable infrastructure for more robust scFMs.

bioinformatics↗

EYKTHYR reveals transcriptional regulators of spatial gene programs

Understanding how transcription factors (TFs) orchestrate gene regulatory networks that define complex tissue structures is central to uncovering tissue organization and disease mechanisms. Although spatial multiome technologies now enable in situ measurement of both transcriptional activity and chromatin accessibility, existing computational methods either overlook spatial tissue context or are hindered by the high dropout rates characteristic of such data. Here, we introduce EO_SCPLOWYKTHYRC_SCPLOW, a computational framework that integrates gene expression and chromatin accessibility within a spatially aware model to identify TFs driving spatial gene programs. EO_SCPLOWYKTHYRC_SCPLOW mitigates dropout effects by leveraging interpretable, low-dimensional embeddings of gene expression and chromatin accessibility - both linear with respect to their input - enabling robust identification and scalable inference of spatial transcriptional regulators. Applied across diverse spatial multiome datasets, EO_SCPLOWYKTHYRC_SCPLOW consistently outperforms existing approaches, accurately identifying TFs that coordinate spatial gene programs in mouse brain development and regulate T-cell states within tumor microenvironments. EO_SCPLOWYKTHYRC_SCPLOW establishes a foundation for decoding how TFs interpret local intercellular signaling to shape tissue structure, offering insights into the regulatory logic underlying spatial organization in health and disease.

bioinformatics↗

POPARI: Modeling multisample variation in spatial transcriptomics

Integrating spatially-resolved transcriptomics (SRT) across biological samples is essential for understanding dynamic changes in tissue architecture and cell-cell interactions in situ. While tools exist for multisample single-cell RNA-seq, methods tailored to multisample SRT remain limited. Here, we introduce PO_SCPLOWOPARIC_SCPLOW, a probabilistic graphical model for factor-based decomposition of multisample SRT that captures condition-specific changes in spatial organization. PO_SCPLOWOPARIC_SCPLOW jointly learns spatial metagenes - linear gene expression programs - and their spatial affinities across samples. Its key innovations include a differential prior to regularize spatial accordance and spatial downsampling to enable multiresolution, hierarchical analysis. Simulations show PO_SCPLOWOPARIC_SCPLOW outperforms existing methods on multisample and multi-resolution spatial metrics. Applications to real datasets uncover spatial metagene dynamics, spatial accordance, and cell identities. In mouse brain (STARmap PLUS), PO_SCPLOWOPARIC_SCPLOW identifies spatial metagenes linked to AD; in thymus (Slide-TCR-seq), it captures increasing colocalization of V(D)J recombination and T cell proliferation; and in ovarian cancer (CosMx), it reveals sample-specific malignant-immune interactions. Overall, PO_SCPLOWOPARIC_SCPLOW provides a general, interpretable framework for analyzing variation in multisample SRT.

bioinformatics↗

Unified integration of spatial transcriptomics across platforms

Spatial transcriptomics (ST) has transformed our understanding of tissue architecture and cellular interactions, but integrating ST data across platforms remains challenging due to differences in gene panels, data sparsity, and technical variability. Here, we introduce LO_SCPLOWLOKIC_SCPLOW, a novel framework for integrating imaging-based ST data from diverse platforms without requiring shared gene panels. LO_SCPLOWLOKIC_SCPLOW addresses ST integration through two key alignment tasks: feature alignment across technologies and batch alignment across datasets. Optimal transport-guided feature propagation adjusts data sparsity to match scRNA-seq references through graph-based imputation, enabling single-cell foundation models such as scGPT to generate unified features. Batch alignment then refines scGPT-transformed embeddings, mitigating batch effects while preserving biological variability. Evaluations on mouse brain samples from five different technologies demonstrate that LO_SCPLOWLOKIC_SCPLOW outperforms existing methods and is effective for cross-technology spatial gene program identification and tissue slice alignment. Applying LO_SCPLOWLOKIC_SCPLOW to five ovarian cancer datasets, we identify an integrated gene program indicative of tumor-infiltrating T cells across gene panels. Together, LO_SCPLOWLOKIC_SCPLOW provides a robust foundation for cross-platform ST studies, with the potential to scale to large atlas datasets, enabling deeper insights into cellular organization and tissue environments.

bioinformatics↗