bioRxiv Science⌕ Search

Biology subjects

Burke, D. P.

Publications and source records attributed to Burke, D. P..

5 recordsLinked to original sources

BioReason-Pro: Advancing Protein Function Prediction with Multimodal Biological Reasoning

Protein function annotation is fundamental to understanding biological mechanisms, designing therapeutics, and advancing biomedical research. Current computational methods either rely on shallow sequence similarity or treat function prediction as isolated classification tasks, failing to capture the integrative reasoning across sequence, structure, domains, and interactions that expert biologists perform to infer function. We introduce BioReason-Pro, the first multimodal reasoning large language model (LLM) for protein function prediction that integrates protein embeddings with biological context to generate structured reasoning traces. A key input into BioReason-Pro is the set of GO term predictions made by GO-GPT, our autoregressive transformer that captures hierarchical and cross-aspect dependencies of GO terms. BioReason-Pro is trained via supervised fine-tuning on synthetic reasoning traces generated by GPT-5 for over 130K proteins and further optimized through reinforcement learning. It achieves 73.6% Fmax on GO term prediction and an LLM judge score of 8/10 on functional summaries, substantially outperforming previous methods. Evaluations with human protein experts show that BioReason-Pro annotations are preferred over ground truth UniProt annotations in 79% of cases. Remarkably, BioReason-Pro predicted a novel interaction partner for the renal cancer biomarker RCDG1, which we confirmed in the lab by co-immunoprecipitation. In other binding-partner predictions, its per-residue attention localized to the exact contact residues resolved in cryo-EM structures. Together, GO-GPT and BioReason-Pro establish a framework for protein function prediction that combines precise ontology modeling with interpretable biological reasoning.

molecular biology↗

Scalable probe-based single-cell transcriptional profiling for virtual cell perturbation mapping and synthetic biology phenotyping

Large-scale single-cell transcriptional phenotyping of genetic perturbations (perturb-seq) links genes to phenotypes and should enable virtual cell predictive modeling and cellular engineering. However, current perturb-seq single-cell methods are costly, information sparse and require barcodes for many applications. We developed ProPer-seq, a perturb-seq method that uses multiplexed custom DNA probe panels to measure and phenotype synthetic biology perturbations at single-cell resolution without barcodes, including multidomain proteins and sgRNAs. ProPer-seq faithfully reproduces gold-standard perturb-seq phenotypes while achieving 4-fold cost reduction and 50% increased gene detection per cell. As a scalable fixed-cell profiling method, ProPer-seq enables atlas-scale profiling for virtual-cell initiatives and demonstrates data quality suitable for training and validating predictive models. Lastly, ProPer-seqs targeted detection of modular transgenes enables library-on-library perturbation profiling of combinatorial synthetic protein design spaces. We applied this to 3,550 sgRNA x dCas9 effector combinations as well as 260 CAR x ORF combinations dynamically profiled in primary T cells, revealing principles of transcriptional control and cell state modulation by multidomain synthetic transgenes.

synthetic biology↗

Stack: In-Context Learning of Single-Cell Biology

Foundation models trained on single-cell transcriptomic data offer the promise of identifying and predicting the diversity of cellular phenotypes across species, diseases, and other biological conditions. However, the current models are limited to their supervised training conditions and tasks, which limits their utility for biological discovery. Here, we present SO_SCPLOWTACKC_SCPLOW, a foundation model trained on 149 million uniformly preprocessed human single cells that leverages tabular attention to generate representations for each cell informed by the cells in its context. SO_SCPLOWTACKC_SCPLOW offers substantial improvements for downstream tasks in the zero-shot setting compared to baselines, whether they are zero-shot, fine-tuned, or trained from scratch on the target dataset. SO_SCPLOWTACKC_SCPLOW can perform in-context learning from unlabeled cells representing arbitrary conditions, such as a chemical perturbation or a different donor, and predict the effect of those conditions on a target cell population without requiring data-specific fine-tuning. We apply SO_SCPLOWTACKC_SCPLOW to generate Perturb Sapiens, the first human whole-organism atlas of perturbed cells, spanning 28 tissues, 40 cell types, and 892 drug, cytokine, and genetic perturbations. We validated subsets of Perturb Sapiens using in vitro stimulation profiles. SO_SCPLOWTACKC_SCPLOW uniquely empowers prioritization of donor-specific perturbation effects, a capability we validated in our newly collected DiseasePert-3M data, comprising T cells from 40 donors across 14 diseases, stimulated with 11 cytokines. Overall, SO_SCPLOWTACKC_SCPLOW presents a new modeling framework where cells themselves act as guiding examples at inference time, unlocking general-purpose in-context learning capabilities for single-cell biology.

bioinformatics↗

Predicting cellular responses to perturbation across diverse contexts with STATE

Cellular responses to perturbations are a cornerstone for understanding biological mechanisms and selecting drug targets. While machine learning models offer tremendous potential for predicting perturbation effects, they currently struggle to generalize to unobserved cellular contexts. Here, we introduce SO_SCPLOWTATEC_SCPLOW, a transformer model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. SO_SCPLOWTATEC_SCPLOW predicts perturbation effects across sets of cells and is trained using gene expression data from over 100 million perturbed cells. SO_SCPLOWTATEC_SCPLOW improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling and chemical perturbations with significantly improved accuracy. Using its cell embedding trained on observational data from 167 million cells, SO_SCPLOWTATEC_SCPLOW identified strong perturbations in novel cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that highlights SO_SCPLOWTATEC_SCPLOWs ability to detect cell type-specific perturbation responses, such as cell survival. Overall, the performance and flexibility of SO_SCPLOWTATEC_SCPLOW sets the stage for scaling the development of virtual cell models.

systems biology↗

scBaseCamp: An AI agent-curated, uniformly processed, and continually expanding single cell data repository

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata and the cost of processing reads. Here, we introduce scBaseCount, a single-cell RNA sequencing database that leverages an AI agent to automate discovery and metadata extraction, and standardize data processing. Built by directly mining all 10x Genomics datasets from SRA, scBaseCount is the largest freely accessible public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues, offering an unbiased view of the composition of data within SRA. Uniform processing enables measurement of both intronic and exonic reads, non-coding gene expression and improves alignment across experiments as well as the performance of AI models trained on this phenotypically diverse data. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to curate and autonomously update large biological data repositories.

bioinformatics↗