bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.10.23.563518

scBoolSeq: Linking scRNA-Seq Statistics and Boolean Dynamics

Abstract

Boolean networks are largely employed to model the qualitative dynamics of cell fate processes by describing the change of binary activation states of genes and transcription factors with time. Being able to bridge such qualitative states with quantitative measurements of gene expressions in cells, as scRNA-Seq, is a cornerstone for data-driven model construction and validation. On one hand, scRNA-Seq binarisation is a key step for inferring and validating Boolean models. On the other hand, the generation of synthetic scRNA-Seq data from baseline Boolean models provides an important asset to benchmark inference methods. However, linking characteristics of scRNA-Seq datasets, including dropout events, with Boolean states is a challenging task. We present O_SCPLOWSCC_SCPLOWBO_SCPLOWOOLC_SCPLOWSO_SCPLOWEQC_SCPLOW, a method for the bidirectional linking of scRNA-Seq data and Boolean activation state of genes. Given a reference scRNA-Seq dataset, O_SCPLOWSCC_SCPLOWBO_SCPLOWOOLC_SCPLOWSO_SCPLOWEQC_SCPLOW computes statistical criteria to classify the empirical gene pseudocount distributions as either unimodal, bimodal, or zero-inflated, and fit a probabilistic model of dropouts, with gene-dependent parameters. From these learnt distributions, O_SCPLOWSCC_SCPLOWBO_SCPLOWOOLC_SCPLOWSO_SCPLOWEQC_SCPLOW can perform both binarisation of scRNA-Seq datasets, and generate synthetic scRNA-Seq datasets from Boolean trajectories, as issued from Boolean networks, using biased sampling and dropout simulation. We present a case study demonstrating the application of O_SCPLOWSCC_SCPLOWBO_SCPLOWOOLC_SCPLOWSO_SCPLOWEQC_SCPLOWs binarisation scheme in data-driven model inference. Furthermore, we compare synthetic scRNA-Seq data generated by O_SCPLOWSCC_SCPLOWBO_SCPLOWOOLC_SCPLOWSO_SCPLOWEQC_SCPLOW with BO_SCPLOWOOLC_SCPLOWODE from the same Boolean Network model. The comparison shows that our method better reproduces the statistics of real scRNA-Seq datasets, such as the mean-variance and mean-dropout relationships while exhibiting clearly defined trajectories in a two-dimensional projection of the data. Author summaryThe qualitative and logical modeling of cell dynamics has brought precious insight on gene regulatory mechanisms that drive cellular differentiation and fate decisions by predicting cellular trajectories and mutations for their control. However, the design and validation of these models is impeded by the quantitative nature of experimental measurements of cellular states. In this paper, we provide and assess a new methodology, O_SCPLOWSCC_SCPLOWBO_SCPLOWOOLC_SCPLOWSO_SCPLOWEQC_SCPLOW for bridging single-cell level pseudocounts of RNA transcripts with Boolean classification of gene activity levels. Our method, implemented as a Python package, enables both to binarise scRNA-Seq data in order to match quantitative measurements with states of logicals models, and to generate synthetic data from Boolean trajectories in order to benchmark inference methods. We show that O_SCPLOWSCC_SCPLOWBO_SCPLOWOOLC_SCPLOWSO_SCPLOWEQC_SCPLOW accurately captures main statistical features of scRNA-Seq data, including measurement dropouts, improving significantly the state of the art. Overall, scBoolSeq brings a statistically-grounded method for enabling the inference and validation of qualitative models from scRNA-Seq data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Magna Lopez, G., Calzone, L., Zinovyev, A., Pauleve, L.. 2023-10-25. scBoolSeq: Linking scRNA-Seq Statistics and Boolean Dynamics. https://doi.org/10.1101/2023.10.23.563518

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

INFORME: coupling information-theoretic experimental design with nonlinear mixed-effects modeling for efficient observation scheduling

Mathematical models of treatment response can inform individualized therapy, but their calibration often requires longitudinal measurements that are costly, burdensome, and collected on fixed schedules. Such schedules may be inefficient, over-sampling patients whose response is already well characterized while delaying informative measurements for those whose model parameters remain uncertain. We present INFORME (INFORmation-theoretic design with Mixed Effects), a framework that combines Bayesian information-theoretic experimental design with nonlinear mixed-effects modeling to adaptively select each patients next measurement time. Population and response-subgroup parameter distributions learned from an existing cohort provide informative priors, allowing candidate measurement times to be ranked by their expected reduction in patient-specific parameter uncertainty. As observations accumulate, priors can be updated to reflect the response subgroup most consistent with the patients data. We evaluate INFORME in two radiotherapy datasets: 150 synthetic tumor volume trajectories from a hybrid cellular automaton model of prostate cancer spheroids (HD1) and longitudinal tumor volumes from 39 patients with head-and-neck cancer (HD2). In HD1, population priors allowed omission of both pretreatment scans, while adaptive scheduling reduced the protocol from nine scans to three or four, with the response group identified from a single post-treatment scan on day 27. In HD2, the adaptive schedule used three scans instead of six and improved prediction by delaying the first on-treatment scan from week 1 to week 2, avoiding transient dynamics that produced false-positive and false-negative response projections. Across both datasets, the adaptive schedules used a mean of 2.7 scans in stead of seven and advanced completion of the patient-specific prediction by a mean of 15.5 days (95% CI, 6.7-24.3) relative to the equidistant protocol, while treatment duration remained unchanged. INFORME therefore reduces measurement burden and accelerates patient-specific prediction by concentrating observations at times that are most informative for model calibration.

systems biology↗

Sobetirome, a thyroid hormone receptor beta agonist, is a potential therapeutic agent for pulmonary fibrosis

Idiopathic pulmonary fibrosis (IPF) is a progressive and fatal disease with limited treatment options. Our group previously identified the antifibrotic potential of thyroid hormone, triiodothyronine (T3); however, clinical translation of thyroid hormone therapy is limited by its systemic adverse effects. In this study, we investigate whether sobetirome, a selective and well tolerated thyroid hormone receptor beta (THRB) agonist, offers antifibrotic benefits of thyroid hormone while minimizing systemic toxicity. Our study reveals that sobetirome, administered via intraperitoneal or inhalational routes, effectively mitigates bleomycin-induced pulmonary fibrosis in mice, with no evidence of toxicity. We identified that sobetirome restores mitochondrial homeostasis via activating the THRB-PPARGC1a axis. This protects alveolar type II epithelial cells from injury-induced apoptosis while selectively inducing apoptosis and metabolic reprogramming in apoptosis resistant IPF fibroblasts. Cell-specific deletion of Ppargc1a in either alveolar epithelial cells or fibroblasts abolishes sobetirome-mediated protection, establishing PPARGC1a as an essential mediator of therapeutic response. Importantly, sobetirome reverses fibrosis-associated transcriptional programs in human IPF lung tissue, reducing expression of key fibrosis-associated genes, including collagen I alpha 1 (COL1A1), collagen III alpha 1 (COL3A1), periostin (POSTN), cathepsin K (CTSK), and Chitinase 3 Like 1 (CHI3L1), while promoting extracellular matrix remodeling, epithelial restoration, and tissue homeostasis. Collectively, our findings identify THRB activation as a novel metabolic strategy for reversing pulmonary fibrosis. Across complementary in vitro, in vivo, and human ex vivo models, sobetirome restores mitochondrial function, modulates apoptotic pathways in pathogenic cells, and promotes fibrosis resolution, highlighting its potential as a lung-targeted therapeutic approach for IPF and other fibrotic lung diseases.

systems biology↗

Mechanistic modeling of bacterial translation initiation across growth conditions

Translation frequency in bacteria depends on how ribosomes, mRNAs, and initiation factors are allocated across growth conditions. Here, we developed a mechanistic ODE-based model of Escherichia coli translation that represents initiation, elongation, termination, and coupled auxiliary processes. Growth-dependent abundances were derived from physiological relationships and reprocessed omics data, and simulated outputs were compared with translation-frequency and active-ribosome references. The model predicts a continuous shift from complex-formation-limited toward ribosome-limited behavior as growth increases. This shift is characterized by a decline in free-ribosome abundance, whereas initiation-factor pools remain largely unbound and do not become depleted in parallel. Together with the implemented IF-dependent kinetic term, this preserved availability provides a model-internal route through which productive initiation can be maintained despite increasing ribosome utilization. Consistently, transcript-wide ribosome loading remains below its theoretical maximum, while COG-level simulations reveal distinct sector-specific translation-frequency trajectories. The study therefore provides a resource-allocation framework for interpreting how mRNA--ribosome interactions shape bacterial translation across growth conditions.

systems biology↗