bioRxiv Science⌕ Search

Biology subjects

Bao, S. C.

Publications and source records attributed to Bao, S. C..

4 recordsLinked to original sources

Combinatorial epigenomic patterns define regulatory programs underlying disease heterogeneity

Complex diseases exhibit substantial variation in clinical presentation and outcome despite shared diagnoses. Current genetic models typically represent inherited risk as a single additive liability, obscuring the diverse biological mechanisms through which variants influence disease. Here, we show that disease-associated variants are organized into recurrent regulatory programs that reveal latent disease mechanisms. Using genome-scale epigenomic maps across human tissues and cell states, we identify regulatory programs that partition disease-associated variants without phenotypic or disease-specific priors. Variants assigned to different programs exert distinct biological effects that translate into divergent clinical outcomes. In type 2 diabetes, these programs reveal previously unrecognized disease subtypes with opposing cardiometabolic profiles that stratify future risk of myocardial infarction and non-alcoholic fatty liver disease. Together, our findings establish that inherited disease risk is organized into latent regulatory programs, revealing a fundamental layer of biological heterogeneity underlying complex disease.

genomics↗

Protocol-dependent cardiomyocyte states determine disease modelling capacity of human iPSCs

Human induced pluripotent stem cell-derived cardiomyocytes (iPSC-CMs) are widely used to model cardiovascular disease, yet numerous differentiation protocols generate cardiomyocytes with heterogeneous molecular and functional properties, complicating experimental design. Here we systematically compare sixteen commonly used cardiomyocyte differentiation protocols and characterize their resulting cell states using single-nucleus RNA sequencing, functional phenotyping and computational integration with human genetic data. Despite similar cardiomyocyte yields, protocols produced distinct transcriptional programs, subtype compositions and physiological properties. By integrating protocol-specific gene expression signatures with genome-wide association studies of cardiovascular traits, we identify cardiomyocyte states enriched for genetic architectures underlying specific diseases. These analyses accurately predict protocols most suitable for modelling particular disease contexts, including electrophysiological defects associated with Brugada syndrome and metabolic vulnerability relevant to myocardial infarction. Our results demonstrate that differentiation protocols encode biologically distinct cardiomyocyte states with differential disease relevance and establish a framework for aligning stem-cell differentiation strategies with human complex trait genetics to guide model selection. This approach enables rational design of iPSC-based disease models and highlights how population-scale genetic data can inform experimental systems in stem cell biology.

systems biology↗

A robust unsupervised clustering approach for high-dimensional biological imaging data reveals shared drug-induced morphological signatures

Modern biology increasingly relies on large-scale screening to generate high dimensional datasets with potential to accelerate discovery. However, analysing these complex datasets remains challenging, particularly in applications where the underlying structure and groupings are unknown, and high dimensionality introduces noise and artifacts that make follow up studies difficult to prioritise. Here, we present an unsupervised consensus clustering tool that quantifies biologically meaningful patterns based on multi-scale data organisation to guide decision-making in high-throughput screening. Using large-scale drug screening data in cancer cell lines and bacterium model, we demonstrate its ability to use diverse data inputs to prioritize robust drug clusters with shared biological mechanisms and conserved drug responses. This method addresses key limitations associated with prioritising robust, actionable hits from scalable screening data.

bioinformatics↗

A community-oriented, data-driven resource to improve protocol design for cardiac modelling from human pluripotent stem cells

Protocol design and benchmarking is central to optimising model development using human pluripotent stem cell derived cardiomyocytes (hPSC-CMs). By applying data mining to decades of research and hundreds of peer reviewed studies, we evaluate how protocol variables associate with common properties of cardiac functional and physiological maturation. This resource is publicly accessible through CMPortal, a community-oriented website that provides data-driven tools for researchers to navigate leverage decades of knowledge for benchmarking protocol designs and outcomes for their dedicated applications in developmental biology, disease modelling, and drug screening.

developmental biology↗