bioRxiv Science⌕ Search

Biology subjects

Hao, S. P.

Publications and source records attributed to Hao, S. P..

2 recordsLinked to original sources

Genular: An Integrated Platform for Defining Cellular Identity and Function through Single-Cell Gene Expression and Multi-Domain Biological Data

Accurately defining cellular identity and function is essential for advancing immunology, understanding disease mechanisms, and developing targeted therapies. However, current bioinformatics tools are limited in their ability to integrate and analyze the vast and diverse single-cell RNA sequencing (scRNA-seq) datasets available, hindering the comprehensive capture of cellular heterogeneity and the identification of subtle genetic changes across immune states, differentiation pathways, and tissue contexts. To overcome these challenges, we introduce genular, an open-source platform that unifies gene expression data analysis across diverse cell types by integrating scRNA-seq data with extensive genomic and proteomic information from 16 databases, including NCBI Gene, Human Protein Atlas, STRING, and UniProt. genular aggregates data from more than 2,893 scRNA-seq experiments, encompassing over 74.5 million unique cells across various tissues and conditions. A key feature of genular is calculating a cell marker score for each gene, enabling the quantification of gene expression across cells to derive unique profiles specific to cell types, states, and lineages. Using genular, we differentiate T cell memory states, map differentiation profiles by tracking gene expression changes as monocytes mature into macrophages and lymphoid progenitor cells develop into T cells, and capture tissue-specific reprogramming of macrophages, revealing distinct gene expression profiles that enable specialized functions in different tissues. By integrating scRNA-seq data with multi-domain biological information and employing advanced statistical methodologies, genular provides a scalable platform that accurately defines cellular identities, functional states, and differentiation pathways. This comprehensive approach facilitates breakthroughs in immunology, gene regulation, cellular differentiation, and disease research, enabling a deeper understanding of immune cell functions and their roles in health and disease.

bioinformatics↗

A harmonized public resource of deeply sequenced diverse human genomes

Underrepresented populations are often excluded from genomic studies due in part to a lack of resources supporting their analyses. The 1000 Genomes Project (1kGP) and Human Genome Diversity Project (HGDP), which have recently been sequenced to high coverage, are valuable genomic resources because of the global diversity they capture and their open data sharing policies. Here, we harmonized a high quality set of 4,094 whole genomes from HGDP and 1kGP with data from the Genome Aggregation Database (gnomAD) and identified over 153 million high-quality SNVs, indels, and SVs. We performed a detailed ancestry analysis of this cohort, characterizing population structure and patterns of admixture across populations, analyzing site frequency spectra, and measuring variant counts at global and subcontinental levels. We also demonstrate substantial added value from this dataset compared to the prior versions of the component resources, typically combined via liftover and variant intersection; for example, we catalog millions of new genetic variants, mostly rare, compared to previous releases. In addition to unrestricted individual-level public release, we provide detailed tutorials for conducting many of the most common quality control steps and analyses with these data in a scalable cloud-computing environment and publicly release this new phased joint callset for use as a haplotype resource in phasing and imputation pipelines. This jointly called reference panel will serve as a key resource to support research of diverse ancestry populations.

genomics↗