bioRxiv Science⌕ Search

Biology subjects

Tan, L. M.

Publications and source records attributed to Tan, L. M..

3 recordsLinked to original sources

Single-cell analysis of human diversity in circulating immune cells

Lack of diversity and proportionate representation in genomics datasets and databases contributes to inequity in healthcare outcomes globally1,2. The relationships of human diversity with biological and biomedical phenotypes are pervasive3, yet remain understudied, particularly in a single-cell genomics context. Here we present the Asian Immune Diversity Atlas (AIDA), a multi-national single-cell RNA-sequencing (scRNA-seq) healthy reference atlas of human immune cells. AIDA comprises 1,265,624 circulating immune cells from 619 healthy donors and 6 controls, spanning 7 population groups across 5 countries. AIDA is one of the largest healthy blood datasets in terms of number of cells, and also the most diverse in terms of number of population groups. Though population groups are frequently compared at the continental level, we identified a pervasive impact of sub-continental diversity on cellular and molecular properties of immune cells. These included cell populations and genes implicated in disease risk and pathogenesis as well as those relevant for diagnostics. We detected single-cell signatures of human diversity not apparent at the level of cell types, as well as modulation of the effects of age and sex by self-reported ethnicity. We discovered functional genetic variants influencing cell type-specific gene expression, including context-dependent effects, which were under-represented in analyses of non-Asian population groups, and which helped contextualise disease-associated variants. We validated our findings using multiple independent datasets and cohorts. AIDA provides fundamental insights into the relationships of human diversity with immune cell phenotypes, enables analyses of multi-ancestry disease datasets, and facilitates the development of precision medicine efforts in Asia and beyond.

genomics↗

Quantification of the escape from X chromosome inactivation with the million cell-scale human single-cell omics datasets reveals heterogeneity of escape across cell types and tissues

One of the two X chromosomes of females is silenced through X chromosome inactivation (XCI) to compensate for the difference in the dosage between sexes. Among the X-linked genes, several genes escape from XCI, which could contribute to the differential gene expression between the sexes. However, the differences in the escape across cell types and tissues are still poorly characterized because no methods could directly evaluate the escape under a physiological condition at the cell-cluster resolution with versatile technology. Here, we developed a method, single-cell Level inactivated X chromosome mapping (scLinaX), which directly quantifies relative gene expression from the inactivated X chromosome with droplet-based single-cell RNA-sequencing (scRNA-seq) data. The scLinaX and differentially expressed genes analyses with the scRNA-seq datasets of [~]1,000,000 blood cells consistently identified the relatively strong degree of escape in lymphocytes compared to myeloid cells. An extension of scLinaX for multi-modal datasets, scLinaX-multi, suggested a stronger degree of escape in lymphocytes than myeloid cells at the chromatin-accessibility level with a 10X multiome dataset. The scLinaX analysis with the human multiple-organ scRNA-seq datasets also identified the relatively strong degree of escape from XCI in lymphoid tissues and lymphocytes. Finally, effect size comparisons of genome-wide association studies between sexes identified the larger effect sizes of the PRKX gene locus-lymphocyte counts association in females than males. This could suggest evidence of the underlying impact of escape on the genotype-phenotype association in humans. Overall, scLinaX and the quantified catalog of escape identified the heterogeneity of escape across cell types and tissues and would contribute to expanding the current understanding of the XCI, escape, and sex differences in gene regulation.

genomics↗

Monopogen: single nucleotide variant calling from single cell sequencing

Distinguishing how genetics impact cellular processes can improve our understanding of variable risk for diseases. Although single-cell omics have provided molecular characterization of cell types and states on diverse tissue samples, their genetic ancestry and effects on cellular molecular traits are largely understudied. Here, we developed Monopogen, a computational tool enabling researchers to detect single nucleotide variants (SNVs) from a variety of single cell transcriptomic and epigenomic sequencing data. It leverages linkage disequilibrium from external reference panels to identify germline SNVs from sparse sequencing data and uses Monovar to identify novel SNVs at cluster (or cell type) levels. Monopogen can identify 100K~3M germline SNVs from various single cell sequencing platforms (scRNA-seq, snRNA-seq, snATAC-seq etc), with genotyping accuracy higher than 95%, when compared against matched whole genome sequencing data. We applied Monopogen on human retina, normal breast and Asian immune diversity atlases, showing that that derived genotypes enable accurate global and local ancestry inference and identification of admixed samples from ancestrally diverse donors. In addition, we applied Monopogen on ~4M cells from 65 human heart left ventricle single cell samples and identified novel variants associated with cardiomyocyte metabolic levels and epigenomic programs. In summary, Monopogen provides a novel computational framework that brings together population genetics and single cell omics to uncover genetic determinants of cellular quantitative traits.

bioinformatics↗