bioRxiv ScienceSearch

Biology subjects

Yuan, G.-C.

Publications and source records attributed to Yuan, G.-C..

8 recordsLinked to original sources

Revealing the critical regulators of cell identity in the mouse cell atlas

Recent progress in single-cell technologies has enabled the identification of all major cell types in mouse. However, for most cell types, the regulatory mechanism underlying their identity remains poorly understood. By computational analysis of the recently published mouse cell atlas data, we have identified 202 gene regulatory networks whose activities are highly variable across different cell types, and more importantly, predicted a small set of essential regulators for each of over 800 cell types in mouse for the first time. Systematic validation by automated literature- and data-mining provides strong additional support for our predictions. Thus, these predictions serve as a valuable resource that would be useful for the broad biological community. Finally, we have built a user-friendly, interactive, web-portal to enable users to navigate this mouse cell network atlas.

bioinformatics

STREAM: Single-cell Trajectories Reconstruction, Exploration And Mapping of omics data

Single-cell transcriptomic assays have enabled the de novo reconstruction of lineage differentiation trajectories, along with the characterization of cellular heterogeneity and state transitions. Several methods have been developed for reconstructing developmental trajectories from single-cell transcriptomic data, but efforts on analyzing single-cell epigenomic data and on trajectory visualization remain limited. Here we present STREAM, an interactive pipeline capable of disentangling and visualizing complex branching trajectories from both single-cell transcriptomic and epigenomic data.

genomics

Decomposing spatially dependent and cell type specific contributions to cellular heterogeneity

Both the intrinsic regulatory network and spatial environment are contributors of cellular identity and result in cell state variations. However, their individual contributions remain poorly understood. Here we present a systematic approach to integrate both sequencing-and imaging-based single-cell transcriptomic profiles, thereby combining whole-transcriptomic and spatial information from these assays. We applied this approach to dissect the cell-type and spatial domain associated heterogeneity within the mouse visual cortex region. Our analysis identified distinct spatially associated signatures within glutamatergic and astrocyte cell compartments, indicating strong interactions between cells and their spatial environment. Using these signatures as a guide to analyze single cell RNAseq data, we identified previously unknown, but spatially associated subpopulations. As such, our integrated approach provides a powerful tool for dissecting the roles of intrinsic regulatory networks and spatial environment in the maintenance of cellular states.

bioinformatics

A cluster-aware, weighted ensemble clustering method for cell-type detection

Single-cell analysis is a powerful tool for dissecting the cellular composition within a tissue or organ. However, it remains difficult to detect rare and common cell types at the same time. Here we present a new computational method, called GiniClust2, to overcome this challenge. GiniClust2 combines the strengths of two complementary approaches, using the Gini index and Fano factor, respectively, through a cluster-aware, weighted ensemble clustering technique. GiniClust2 successfully identifies both common and rare cell types in diverse datasets, outperforming existing methods. GiniClust2 is scalable to very large datasets.

bioinformatics

Haystack: systematic analysis of the variation of epigenetic states and cell-type specific regulatory elements

MotivationWith the increasing amount of genomic and epigenomic data in the public domain, a pressing challenge is how to integrate these data to investigate the role of epigenetic mechanisms in regulating gene expression and maintenance of cell-identity. To this end, we have implemented a computational pipeline to systematically study epigenetic variability and uncover regulatory DNA sequences that play a role in gene regulation.\n\nResultsHaystack is a bioinformatics pipeline to characterize hotspots of epigenetic variability across different cell-types as well as cell-type specific cis-regulatory elements along with their corresponding transcription factors. Our approach is generally applicable to any epigenetic mark and provides an important tool to investigate cell-type identity and the mechanisms underlying epigenetic switches during development. Additionally, we make available a set of precomputed tracks for a number of epigenetic marks across several cell types. These precomputed results may be used as an independent resource for functional annotation of the human genome.\n\nAvailabilityThe Haystack pipeline is implemented as an open-source, multiplatform, Python package called haystack_bio available at https://github.com/pinellolab/haystack_bio.\n\nContactlpinello@mgh.harvard.edu, gcyuan@jimmy.harvard.edu

bioinformatics

Dissecting super-enhancer hierarchy based on chromatin interactions

Recent studies have highlighted super-enhancers (SEs) as important regulatory elements for gene expression, but their intrinsic properties remain incompletely characterized. Through an integrative analysis of Hi-C and ChIP-seq data, we find that a significant fraction of SEs are hierarchically organized, containing both hub and non-hub enhancers. Hub enhancers share similar histone marks with non-hub enhancers, but are distinctly associated with cohesin and CTCF binding sites and disease-associated genetic variants. Genetic ablation of hub enhancers results in profound defects in gene activation and local chromatin landscape. As such, hub enhancers are the major constituents responsible for SE functional and structural organization.

bioinformatics

Challenges And Emerging Directions In Single-Cell Analysis

Single-cell analysis is a rapidly evolving approach to characterize genome-scale molecular information at the individual cell level. Development of single-cell technologies and computational methods has enabled systematic investigation of cellular heterogeneity in a wide range of tissues and cell populations, yielding fresh insights into the composition, dynamics, and regulatory mechanisms of cell states in development and disease. Despite substantial advances, significant challenges remain in the analysis, integration, and interpretation of single-cell omics data. Here, we discuss the state of the field and recent advances, and look to future opportunities.

genomics

Integrated Computational Guide Design, Execution, And Analysis Of Arrayed And Pooled CRISPR Genome Editing Experiments

CRISPR genome editing experiments offer enormous potential for the evaluation of genomic loci using arrayed single guide RNAs (sgRNAs) or pooled sgRNA libraries. Numerous computational tools are available to help design sgRNAs with optimal on-target efficiency and minimal off-target potential. In addition, computational tools have been developed to analyze deep sequencing data resulting from genome editing experiments. However, these tools are typically developed in isolation and oftentimes not readily translatable into laboratory-based experiments. Here we present a protocol that describes in detail both the computational and benchtop implementation of an arrayed and/or pooled CRISPR genome editing experiment. This protocol provides instructions for sgRNA design with CRISPOR, experimental implementation, and analysis of the resulting high-throughput sequencing data with CRISPResso. This protocol allows for design and execution of arrayed and pooled CRISPR experiments in 4-5 weeks by non-experts as well as computational data analysis in 1-2 days that can be performed by both computational and non-computational biologists alike.

molecular biology