bioRxiv ScienceSearch

Biology subjects

Lee, C.

Publications and source records attributed to Lee, C..

22 records · Page 2Linked to original sources

Comprehensive statistical inference of the clonal structure of cancer from multiple biopsies

A comprehensive characterization of tumor genetic heterogeneity is critical for understanding how cancers evolve and escape treatment. Although many algorithms have been developed for capturing tumor heterogeneity, they are designed for analyzing either a single type of genomic aberration or individual biopsies. Here we present THEMIS (Tumor Heterogeneity Extensible Modeling via an Integrative System), which allows for the joint analysis of different types of genomic aberrations from multiple biopsies taken from the same patient, using a dynamic graphical model. Simulation experiments demonstrate higher accuracy of THEMIS over its ancestor, TITAN. The heterogeneity analysis results from THEMIS are validated with single cell DNA sequencing from a clinical tumor biopsy. When THEMIS is used to analyze tumor heterogeneity among multiple biopsies from the same patient, it helps to reveal the mutation accumulation history, track cancer progression, and identify the mutations related to treatment resistance. We implement our model via an extensible modeling platform, which makes our approach open, reproducible, and easy for others to extend.

bioinformatics

Comprehensive single cell transcriptional profiling of a multicellular organism by combinatorial indexing

Conventional methods for profiling the molecular content of biological samples fail to resolve heterogeneity that is present at the level of single cells. In the past few years, single cell RNA sequencing has emerged as a powerful strategy for overcoming this challenge. However, its adoption has been limited by a paucity of methods that are at once simple to implement and cost effective to scale massively. Here, we describe a combinatorial indexing strategy to profile the transcriptomes of large numbers of single cells or single nuclei without requiring the physical isolation of each cell (Single cell Combinatorial Indexing RNA-seq or sci-RNA-seq). We show that sci-RNA-seq can be used to efficiently profile the transcriptomes of tens-of-thousands of single cells per experiment, and demonstrate that we can stratify cell types from these data. Key advantages of sci-RNA-seq over contemporary alternatives such as droplet-based single cell RNA-seq include sublinear cost scaling, a reliance on widely available reagents and equipment, the ability to concurrently process many samples within a single workflow, compatibility with methanol fixation of cells, cell capture based on DNA content rather than cell size, and the flexibility to profile either cells or nuclei. As a demonstration of sci-RNA-seq, we profile the transcriptomes of 42,035 single cells from C. elegans at the L2 stage, effectively 50-fold \"shotgun cellular coverage\" of the somatic cell composition of this organism at this stage. We identify 27 distinct cell types, including rare cell types such as the two distal tip cells of the developing gonad, estimate consensus expression profiles and define cell-type specific and selective genes. Given that C. elegans is the only organism with a fully mapped cellular lineage, these data represent a rich resource for future methods aimed at defining cell types and states. They will advance our understanding of developmental biology, and constitute a major step towards a comprehensive, single-cell molecular atlas of a whole animal.

genomics

Paired CRISPR/Cas9 guide-RNAs enable high-throughput deletion scanning (ScanDel) of a Mendelian disease locus for functionally critical non-coding elements

The extent to which distal non-coding mutations contribute to Mendelian disease remains a major unknown in human genetics. Given that a genes in vivo function can be appropriately modeled in vitro, CRISPR/Cas9 genome editing enables the large-scale perturbation of distal non-coding regions to identify functional elements in their native context. However, early attempts at such screens have relied on one individual guide RNA (gRNA) per cell, resulting in sparse mutagenesis with minimal redundancy across regions of interest. To address this, we developed a system that uses pairs of gRNAs to program thousands of kilobase-scale deletions that scan across a targeted region in a tiling fashion (\"ScanDel\"). As a proof-of-concept, we applied ScanDel to program 4,342 overlapping 1- and 2- kilobase (Kb) deletions that tile a 206 Kb region centered on HPRT1, the gene underlying Lesch-Nyhan syndrome, with median 27-fold redundancy per base. Programmed deletions were functionally assayed by selecting for loss of HPRT1 function with 6-thioguanine. HPRT1 exons served as positive controls, and all were successfully identified as functionally critical by the screen. Remarkably, HPRT1 function appeared robust to deletion of any intergenic or deeply intronic non-coding region across the 206 Kb locus, indicating that proximal regulatory sequences are sufficient for its expression. A sparser mutagenesis screen of the same 206 Kb with individual gRNAs also failed to identify critical distal regulatory elements. Although our screen did find programmed deletions and individual gRNAs with putative functional consequences that targeted exon-proximal non-coding sequences (e.g. the promoter), long-read sequencing revealed that this signal was driven almost entirely by rare, unexpected deletions that extended into exonic sequence. These targeted validation experiments defined a small region surrounding the transcriptional start site as the only non-coding sequence essential to HPRT1 function. Overall, our results suggest that distal regulatory elements are not critical for HPRT1 expression, and underscore the necessity of comprehensive edited-locus genotyping for validating the results of CRISPR screens. The application of ScanDel to additional loci will enable more insight into the extent to which the disruption of distal non-coding elements contributes to Mendelian diseases. In addition, dense, redundant, large-scale deletion scanning with gRNA pairs will facilitate a deeper understanding of endogenous gene regulation in the human genome.

genomics

A synthesis of over 9,000 mass spectrometry experiments reveals the core set of human protein complexes

Macromolecular protein complexes carry out many of the essential functions of cells, and many genetic diseases arise from disrupting the functions of such complexes. Currently there is great interest in defining the complete set of human protein complexes, but recent published maps lack comprehensive coverage. Here, through the synthesis of over 9,000 published mass spectrometry experiments, we present hu.MAP, the most comprehensive and accurate human protein complex map to date, containing >4,600 total complexes, >7,700 proteins and >56,000 unique interactions, including thousands of confident protein interactions not identified by the original publications. hu.MAP accurately recapitulates known complexes withheld from the learning procedure, which was optimized with the aid of a new quantitative metric (k-cliques) for comparing sets of sets. The vast majority of complexes in our map are significantly enriched with literature annotations and the map overall shows improved coverage of many disease-associated proteins, as we describe in detail for ciliopathies. Using hu.MAP, we predicted and experimentally validated candidate ciliopathy disease genes in vivo in a model vertebrate, discovering CCDC138, WDR90, and KIAA1328 to be new cilia basal body/centriolar satellite proteins, and identifying ANKRD55 as a novel member of the intraflagellar transport machinery. By offering significant improvements to the accuracy and coverage of human protein complexes, hu.MAP (http://proteincomplexes.org) serves as a valuable resource for better understanding the core cellular functions of human proteins and helping to determine mechanistic foundations of human disease.

systems biology