bioRxiv ScienceSearch

Biology subjects

Ji, H. P.

Publications and source records attributed to Ji, H. P..

7 recordsLinked to original sources

Joint single cell DNA-Seq and RNA-Seq of gastric cancer reveals subclonal signatures of genomic instability and gene expression

Sequencing the genomes of individual cancer cells provides the highest resolution of intratumoral heterogeneity. To enable high throughput single cell DNA-Seq across thousands of individual cells per sample, we developed a droplet-based, automated partitioning technology for whole genome sequencing. We applied this approach on a set of gastric cancer cell lines and a primary gastric tumor. In parallel, we conducted a separate single cell RNA-Seq analysis on these same cancers and used copy number to compare results. This joint study, covering thousands of single cell genomes and transcriptomes, revealed extensive cellular diversity based on distinct copy number changes, numerous subclonal populations and in the case of the primary tumor, subclonal gene expression signatures. We found genomic evidence of positive selection - where the percentage of replicating cells per clone is higher than expected - indicating ongoing tumor evolution. Our study demonstrates that joining single cell genomic DNA and transcriptomic features provides novel insights into cancer heterogeneity and biology. SIGNIFICANCEWe conducted a massively parallel DNA sequencing analysis on a set of gastric cancer cell lines and a primary gastric tumor in combination with a joint single cell RNA-Seq analysis. This joint study, covering thousands of single cell genomes and transcriptomes, revealed extensive cellular diversity based on distinct copy number changes, numerous subclonal populations and in the case of the primary tumor, subclonal gene expression signatures. We found genomic evidence of positive selection where the percentage of replicating cells per clone is higher than expected indicating ongoing tumor evolution. Our study demonstrates that combining single cell genomic DNA and transcriptomic features provides novel insights into cancer heterogeneity and biology.

genomics

CRISPRpic: Fast and precise analysis for CRISPR-induced mutations via prefixed index counting

Analysis of CRISPR-induced mutations at targeted loci can be achieved by PCR amplification followed by massively parallel sequencing. We developed a novel algorithm, called CRISPRpic, to analyze sequencing reads from CRISPR experiments via counting exact-matches and pattern-searching. Compared to other methods that are based on sequence alignment, CRISPRpic provides precise mutation calling and ultrafast analysis of sequencing results. The Python script for CRISPRpic is available at https://github.com/compbio/CRISPRpic.

bioinformatics

Assembly of Mb-size genome segments from linked read sequencing of CRISPR DNA targets

We developed a targeted sequencing method for intact high molecular weight (HMW) DNA targets as large as 0.2 Mb. This process uses HMW DNA isolated from intact cells, custom designed Cas9-guide RNA complexes to generate 0.1 - 0.2 Mb DNA targets, electrophoretic isolation of the DNA targets and sequencing with barcode linked reads. We used alignment methods as well as local assembly of the target regions to identify haplotypes and structural variants (SVs) across multi-Megabase genomic regions. To demonstrate the performance of this approach, we designed three assays that covered a 0.2 Mb region surrounding the BRCA1 gene, a set of 40 overlapping 0.2 Mb targets covering the entire 4-Mb MHC locus, and 18 well-characterized structural variants. Using the highly characterized NA12878 genome, we achieved on-target coverage of more than 50X, while overall whole genome coverage was approximately 4X. We generated haplotypes that completely covered each targeted locus, with a maximum size of 4 Mb (for the MHC region). This method detected structural variants such as deletions and inversions with determination of the exact breakpoints and genotypes. Even breakpoints inside highly homologous segmental duplications are precisely determined with our high-quality assemblies. Overall, this is a new method to sequence large DNA segments.

genomics

Haplotype-resolved and integrated genome analysis of ENCODE cell line HepG2

The HepG2 cancer cell line is one of the most widely-used biomedical research and one of the main cell lines of ENCODE. Vast numbers of functional genomics and epigenomics datasets have been produced to characterize its biology. However, the correct interpretation such data requires an understanding of the cell lines genome sequence and genome structure. Using a variety of sequencing and analysis methods, we identified a wide spectrum of HepG2 genome characteristics: copy numbers of chromosomal segments, SNVs and Indels (corrected for aneuploidy), phased haplotypes extending to entire chromosome arms, loss of heterozygosity, retrotransposon insertions, structural variants (SVs) including complex and somatic genomic rearrangements. We also identified allele-specific expression and DNA methylation genome-wide and assembled an allele-specific CRISPR/Cas9 targeting map.\n\nSIGNIFICANCEHaplotype-resolved and comprehensive whole-genome analysis of a widely-used cell line for cancer research and ENCODE, HepG2, serves as an essential resource for unlocking complex cancer gene regulation using a genome-integrated framework and also provides genomic context for the analysis of ~1,000 functional datasets to date on ENCODE for biological discovery. We also demonstrate how deeper insights into genomic regulatory complexity are gained by adopting a genome-integrated framework.

genomics

SVEngine: an efficient and versatile simulator of genome structural variations with features of cancer clonal evolution

BackgroundSimulating genome sequence data with features can facilitate the development and benchmarking of structural variant analysis programs. However, there are a limited number of data simulators that provide structural variants in silico. Moreover, there are a paucity of programs that generate structural variants with different allelic fraction and haplotypes.\n\nFindingsWe developed SVEngine, an open source tool to address this need. SVEngine simulates next generation sequencing data with embedded structural variations. As input, SVEngine takes template haploid sequences (FASTA) and an external variant file, a variant distribution file and/or a clonal phylogeny tree file (NEWICK) as input. Subsequently, it simulates and outputs sequence contigs (FASTAs), sequence reads (FASTQs) and/or post-alignment files (BAMs). All of the files contain the desired variants, along with BED files containing the ground truth. SVEngines flexible design process enables one to specify size, position, and allelic fraction for deletion, insertion, duplication, inversion and translocation variants. Finally, SVEngine simulates sequence data that replicates the characteristics of a sequencing library with mixed sizes of DNA insert molecules. To improve the compute speed, SVEngine is highly parallelized to reduce the simulation time.\n\nConclusionsWe demonstrated the versatile features of SVEngine and its improved runtime comparisons with other available simulators. SVEngines features include the simulation of locus-specific variant frequency designed to mimic the phylogeny of cancer clonal evolution. We validated the accuracy of the simulations. Our evaluation included checking various sequencing mapping features such as coverage change, read clipping, insert size shift and neighbouring hanging read pairs for representative variant types. SVEngine is implemented as a standard Python package and is freely available for academic use at: https://bitbucket.org/charade/svengine.

bioinformatics

Single-cell transcriptome analysis identifies distinct cell types and intercellular niche signaling in a primary gastric organoid model

The diverse cellular milieu of the gastric tissue microenvironment plays a critical role in normal tissue homeostasis and tumor development. However, few cell culture model can recapitulate the tissue microenvironment and intercellular signaling in vitro. Here we applied an air-liquid interface method to culture primary gastric organoids that contains epithelium with endogenous stroma. To characterize the microenvironment and intercellular signaling in this model, we analyzed the transcriptomes of over 5,000 individual cells from primary gastric organoids cultured at different time points. We identified epithelial cells, fibroblasts and macrophages at the early stage of organoid formation, and revealed that macrophages were polarized towards wound healing and tumor promotion. The organoids maintained both epithelial and fibroblast lineages during the course of time, and a subset of cells in both lineages expressed the stem cell marker Lgr5. We identified that Rspo3 was specifically expressed in the fibroblast lineage, providing an endogenous source of the R-spondin to activate Wnt signaling. Our studies demonstrate that air-liquid-interface-derived organoids provide a novel platform to study intercellular signaling and immune response in vitro.

genomics

A Robust Targeted Sequencing Approach For Low Input And Variable Quality DNA From Clinical Samples

Next-generation deep sequencing of gene panels is being adopted as a diagnostic test to identify actionable mutations in cancer patient samples. However, clinical samples, such as formalin-fixed, paraffin-embedded specimens, frequently provide low quantities of degraded, poor quality DNA. To overcome these issues, many sequencing assays rely on extensive PCR amplification leading to an accumulation of bias and artifacts. Thus, there is a need for a targeted sequencing assay that performs well with DNA of low quality and quantity without relying on extensive PCR amplification. We evaluate the performance of a targeted sequencing assay based on Oligonucleotide Selective Sequencing, which permits the enrichment of genes and regions of interest and the identification of sequence variants from low amounts of damaged DNA. This assay utilizes a repair process adapted to clinical FFPE samples, followed by adaptor ligation to single stranded DNA and a primer-based capture technique. Our approach generates sequence libraries of high fidelity with reduced reliance on extensive PCR amplification - this facilitates the accurate assessment of copy number alterations in addition to delivering accurate SNV and indel detection. We apply this method to capture and sequence the exons of a panel of 130 cancer-related genes, from which we obtain high read coverage uniformity across the targeted regions at starting input DNA amounts as low as 10 ng per sample. We further demonstrate the performance of this assay using a series of reference DNA samples, and by identifying sequence variants in DNA from matched clinical samples originating from different tissue types.

genomics