bioRxiv ScienceSearch

Biology subjects

Chan, T.

Publications and source records attributed to Chan, T..

4 recordsLinked to original sources

Resource: Scalable whole genome sequencing of 40,000 single cells identifies stochastic aneuploidies, genome replication states and clonal repertoires

Essential features of cancer tissue cellular heterogeneity such as negatively selected genome topologies, sub-clonal mutation patterns and genome replication states can only effectively be studied by sequencing single-cell genomes at scale and high fidelity. Using an amplification-free single-cell genome sequencing approach implemented on commodity hardware (DLP+) coupled with a cloud-based computational platform, we define a resource of 40,000 single-cell genomes characterized by their genome states, across a wide range of tissue types and conditions. We show that shallow sequencing across thousands of genomes permits reconstruction of clonal genomes to single nucleotide resolution through aggregation analysis of cells sharing higher order genome structure. From large-scale population analysis over thousands of cells, we identify rare cells exhibiting mitotic mis-segregation of whole chromosomes. We observe that tissue derived scWGS libraries exhibit lower rates of whole chromosome anueploidy than cell lines, and loss of p53 results in a shift in event type, but not overall prevalence in breast epithelium. Finally, we demonstrate that the replication states of genomes can be identified, allowing the number and proportion of replicating cells, as well as the chromosomal pattern of replication to be unambiguously identified in single-cell genome sequencing experiments. The combined annotated resource and approach provide a re-implementable large scale platform for studying lineages and tissue heterogeneity.

genomics

ATRX, DAXX or MEN1 mutant pancreatic neuroendocrine tumors are a distinct alpha-cell signature subgroup

The most commonly mutated genes in pancreatic neuroendocrine tumors (PanNETs) are ATRX, DAXX, and MEN1. Little is known about the cells-of-origin for non-functional neuroendocrine tumors. Here, we genotyped 64 PanNETs for mutations in ATRX, DAXX, and MEN1 and found 37 tumors (58%) carry mutations in these three genes (A-D-M mutant PanNETs) and this correlates with a worse clinical outcome than tumors carrying the wild-type alleles of all three genes (A-D-M WT PanNETs). We performed RNA sequencing and DNA-methylation analysis on 33 randomly selected cases to reveal two distinct subgroups with one group consisting entirely of A-D-M mutant PanNETs. Two biomarkers differentiating A-D-M mutant from A-D-M WT PanNETs were high ARX gene expression and low PDX1 gene expression with PDX1 promoter hyper-methylation in the A-D-M mutant PanNETs. Moreover, A-D-M mutant PanNETs had a gene expression signature related to that of alpha cells (pval < 0.009) of pancreatic islets including increased expression of HNF1A and its transcriptional target genes. This gene expression profile suggests that A-D-M mutant PanNETs originate from or transdifferentiate into a distinct cell type similar to alpha cells.

cancer biology

Multi-dimensional genomic analysis of myoepithelial carcinoma identifies prevalent oncogenic gene fusions

Myoepithelial carcinoma (MECA) is an aggressive type of salivary gland cancer with largely unknown molecular features. MECA may arise de novo or result from oncogenic transformation of a pre-existing pleomorphic adenoma (MECA ex-PA). We comprehensively analyzed the molecular alterations in MECA with integrated genomic analyses. We identified a low mutational load (0.5/MB), but a high prevalence of fusion oncogenes (28/40 tumors; 70%). We found FGFR1-PLAG1 in 7 (18%) cases, and the novel TGFBR3-PLAG1 fusion in 6 (15%) cases. TGFBR3-PLAG1 was specific for MECA de novo tumors or the malignant component of MECA ex-PA, was absent in 723 other salivary gland tumors, and promoted a tumorigenic phenotype in vitro. We discovered other novel PLAG1 fusions, including ND4-PLAG1, which is an oncogenic fusion between mitochondrial and nuclear DNA. One tumor harbored an MSN-ALK fusion, which was tumorigenic in vitro, and targetable with ALK inhibitors. Certain gene fusions were predicted to result in neoantigens with high MHC binding affinity. A high number of copy number alterations was associated with poorer prognosis. Our findings indicate that MECA is a fusion-driven disease, nominate TGFBR3-PLAG1 as a hallmark of MECA, and provide a framework for future steps of diagnostic and therapeutic research in this lethal cancer.

cancer biology

Software For The Integration Of Multi-Omics Experiments In Bioconductor

Multi-omics experiments are increasingly commonplace in biomedical research, and add layers of complexity to experimental design, data integration, and analysis. R and Bioconductor provide a generic framework for statistical analysis and visualization, as well as specialized data classes for a variety of high-throughput data types, but methods are lacking for integrative analysis of multi-omics experiments. The MultiAssayExperiment software package, implemented in R and leveraging Bioconductor software and design principles, provides for the coordinated representation of, storage of, and operation on multiple diverse genomics data. We provide all of the multiple omics data for each cancer tissue in The Cancer Genome Atlas (TCGA) as ready-to-analyze MultiAssayExperiment objects, and demonstrate in these and other datasets how the software simplifies data representation, statistical analysis, and visualization. The MultiAssayExperiment Bioconductor package reduces major obstacles to efficient, scalable and reproducible statistical analysis of multi-omics data and enhances data science applications of multiple omics datasets.

bioinformatics