bioRxiv ScienceSearch

Biology subjects

Su-In Lee

Publications and source records attributed to Su-In Lee.

3 recordsLinked to original sources

Extracting a low-dimensional description of multiple gene expression datasets reveals a potential driver for tumor-associated stroma in ovarian cancer

BackgroundDiscovering patient subtypes and molecular drivers of a subtype are difficult and driving problems underlying most modern disease expression studies collected across patient populations. Expression patterns conserved across multiple expression datasets from independent disease studies are likely to represent important molecular events underlying the disease.\n\nMethodsWe present the INSPIRE (INferring Shared modules from multiPle gene expREssion datasets) method to infer highly coherent and robust modules of co-expressed genes and the dependencies among the modules from multiple expression datasets. Focusing on inferring modules and their dependencies conserved across multiple expression datasets is important for several reasons. First, using multiple datasets will increase the power to detect robust and relevant patterns (modules and dependencies among modules). Second, INSPIRE enables the use of multiple datasets that contain different sets of genes due to, e.g., the difference in microarray platforms. Many methods designed for expression data analysis cannot integrate multiple datasets with variable discrepancy to infer a single combined model, whereas INSPIRE can naturally model the dependencies among the modules even when a large proportion of genes are not observed on a certain platform.\n\nResultsWe evaluated INSPIRE on synthetically generated datasets with known underlying network structure among modules, and gene expression datasets from multiple ovarian cancer studies. We show that the model learned by INSPIRE can explain unseen data better and can reveal prior knowledge on gene functions more accurately than alternative methods. We demonstrate that applying INSPIRE to nine ovarian cancer datasets leads to the identification of a new marker and potential molecular driver of tumor-associated stroma - HOPX. We also demonstrate that the HOPXmodule strongly overlaps with the genes defining the mesenchymal patient subtype identified in The Cancer Genome Atlas (TCGA) ovarian cancer data. We provide evidence for a previously unknown molecular basis of tumor resectability efficacy involving tumor-associated mesenchymal stem cells represented by HOPX.\n\nConclusionsINSPIRE extracts a low-dimensional description from multiple gene expression data, which consists of modules and their dependencies. The discovery of a new tumor-associated stroma marker, HOPX, and its module suggests a previously unknown mechanism underlying tumor-associated stroma.

Systems Biology

Identifying Network Perturbation in Cancer

We present a computational framework, called DISCERN (DIfferential SparsE Regulatory Network), to identify informative topological changes in gene-regulator dependence networks inferred on the basis of mRNA expression datasets within distinct biological states. DISCERN takes two expression datasets as input: an expression dataset of diseased tissues from patients with a disease of interest and another expression dataset from matching normal tissues. DISCERN estimates the extent to which each gene is perturbed - having distinct regulator connectivity in the inferred gene-regulatory dependencies between the disease and normal conditions. This approach has distinct advantages over existing methods. First, DISCERN infers conditional dependencies between candidate regulators and genes, where conditional dependence relationships discriminate the evidence for direct interactions from indirect interactions more precisely than pairwise correlation. Second, DISCERN uses a new likelihood-based scoring function to alleviate concerns about accuracy of the specific edges inferred in a particular network. DISCERN identifies perturbed genes more accurately in synthetic data than existing methods to identify perturbed genes between distinct states. In expression datasets from patients with acute myeloid leukemia (AML), breast cancer and lung cancer, genes with high DISCERN scores in each cancer are enriched for known tumor drivers, genes associated with the biological processes known to be important in the disease, and genes associated with patient prognosis, in the respective cancer. Finally, we show that DISCERN can uncover potential mechanisms underlying network perturbation by explaining observed epigenomic activity patterns in cancer and normal tissue types more accurately than alternative methods, based on the available epigenomic from the ENCODE project.

Genomics

Learning the human chromatin network from all ENCODE ChIP-seq data

Introduction: A cells epigenome arises from interactions among regulatory factors -- transcription factors, histone modifications, and other DNA-associated proteins -- co-localized at particular genomic regions. Identifying the network of interactions among regulatory factors, the chromatin network, is of paramount importance in understanding epigenome regulation.\n\nMethods: We developed a novel computational approach, ChromNet, to infer the chromatin network from a set of ChIP-seq datasets. ChromNet has four key features that enable its use on large collections of ChIP-seq data. First, rather than using pairwise co-localization of factors along the genome, ChromNet identifies conditional dependence relationships that better discriminate direct and indirect interactions. Second, our novel statistical technique, the group graphical model, improves inference of conditional dependence on highly correlated datasets. Such datasets are common because some transcription factors form a complex and the same transcription factor is often assayed in different laboratories or cell types. Third, ChromNets computationally efficient method and the group graphical model enable the learning of a joint network across all cell types, which greatly increases the scope of possible interactions. We have shown that this results in a significantly higher fold enrichment for validated protein interactions. Fourth, ChromNet provides an efficient way to identify the genomic context that drives a particular network edge, which provides a more comprehensive understanding of regulatory factor interactions.\n\nResults: We applied ChromNet to all available ChIP-seq data from the ENCODE Project, consisting of 1451 ChIP-seq datasets, which revealed previously known physical interactions better than alternative approaches. ChromNet also identified previously unreported regulatory factor interactions. We experimentally validated one of these interactions, between the MYC and HCFC1 transcription factors.\n\nDiscussion: ChromNet provides a useful tool for understanding the interactions among regulatory factors and identifying novel interactions. We have provided an interactive web-based visualization of the full ENCODE chromatin network and the ability to incorporate custom datasets at http://chromnet.cs.washington.edu.

Genomics