bioRxiv ScienceSearch

Biology subjects

Noushmehr, H.

Publications and source records attributed to Noushmehr, H..

9 recordsLinked to original sources

Analyses of cancer data in the Genomic Data Commons Data Portal with new functionalities in the TCGAbiolinks R/Bioconductor package

The advent of Next Generation Sequencing (NGS) technologies has opened new perspectives in deciphering the genetic mechanisms underlying complex diseases. Nowadays, the amount of genomic data is massive and substantial efforts and new tools are required to unveil the information hidden in the data.\n\nThe Genomic Data Commons (GDC) Data Portal is a large data collection platform that includes different genomic studies included the ones from The Cancer Genome Atlas (TCGA) and the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) initiatives, accounting for more than 40 tumor types originating from nearly 30000 patients. Such platforms, although very attractive, must make sure the stored data are easily accessible and adequately harmonized. Moreover, they have the primary focus on the data storage in a unique place, and they do not provide a comprehensive toolkit for analyses and interpretation of the data. To fulfill this urgent need, comprehensive but easily accessible computational methods for integrative analyses of genomic data without renouncing a robust statistical and theoretical framework are needed. In this context, the R/Bioconductor package TCGAbiolinks was developed, offering a variety of bioinformatics functionalities. Here we introduce new features and enhancements of TCGAbiolinks in terms of i) more accurate and flexible pipelines for differential expression analyses, ii) different methods for tumor purity estimation and filtering, iii) integration of normal samples from the Genotype-Tissue-Expression (GTEx) platform iv) support for other genomics datasets, here exemplified by the TARGET data.\n\nEvidence has shown that accounting for tumor purity is essential in the study of tumorigenesis, as these factors promote confounding behavior regarding differential expression analysis. Henceforth, we implemented these filtering procedures in TCGAbiolinks. Moreover, a limitation of some of the TCGA datasets is the unavailability or paucity of corresponding normal samples. We thus integrated into TCGAbiolinks the possibility to use normal samples from the Genotype-Tissue Expression (GTEx) project, which is another large-scale repository cataloging gene expression from healthy individuals. The new functionalities are available in the TCGABiolinks v 2.8 and higher released in Bioconductor version 3.7.

bioinformatics

Integrated Molecular Profiling Studies to Characterize the Cellular Origins of High-Grade Serous Ovarian Cancer

Historically, high-grade serous ovarian cancers (HGSOCs) were thought to arise from ovarian surface epithelial cells (OSECs) but recent data implicate fallopian tube secretory epithelial cells (FTSECs) as the major precursor. We performed transcriptomic and epigenomic profiling to characterize molecular similarities between OSECs, FTSECs and HGSOCs. Transcriptomic signatures of FTSECs were preserved in most HGSOCs reinforcing FTSECs as the predominant cell-of-origin; though an OSEC-like signature was associated with increased chemosensitivity (Padj = 0.03) and was enriched in proliferative-type tumors, suggesting a dualistic model for HGSOC origins. More super-enhancers (SEs) were shared between FTSECs and HGSOCs than between OSECS and HGSOCs (P < 2.2 x 10-16). SOX18, ELF3 and EHF transcription factors (TFs) coincided with HGSOC SEs and represent putative novel drivers of tumor development. Our integrative analyses support a predominantly fallopian origin for HGSOCs and indicate tumorigenesis may be driven by different TFs according to cell-of-origin.

cancer biology

Multi-Tissue Transcriptome-Wide Association Studies Identify 21 Novel Candidate Susceptibility Genes for High Grade Serous Epithelial Ovarian Cancer

Genome-wide association studies (GWASs) have identified about 30 different susceptibility loci associated with high grade serous ovarian cancer (HGSOC) risk. We sought to identify potential susceptibility genes by integrating the risk variants at these regions with genetic variants impacting gene expression and splicing of nearby genes. We compiled gene expression and genotyping data from 2,169 samples for 6 different HGSOC-relevant tissue types. We integrated these data with GWAS data from 13,037 HGSOC cases and 40,941 controls, and performed a transcriptome-wide association study (TWAS) across >70,000 significantly heritable gene/exon features. We identified 24 transcriptome-wide significant associations for 14 unique genes, plus 90 significant exon-level associations in 20 unique genes. We implicated multiple novel genes at risk loci, e.g. LRRC46 at 19q21.32 (TWAS P=1x10-9) and a PRC1 splicing event (TWAS P=9x10-8) which was splice-variant specific and exhibited no eQTL signal. Functional analyses in HGSOC cell lines found evidence of essentiality for GOSR2, INTS1, KANSL1 and PRC1; with the latter gene showing levels of essentiality comparable to that of MYC. Overall, gene expression and splicing events explained 41% of SNP-heritability for HGSOC (s.e. 11%, P=2.5x10-4), implicated at least one target gene for 6/13 distinct genome-wide significant regions and revealed 2 known and 26 novel candidate susceptibility genes for HGSOC.\n\nSTATEMENT OF SIGNIFICANCEFor many ovarian cancer risk regions, the target genes regulated by germline genetic variants are unknown. Using expression data from >2,100 individuals, this study identified novel associations of genes and splicing variants with ovarian cancer risk; with transcriptional variation now explaining over one-third of the SNP-heritability for this disease.

genomics

Moonlight: a tool for biological interpretation and driver genes discovery

Cancer is a complex and heterogeneous disease. It is crucial to identify the key driver genes and their role in cancer mechanisms with attention to different cancer stages, types or subtypes. Cancer driver genes are elusive and their discovery is complicated by the fact that the same gene can play a diverse role in different contexts. Key biological processes, such as cell proliferation and cell death, have been linked to cancer progression. Thus, in principle, they can be exploited to classify the cancer genes and unveil their role. Here, we present a new method, Moonlight, that exploit expression data to classify cancer genes. Moonlight relies on the integration of functional enrichment analysis, gene regulatory networks and upstream regulator analysis from expression data to score the importance of biological cancer-related processes taking into account either the inter- or intra-tumor heterogeneity. We then employed these scores to predict if each gene acts as a tumor suppressor gene (TSG) or as an oncogene (OCG). Our methodology also allow to predict genes with dual role, i.e. the moonlight genes (TSG in one cancer type or stage and OCG in another), as well as to elucidate the underlying biological processes. Availability: https://bioconductor.org/packages/MoonlightR & https://github.com/ibsquare/MoonlightR/

bioinformatics

Glioma CpG Island Methylator Phenotype (G-CIMP): Biological and Clinical Implications

Gliomas are a heterogeneous group of brain tumors with distinct biological and clinical properties. Despite advances in surgical techniques and clinical regimens, treatment of high-grade glioma remains challenging and carries dismal rates of therapeutic success and overall survival. Challenges include the molecular complexity of gliomas, as well as inconsistencies in histopathological grading, resulting in an inaccurate prediction of disease progression and failure in the use of standard therapy. The updated 2016 World Health Organization (WHO) classification of tumors of the central nervous system reflects a refinement of tumor diagnostics by integrating the genotypic and phenotypic features, thereby narrowing the defined subgroups. The new classification recommends the molecular diagnosis of the IDH mutational status in gliomas. IDH-mutant gliomas manifest the CpG Island Methylator Phenotype (G-CIMP). Notably, the recent identification of clinically relevant subsets of G-CIMP tumors (G-CIMP-high and G-CIMP-low) provide a further refinement in glioma classification that is independent of grade and histology. This scheme may be useful for predicting patient outcome and may be translated into effective therapeutic strategies tailored to each patient. In this review, we highlight the evolution of our understanding of the G-CIMP subsets and how recent advances in characterizing the genome and epigenome of gliomas may influence future basic and translational research.\n\nConflict of interestThe authors declare no conflict of interest

cancer biology

Distinct epigenetic shift in a subset of Glioma CpG island methylator phenotype (G-CIMP) during tumor recurrence

Histomorphology and current grading schemes are unable to predict glioma relapse and malignant tumor progression. We reported that the IDH-mutant associated Glioma-CpG Island Methylator Phenotype (G-CIMP) can be further divided into two clinically distinct subtypes independent of histopathological grading (G-CIMP-high and -low) with evidence of correlation with tumor progression. Here we performed a comprehensive epigenomic analysis of 74 longitudinally collected glioma samples (grade II-IV) to understand malignant recurrence from G-CIMP-high to G-CIMP-low. G-CIMP-low recurrence appeared in 12% of all gliomas and resemble IDH-wildtype primary glioblastoma. G-CIMP-low recurrence can be characterized by distinct epigenetic changes at candidate functional tissue enhancers with AP-1/SOX binding elements, stem cell-like epigenomic phenotype, and genomic instability. Finally, we defined a set of candidate biomarker signatures that predict recurrence of G-CIMP-low with clinically relevance on patient outcomes. Our study provides opportunity for refined clinical trial designs and therapeutic targets that limit progression to more aggressive G-CIMP-low phenotype.\n\nHIGHLIGHTSO_LIIndolent G-CIMP-high progresses to aggressive G-CIMP-low phenotype\nC_LIO_LIIncidence of G-CIMP-low recurrent tumors are 3 times greater than G-CIMP-low primary\nC_LIO_LIG-CIMP-low recurrent tumors share epigenomic features with IDH-wildtype primary GBM\nC_LIO_LIPredictive biomarkers of G-CIMP-low progression at primary diagnosis\nC_LI

cancer biology

Enhancer Linking by Methylation/Expression Relationships with the R package ELMER version 2

MotivationDNA methylation has been used to identify functional changes at transcriptional enhancers and other cis-regulatory modules (CRMs) in tumors and other disease tissues. Our R/Bioconductor package ELMER (Enhancer Linking by Methylation/Expression Relationships) provides a systematic approach that reconstructs altered gene regulatory networks (GRNs) by combining enhancer methylation and gene expression data derived from the same sample set.\n\nResultsWe present a completely revised version 2 of ELMER that provides numerous new features including an optional web-based interface and a new Supervised Analysis mode to use pre-defined sample groupings. We show that this approach can identify GRNs associated with many new Master Regulators including KLF5 in breast cancer.\n\nAvailabilityELMER v.2 is available as an R/Bioconductor package at http://bioconductor.org/packages/ELMER/

bioinformatics

TCGAbiolinksGUI: A graphical user interface to analyze cancer molecular and clinical data

BackgroundThe GDC (Genomic Data Commons) data portal provides users with data from cancer genomics studies. Recently, we developed the R/Bioconductor TCGAbiolinks package, which allows users to search, download and prepare cancer genomics data for integrative data analysis. The use of this package requires users to have advanced knowledge of R thus limiting the number of users.\n\nResultsTo overcome this obstacle and improve the accessibility of the package by a wider range of users, we developed TCGAbiolinksGUI that uses shiny graphical user interface (GUI) available through the R/Bioconductor package.\n\nConclusionThe TCGAbiolinksGUI package is freely available within the Bioconductor project at http://bioconductor.org/packages/TCGAbiolinksGUI/. Links to the GitHub repository, a demo version of the tool, a docker image and PDF/video tutorials are available at http://bit.do/TCGAbiolinksDocs.

bioinformatics

RGBM: Regularized Gradient Boosting Machines For The Identification of Transcriptional Regulators Of Discrete Glioma Subtypes

The transcription factors (TF) which regulate gene expressions are key determinants of cellular phenotypes. Reconstructing large-scale genome-wide networks which capture the influence of TFs on target genes are essential for understanding and accurate modelling of living cells. We propose RGBM: a gene regulatory network (GRN) inference algorithm, which can handle data from heterogeneous information sources including dynamic time-series, gene knockout, gene knockdown, DNA microarrays and RNA-Seq expression profiles. RGBM allows to use an a priori mechanistic of active biding network consisting of TFs and corresponding target genes. RGBM is evaluated on the DREAM challenge datasets where it surpasses the winners of the competitions and other established methods for two evaluation metrics by about 10-15%.\n\nWe use RGBM to identify the main regulators of the molecular subtypes of brain tumors. Our analysis reveals the identity and corresponding biological activities of the master regulators driving transformation of the G-CIMP-high into the G-CIMP-low subtype of glioma and PA-like into LGm6-GBM, thus, providing a clue to the yet undetermined nature of the transcriptional events driving the evolution among these novel glioma subtypes.\n\nRGBM is available for download on CRAN at https://cran.rproject.org/web/packages/RGBM/index.html

bioinformatics