bioRxiv ScienceSearch

Biology subjects

Barillot, E.

Publications and source records attributed to Barillot, E..

11 recordsLinked to original sources

Stabilized Independent Component Analysis outperforms other methods in finding reproducible signals in tumoral transcriptomes

MotivationMatrix factorization methods are widely exploited in order to reduce dimensionality of transcriptomic datasets to the action of few hidden factors (metagenes). Applying such methods to similar independent datasets should yield reproducible inter-series outputs, though it was never demonstrated yet.\n\nResultsWe systematically test state-of-art methods of matrix factorization on several transcriptomic datasets of the same cancer type. Inspired by concepts of evolutionary bioinformatics, we design a new framework based on Reciprocally Best Hit (RBH) graphs in order to benchmark the methods reproducibility. We show that a particular protocol of application of Independent Component Analysis (ICA), accompanied by a stabilisation procedure, leads to a significant increase in the inter-series output reproducibility. Moreover, we show that the signals detected through this method are systematically more interpretable than those of other state-of-art methods. We developed a user-friendly tool BIODICA for performing the Stabilized ICA-based RBH meta-analysis. We apply this methodology to the study of colorectal cancer (CRC) for which 14 independent publicly available transcriptomic datasets can be collected. The resulting RBH graph maps the landscape of interconnected factors that can be associated to biological processes or to technological artefacts. These factors can be used as clinical biomarkers or robust and tumor-type specific transcriptomic signatures of tumoral cells or tumoral microenvironment. Their intensities in different samples shed light on the mechanistic basis of CRC molecular subtyping.\n\nAvailabilityThe BIODICA tool is available from https://github.com/LabBandSB/BIODICA.\n\nContactlaura.cantini@curie.fr and andrei.zinovyev@curie.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Identification of microRNA clusters cooperatively acting on Epithelial to Mesenchymal Transition in Triple Negative Breast Cancer

MicroRNAs play important roles in many biological processes. Their aberrant expression can have oncogenic or tumor suppressor function directly participating to carcinogenesis, malignant transformation, invasiveness and metastasis. Indeed, miRNA profiles can distinguish not only between normal and cancerous tissue but they can also successfully classify different subtypes of a particular cancer. Here, we focus on a particular class of transcripts encoding polycistronic miRNA genes that yields multiple miRNA components. We describe clustered MiRNA Master Regulator Analysis (ClustMMRA), a fully redesigned release of the MMRA computational pipeline (MiRNA Master Regulator Analysis), developed to search for clustered miRNAs potentially driving cancer molecular subtyping. Genomically clustered miRNAs are frequently co-expressed to target different components of pro-tumorigenic signalling pathways. By applying ClustMMRA to breast cancer patient data, we identified key miRNA clusters driving the phenotype of different tumor subgroups. The pipeline was applied to two independent breast cancer datasets, providing statistically concordant results between the two analysis. We validated in cell lines the miR-199/miR-214 as a novel cluster of miRNAs promoting the triple negative subtype phenotype through its control of proliferation and EMT.

systems biology

Metabolic and signalling network map integration: application to cross-talk studies and omics data analysis in cancer

BackgroundThe interplay between metabolic processes and signalling pathways remains poorly understood. Global, detailed and comprehensive reconstructions of human metabolism and signalling pathways exist in the form of molecular maps, but they have never been integrated together. We aim at filling in this gap by creating an integrated resource of both signalling and metabolic pathways allowing a visual exploration of multi-level omics data and study of cross-regulatory circuits between these processes in health and in disease.\n\nResultsWe combined two comprehensive manually curated network maps. Atlas of Cancer Signalling Network (ACSN), containing mechanisms frequently implicated in cancer; and ReconMap 2.0, a comprehensive reconstruction of human metabolic network. We linked ACSN and ReconMap 2.0 maps via common players and represented the two maps as interconnected layers using the NaviCell platform for maps exploration. In addition, proteins catalysing metabolic reactions in ReconMap 2.0 were not previously visually represented on the map canvas. This precluded visualisation of omics data in the context of ReconMap 2.0. We suggested a solution for displaying protein nodes on the ReconMap 2.0 map in the vicinity of the corresponding reaction or process nodes. This permits multi-omics data visualisation in the context of both map layers. Exploration and shuttling between the two map layers is possible using Google Maps-like features of NaviCell. The integrated ACSN-ReconMap 2.0 resource is accessible online and allows data visualisation through various modes such as markers, heat maps, bar-plots, glyphs and map staining. The integrated resource was applied for comparison of immunoreactive and proliferative ovarian cancer subtypes using transcriptomic, copy number and mutation multi-omics data. A certain number of metabolic and signalling processes specifically deregulated in each of the ovarian cancer sub-types were identified.\n\nConclusionsAs knowledge evolves and new omics data becomes more heterogeneous, gathering together existing domains of biology under common platforms is essential. We believe that an integrated ACSN-ReconMap 2.0 resource will help in understanding various disease mechanisms and discovery of new interactions at the intersection of cell signalling and metabolism. In addition, the successful integration of metabolic and signalling networks allows broader systems biology approach application for data interpretation and retrieval of intervention points to tackle simultaneously the key players coordinating signalling and metabolism in human diseases.

systems biology

PhysiBoSS: a multi-scale agent based modelling framework integrating physical dimension and cell signalling

Due to the complexity of biological systems, their heterogeneity, and the internal regulation of each cell and its surrounding, mathematical models that take into account cell signalling, cell population behaviour and the extracellular environment are particularly helpful to understand such complex systems. However, very few of these tools, freely available and computationally efficient, are currently available. To fill this gap, we present here our open-source software, PhysiBoSS, which is built on two available software packages that focus on different scales: intracellular signalling using continuous-time markovian Boolean modelling (MaBoSS) and multicellular behaviour using agent-based modelling (PhysiCell).\n\nThe multi-scale feature of PhysiBoSS - its agent-based structure and the possibility to integrate any Boolean network to it - provide a flexible and computationally efficient framework to study heterogeneous cell population growth in diverse experimental set-ups. This tool allows one to explore the effect of environmental and genetic alterations of individual cells at the population level, bridging the critical gap from genotype to phenotype. PhysiBoSS thus becomes very useful when studying population response to treatment, mutations effects, cell modes of invasion or isomorphic morphogenesis events.\n\nTo illustrate potential use of PhysiBoSS, we studied heterogeneous cell fate decisions in response to TNF treatment in a 2-D cell population and in a tumour cell 3-D spheroid. We explored the effect of different treatment regimes and the behaviour and selection of several resistant mutants. We highlighted the importance of spatial information on the population dynamics by considering the effect of competition for resources like oxygen. PhysiBoSS is freely available on GitHub (https://github.com/gletort/PhysiBoSS), and is distributed open source under the BSD 3-clause license. It is compatible with most Unix systems, and a Docker package (https://hub.docker.com/r/gletort/physiboss/) is provided to ease its deployment in other systems.

systems biology

Moonlight: a tool for biological interpretation and driver genes discovery

Cancer is a complex and heterogeneous disease. It is crucial to identify the key driver genes and their role in cancer mechanisms with attention to different cancer stages, types or subtypes. Cancer driver genes are elusive and their discovery is complicated by the fact that the same gene can play a diverse role in different contexts. Key biological processes, such as cell proliferation and cell death, have been linked to cancer progression. Thus, in principle, they can be exploited to classify the cancer genes and unveil their role. Here, we present a new method, Moonlight, that exploit expression data to classify cancer genes. Moonlight relies on the integration of functional enrichment analysis, gene regulatory networks and upstream regulator analysis from expression data to score the importance of biological cancer-related processes taking into account either the inter- or intra-tumor heterogeneity. We then employed these scores to predict if each gene acts as a tumor suppressor gene (TSG) or as an oncogene (OCG). Our methodology also allow to predict genes with dual role, i.e. the moonlight genes (TSG in one cancer type or stage and OCG in another), as well as to elucidate the underlying biological processes. Availability: https://bioconductor.org/packages/MoonlightR & https://github.com/ibsquare/MoonlightR/

bioinformatics

Application of Atlas of Cancer Signalling Network in pre-clinical studies

Initiation and progression of cancer involve multiple molecular mechanisms. The knowledge on these mechanisms is expanding and should be converted into guidelines for tackling the disease. We discuss here formalization of biological knowledge into a comprehensive resource Atlas of Cancer Signalling Network (ACSN) and Google Maps-based tool NaviCell that supports map navigation. The application of maps for omics data visualisation in the context of signalling maps is possible using NaviCell Web Service module and NaviCom tool for generation of network-based molecular portraits of cancer using multi-level omics data. We review how these resources and tools are applied for cancer pre-clinical studies among others for rationalizing synergistic effect of drugs and designing complex disease stage-specific druggable interventions following structural analysis of the maps together with omics data. Modules and maps of ACSN as signatures of biological functions, can help in cancer data analysis and interpretation. In addition, they can also be used to find association between perturbations in particular molecular mechanisms to the risk of a specific cancer type development. These approaches and beyond help to study interplay between molecular mechanisms of cancer, deciphering how gene interactions govern hallmarks of cancer in specific context. We discuss a perspective to develop a flexible methodology and a pipeline to enable systematic omics data analysis in the context of signalling network maps, for stratifying patients and suggesting interventions points and drug repositioning in cancer and other human diseases.

systems biology

Determining the optimal number of independent components for reproducible transcriptomic data analysis

BackgroundIndependent Component Analysis (ICA) is a method that models gene expression data as an action of a set of statistically independent hidden factors. The output of ICA depends on a fundamental parameter: the number of components (factors) to compute. The optimal choice of this parameter, related to determining the effective data dimension, remains an open question in the application of blind source separation techniques to transcriptomic data.\n\nResultsHere we address the question of optimizing the number of statistically independent components in the analysis of transcriptomic data for reproducibility of the components in multiple runs of ICA (within the same or within varying effective dimensions) and in multiple independent datasets. To this end, we introduce ranking of independent components based on their stability in multiple ICA computation runs and define a distinguished number of components (Most Stable Transcriptome Dimension, MSTD) corresponding to the point of the qualitative change of the stability profile. Based on a large body of data, we demonstrate that a sufficient number of dimensions is required for biological interpretability of the ICA decomposition and that the most stable components with ranks below MSTD have more chances to be reproduced in independent studies compared to the less stable ones. At the same time, we show that a transcriptomics dataset can be reduced to a relatively high number of dimensions without losing the interpretability of ICA, even though higher dimensions give rise to components driven by small gene sets.\n\nConclusionsWe suggest a protocol of ICA application to transcriptomics data with a possibility of prioritizing components with respect to their reproducibility that strengthens the biological interpretation. Computing too few components (much less than MSTD) is not optimal for interpretability of the results. The components ranked within MSTD range have more chances to be reproduced in independent studies.

bioinformatics

Differential epigenetic landscapes and transcription factors explain X-linked gene behaviours during X-chromosome reactivation in the mouse inner cell mass

X-chromosome inactivation (XCI) is established in two waves during mouse development. First, silencing of the paternal X chromosome (Xp) is triggered, with transcriptional repression of most genes and enrichment of epigenetic marks such as H3K27me3 being achieved in all cells by the early blastocyst stage. XCI is then reversed in the inner cell mass (ICM), followed by a second wave of maternal or paternal XCI, in the embryo-proper. Although the role of Xist RNA in triggering XCI is now clear, the mechanisms underlying Xp reactivation in the inner cell mass have remained enigmatic. Here we use in vivo single cell approaches (allele-specific RNAseq, nascent RNA FISH and immunofluorescence) and find that different genes show very different timing of reactivation. We observe that the genes reactivate at different stages and that initial enrichment in H3K27me3 anti-correlates with the speed of reactivation. To define whether this repressive histone mark is lost actively or passively, we investigate embryos mutant for the X-encoded H3K27me3 demethylase, UTX. Xp genes that normally reactivate slowly are retarded in their reactivation in Utx mutants, while those that reactive rapidly are unaffected. Therefore, efficient reprogramming of some X-linked genes in the inner cell mass is very rapid, indicating minimal epigenetic memory and potentially driven by transcription factors, whereas others may require active erasure of chromatin marks such as H3K27me3.

developmental biology

Classification Of Gene Signatures For Their Information Value And Functional Redundancy

Large collections of gene signatures play a pivotal role in interpreting results of omics data analysis but suffer from compositional (large overlap) and functional (redundant read-outs) redundancy, and many gene signatures rarely pop-up in statistical tests. Based on pan-cancer data analysis, here we define a restricted set of 962 so called informative signatures and demonstrate that they have more chances to appear highly enriched in cancer biology studies. We show that the majority of informative signatures conserve their weights for the composing genes (eigengenes) from one cancer type to another. We construct InfoSigMap, an interactive online map showing the structure of compositional and functional redundancies between informative signatures and charting the territories of biological functions accessible through transcriptomic studies. InfoSigMap can be used to visualize in one insightful picture the results of comparative omics data analyses and suggests reconsidering existing annotations of certain reference gene set groups.

systems biology

Knowledge Formalization and High-Throughput Data Visualization Using Signaling Network Maps

Generation and usage of high-quality molecular signalling network maps can be augmented by standardising notations, establishing curation workflows and application of computational biology methods to exploit the knowledge contained in the maps. In this manuscript, we summarize the major aims and challenges of assembling information in the form of comprehensive maps of molecular interactions. Mainly, we share our experience gained while creating the Atlas of Cancer Signalling Network. In the step-by-step procedure, we describe the map construction process and suggest solutions for map complexity management by introducing a hierarchical modular map structure. In addition, we describe the NaviCell platform, a computational technology using Google Maps API to explore comprehensive molecular maps similar to geographical maps, and explain the advantages of semantic zooming principles for map navigation. We also provide the outline to prepare signalling network maps for navigation using the NaviCell platform. Finally, several examples of cancer high-throughput data analysis and visualization in the context of comprehensive signalling maps are presented.

systems biology

NaviCom: A web application to create interactive molecular network portraits using multi-level omics data

Human diseases such as cancer are routinely characterized by high-throughput molecular technologies, and multi-level omics data are accumulated in public databases at increasing rate. Retrieval and visualization of these data in the context of molecular network maps can provide insights into the pattern of molecular functions encompassed by an omics profile. In order to make this task easy, we developed NaviCom, a Python package and web platform for visualization of multi-level omics data on top of biological network maps. NaviCom is bridging the gap between cBioPortal, the most used resource of large-scale cancer omics data and NaviCell, a data visualization web service that contains several molecular network map collections. NaviCom proposes several standardized modes of data display on top of molecular network maps, allowing to address specific biological questions. We illustrate how users can easily create interactive network-based cancer molecular portraits via NaviCom web interface using the maps of Atlas of Cancer Signaling Network (ACSN) and other maps. Analysis of these molecular portraits can help in formulating a scientific hypothesis on the molecular mechanisms deregulated in the studied disease.

systems biology