bioRxiv ScienceSearch

Biology subjects

Subramanian, A.

Publications and source records attributed to Subramanian, A..

8 recordsLinked to original sources

The Carcinogenome Project: In-vitro Gene Expression Profiling of Chemical Perturbations to Predict Long-Term Carcinogenicity

Background: Most chemicals in commerce have not been evaluated for their carcinogenic potential. The current de-facto gold-standard approach to carcinogen testing adopts the two-year rodent bioassay, a time consuming and costly procedure. Alternative approaches, such as high-throughput in-vitro assays, show promise in addressing the limitations in carcinogen screening.\n\nObjectives: We developed a screening process for predicting chemical carcinogenicity and genotoxicity and characterizing modes of actions (MoAs) using in-vitro gene expression assays.\n\nMethods: We generated a large toxicogenomics resource comprising ~6,000 expression profiles corresponding to 330 chemicals profiled in HepG2 cells at multiple doses and in replicates. Predictive models of carcinogenicity were built using a Random Forest classifier. Differential pathway enrichment analysis was performed to identify pathways associated with carcinogen exposure. Signatures of carcinogenicity and genotoxicity were compared with external data sources including Drugmatrix and the Connectivity Map.\n\nResults: Among profiles with sufficient bioactivity, our classifiers achieved 72.2% AUC for predicting carcinogenicity and 82.3% AUC for predicting genotoxicity. Our analysis showed that chemical bioactivity, as measured by the strength and reproducibility of the transcriptional response, is not significantly associated with long-term carcinogenicity, as evidenced by the many carcinogenic chemicals that did not elicit substantial changes in gene expression at doses up to 40 M. However, sufficiently high transcriptional bioactivity is necessary for a chemical to be used for prediction of carcinogenicity. Pathway enrichment analysis revealed several pathways consistent with literature review of pathways that drive cancer, including DNA damage and DNA repair. These data are available for download via https://clue.io/CRCGN_ABC, and a web portal for interactive query and visualization of the data and results is accessible at https://carcinogenome.org.\n\nConclusions: We demonstrated a short-term in-vitro screening approach using gene expression profiling to predict long-term carcinogenicity and infer MoAs of chemical perturbations.

bioinformatics

Open Community Challenge Reveals Molecular Network Modules with Key Roles in Diseases

Identification of modules in molecular networks is at the core of many current analysis methods in biomedical research. However, how well different approaches identify disease-relevant modules in different types of gene and protein networks remains poorly understood. We launched the "Disease Module Identification DREAM Challenge", an open competition to comprehensively assess module identification methods across diverse protein-protein interaction, signaling, gene co-expression, homology, and cancer-gene networks. Predicted network modules were tested for association with complex traits and diseases using a unique collection of 180 genome-wide association studies (GWAS). Our critical assessment of 75 contributed module identification methods reveals novel top-performing algorithms, which recover complementary trait-associated modules. We find that most of these modules correspond to core disease-relevant pathways, which often comprise therapeutic targets and correctly prioritize candidate disease genes. This community challenge establishes benchmarks, tools and guidelines for molecular network analysis to study human disease biology (https://synapse.org/modulechallenge).

bioinformatics

The GCTx format and cmap{Py, R, M} packages: resources for the optimized storage and integrated traversal of dense matrices of data and annotations

MotivationComputational analysis of datasets generated by treating cells with pharmacological and genetic perturbagens has proven useful for the discovery of functional relationships. Facilitated by technological improvements, perturbational datasets have grown in recent years to include millions of experiments. While initial studies, such as our work on Connectivity Map, used gene expression readouts, recent studies from the NIH LINCS consortium have expanded to a more diverse set of molecular readouts, including proteomic and cell morphological signatures. Sharing these diverse data creates many opportunities for research and discovery, but the unprecedented size of data generated and the complex metadata associated with experiments have also created fundamental technical challenges regarding data storage and cross-assay integration.\n\nResultsWe present the GCTx file format and a suite of open-source packages for the efficient storage, serialization, and analysis of dense two-dimensional matrices. The utility of this format is not just theoretical; we have extensively used the format in the Connectivity Map to assemble and share massive data sets comprising 1.7 million experiments. We anticipate that the generalizability of the GCTx format, paired with code libraries that we provide, will stimulate wider adoption and lower barriers for integrated cross-assay analysis and algorithm development.\n\nAvailabilitySoftware packages (available in Matlab, Python, and R) are freely available at https://github.com/cmap\n\nSupplementary informationSupplementary information is available at clue.io/code.\n\nContactoana@broadinstitute.org

bioinformatics

A unified web platform for network-based analyses of genomic data

Functional genomics networks are widely used to identify unexpected pathway relationships in large genomic datasets. However, it is challenging to quantitatively compare the signal-to-noise ratio of different networks, the biology they describe, and to identify the optimal network to interpret a particular genetic dataset. Via GeNets users can train a machine-learning model (Quack) to make such comparisons; and they can execute, store, and share analyses of genetic and RNA sequencing datasets.

genomics

Cell-type heterogeneity in the zebrafish olfactory placode is generated from progenitors within preplacodal ectoderm

Vertebrate olfactory placodes consists of a variety of neuronal populations, which are thought to have distinct embryonic origins. In the zebrafish, while ciliated sensory neurons arise from preplacodal ectoderm (PPE), previous lineage tracing studies suggest that both Gonadotropin releasing hormone 3 (Gnrh3) and microvillous sensory neurons derive from cranial neural crest (CNC). We find that the expression of Islet1/2 is restricted to Gnrh3 neurons associated with the olfactory placode. Unexpectedly, however, we find no change in Islet1/2+ cell numbers in sox10 mutant embryos, calling into question their CNC origin. Lineage reconstruction based on backtracking in time-lapse confocal datasets, and confirmed by photoconversion experiments, reveals that Gnrh3 neurons derive from the anterior/medial PPE. Similarly, all of the microvillous sensory neurons we have traced arise from preplacodal progenitors. Our results suggest that rather than originating from separate ectodermal populations, cell-type heterogeneity is generated from overlapping pools of progenitors within the preplacodal ectoderm.

developmental biology

A Library of Phosphoproteomic and Chromatin Signatures for Characterizing Cellular Responses to Drug Perturbations

Though the added value of proteomic measurements to gene expression profiling has been demonstrated, profiling of gene expression on its own remains the dominant means of understanding cellular responses to perturbation. Direct protein measurements are typically limited due to issues of cost and scale; however, the recent development of high-throughput, targeted sentinel mass spectrometry assays provides an opportunity for proteomics to contribute at a meaningful scale in high-value areas for drug development. To demonstrate the feasibility of a systematic and comprehensive library of perturbational proteomic signatures, we profiled 90 drugs (in triplicate) in six cell lines using two different proteomic assays -- one measuring global changes of epigenetic marks on histone proteins and another measuring a set of peptides reporting on the phosphoproteome -- for a total of more than 3,400 samples. This effort represents a first-of-its-kind resource for proteomics. The majority of tested drugs generated reproducible responses in both phosphosignaling and chromatin states, but we observed differences in the responses that were cell line-and assay-specific. We formalized the process of comparing response signatures within the data using a concept called connectivity, which enabled us to integrate data across cell types and assays. Furthermore, it facilitated incorporation of transcriptional signatures. Consistent connectivity among cell types revealed cellular responses that transcended cell-specific effects, while consistent connectivity among assays revealed unexpected associations between drugs that were confirmed by experimental follow-up. We further demonstrated how the resource could be leveraged against public domain external datasets to recognize therapeutic hypotheses that are consistent with ongoing clinical trials for the treatment of multiple myeloma and acute lymphocytic leukemia (ALL). These data are available for download via the Gene Expression Omnibus (accession GSE101406), and web apps for interacting with this resource are available at https://clue.io/proteomics.\n\nHighlightsO_LIFirst-of-its-kind public resource of proteomic responses to systematically administered perturbagens\nC_LIO_LIDirect proteomic profiling of phosphosignaling and chromatin states in cells for 90 drugs in six different cell lines\nC_LIO_LIExtends Connectivity Map concept to proteomic data for integration with transcriptional data\nC_LIO_LIEnables recognition of unexpected, cell type-specific activities and potential translational therapeutic opportunities\nC_LI

systems biology

Evaluation of RNAi and CRISPR technologies by large-scale gene expression profiling in the Connectivity Map

The application of RNA interference (RNAi) to mammalian cells has provided the means to perform phenotypic screens to determine the functions of genes. Although RNAi has revolutionized loss of function genetic experiments, it has been difficult to systematically assess the prevalence and consequences of off-target effects. The Connectivity Map (CMAP) represents an unprecedented resource to study the gene expression consequences of expressing short hairpin RNAs (shRNAs). Analysis of signatures for over 13,000 shRNAs applied in 9 cell lines revealed that miRNA-like off-target effects of RNAi are far stronger and more pervasive than generally appreciated. We show that mitigating off-target effects is feasible in these datasets via computational methodologies to produce a Consensus Gene Signature (CGS). In addition, we compared RNAi technology to clustered regularly interspaced short palindromic repeat (CRISPR)-based knockout by analysis of 373 sgRNAs in 6 cells lines, and show that the on-target efficacies are comparable, but CRISPR technology is far less susceptible to systematic off-target effects. These results will help guide the proper use and analysis of loss-of-function reagents for the determination of gene function.

bioinformatics

A Next Generation Connectivity Map: L1000 Platform And The First 1,000,000 Profiles

We previously piloted the concept of a Connectivity Map (CMap), whereby genes, drugs and disease states are connected by virtue of common gene-expression signatures. Here, we report more than a 1,000-fold scale-up of the CMap as part of the NIH LINCS Consortium, made possible by a new, low-cost, high throughput reduced representation expression profiling method that we term L1000. We show that L1000 is highly reproducible, comparable to RNA sequencing, and suitable for computational inference of the expression levels of 81% of non-measured transcripts. We further show that the expanded CMap can be used to discover mechanism of action of small molecules, functionally annotate genetic variants of disease genes, and inform clinical trials. The 1.3 million L1000 profiles described here, as well as tools for their analysis, are available at https://clue.io.\n\nHIGHLIGHTSO_LIA new gene expression profiling method, L1000, dramatically lowers cost\nC_LIO_LIThe Connectivity Map database now includes 1.3 million publicly accessible L1000 perturbational profiles\nC_LIO_LIThis expanded Connectivity Map facilitates discovery of small molecule mechanism of action and functional annotation of genetic variants\nC_LIO_LIThe work establishes feasibility and utility of a truly comprehensive Connectivity Map\nC_LI

genomics