bioRxiv ScienceSearch

Biology subjects

Chris Sander

Publications and source records attributed to Chris Sander.

14 recordsLinked to original sources

3D RNA from evolutionary couplings

Non-protein-coding RNAs are ubiquitous in cell physiology, with a diverse repertoire of known functions. In fact, the majority of the eukaryotic genome does not code for proteins, and thousands of conserved long non-protein-coding RNAs of currently unkown function have been identified. When available, knowledge of their 3D structure is very helpful in elucidating the function of these RNAs. However, despite some outstanding structure elucidation of RNAs using X-ray crystallography, NMR and cryoEM, learning RNA 3D structures remains low-throughput. RNA structure prediction in silico is a promising alternative approach and works well for double-helical stems, but full 3D structure determination requires tertiary contacts outside of secondary structures that are difficult to infer from sequence information. Here, based only on information from RNA multiple sequence alignments, we use a global statistical sequence probability model of co-variation in a pairs of nucleotide positions to detect 3D contacts, in analogy to recently developed breakthrough methods for computational protein folding. In blinded tests on 22 known RNA structures ranging in size from 65 to 1800 nucleotides, the predicted contacts matched physical nucleotide interactions with 65-95% true positive prediction accuracy. Importantly, we infer many long-range tertiary contacts, including non-Watson-Crick interactions, where secondary structure elements assemble in 3D. When used as restraints in molecular dynamics simulations, the inferred contacts improve RNA 3D structure prediction to a coordinate error as low as 6 - 10 [A] rmsd deviation in atom positions, with potential for further refinement by molecular dynamics. These contacts include functionally important interactions, such as those that distinguish the active and inactive conformations of four riboswitches. In blind prediction mode, we present evolutionary couplings suitable for folding simulations for 180 RNAs of unknown structure, available at https://marks.hms.harvard.edu/ev_rna/. We anticipate that this approach can help shed light on the structure and function of non-protein-coding RNAs as well as 3D-structured mRNAs.

Bioinformatics

Comparing cancer cell lines and tumor samples by genomic profiles

Cancer cell lines are often used in laboratory experiments as models of tumors, although they can have substantially different genetic and epigenetic profiles compared to tumors. We have developed a general computational method - TumorComparer - to systematically quantify similarities and differences between tumor material when detailed genetic and molecular profiles are available. The comparisons can be flexibly tailored to a particular biological question by placing a higher weight on functional alterations of interest ( weighted similarity). In a first pan-cancer application, we have compared 260 cell lines from the Cancer Cell Line Encyclopaedia (CCLE) and 1914 tumors of six different cancer types from The Cancer Genome Atlas (TCGA), using weights to emphasize genomic alterations that frequently recur in tumors. We report the potential suitability of particular cell lines as tumor models and identify apparently unsuitable outlier cell lines, some of which are in wide use, for each of the six cancer types. In future, this weighted similarity method may be generalized for use in a clinical setting to compare patient profiles consisting of genomic patterns combined with clinical attributes, such as diagnosis, treatment and response to therapy.

Cancer Biology

The landscape of T cell infiltration in human cancer and its association with antigen presenting gene expression

One sentence summaryIn silico decomposition of the immune microenvironment among common tumor types identified clear cell renal cell carcinoma as the most highly infiltrated by T-cells and further analysis of this tumor type revealed three distinct and clinically relevant clusters which were validated in an independent cohort.\n\nAbstractInfiltrating T cells in the tumor microenvironment have crucial roles in the competing processes of pro-tumor and anti-tumor immune response. However, the infiltration level of distinct T cell subsets and the signals that draw them into a tumor, such as the expression of antigen presenting machinery (APM) genes, remain poorly characterized across human cancers. Here, we define a novel mRNA-based T cell infiltration score (TIS) and profile infiltration levels in 19 tumor types. We find that clear cell renal cell carcinoma (ccRCC) is the highest for TIS and among the highest for the correlation between TIS and APM expression, despite a modest mutation burden. This finding is contrary to the expectation that immune infiltration and mutation burden are linked. To further characterize the immune infiltration in ccRCC, we use RNA-seq data to computationally infer the infiltration levels of 24 immune cell types in a discovery cohort of 415 ccRCC patients and validate our findings in an independent cohort of 101 ccRCC patients. We find three clusters of tumors that are primarily separated by levels of T cell infiltration and APM gene expression. In ccRCC, the levels of Th17 cells and the ratio of CD8+ T/Treg levels are associated with improved survival whereas the levels of Th2 cells and Tregs are associated with negative clinical outcome. Our analysis illustrates the utility of computational immune cell decomposition for solid tumors, and the potential of this method to guide clinical decision-making.

Cancer Biology

EVfold.org: Evolutionary Couplings and Protein 3D Structure Prediction

Recently developed maximum entropy methods infer evolutionary constraints on protein function and structure from the millions of protein sequences available in genomic databases. The EVfold web server (at EVfold.org) makes these methods available to predict functional and structural interactions in proteins. The key algorithmic development has been to disentangle direct and indirect residue-residue correlations in large multiple sequence alignments and derive direct residue-residue evolutionary couplings (EVcouplings or ECs). For proteins of unknown structure, distance constraints obtained from evolutionarily couplings between residue pairs are used to de novo predict all-atom 3D structures, often to good accuracy. Given sufficient sequence information in a protein family, this is a major advance toward solving the problem of computing the native 3D fold of proteins from sequence information alone.\n\nAvailabilityEVfold server at http://evfold.org/\n\nContactevfoldtest@gmail.com\n\nAbbreviations

Bioinformatics

Mitochondrial DNA Copy Number Variation Across Human Cancers

In cancer, mitochondrial dysfunction, through mutations, deletions, and changes in copy number of mitochondrial DNA (mtDNA), contributes to the malignant transformation and progress of tumors. Here, we report the first large-scale survey of mtDNA copy number variation across 21 distinct solid tumor types, examining over 13,000 tissue samples profiled with next-generation sequencing methods. We find a tendency for cancers, especially of the bladder and kidney, to be significantly depleted of mtDNA, relative to matched normal tissue. We show that mtDNA copy number is correlated to the expression of mitochondrially-localized metabolic pathways, suggesting that mtDNA copy number variation reflect gross changes in mitochondrial metabolic activity. Finally, we identify a subset of tumor-type-specific somatic alterations, including IDH1 and NF1 mutations in gliomas, whose incidence is strongly correlated to mtDNA copy number. Our findings suggest that modulation of mtDNA copy number may play a role in the pathology of cancer.

Genomics

PaxtoolsR: Pathway Analysis in R Using Pathway Commons

PurposePaxtoolsR package enables access to pathway data represented in the BioPAX format and made available through the Pathway Commons webservice for users of the R language. Features include the extraction, merging, and validation of pathway data represented in the BioPAX format. This package also provides novel pathway datasets and advanced querying features for R users through the Pathway Commons webservice allowing users to query, extract, and retrieve data and integrate this data with local BioPAX datasets.\n\nAvailabilityThe PaxtoolsR package is compatible with R 3.1.1 on Windows, Mac OS X, and Linux using Bioconductor 3.0 and is available through the Bioconductor R package repository along with source code and a tutorial vignette describing common tasks, such as data visualization and gene set enrichment analysis. Source code and documentation are at http://bioconductor.org/packages/release/bioc/html/paxtoolsr.html. This plugin is free, open-source and licensed under the GNU Lesser General Public License (LGPL) v3.0.\n\nContactpaxtools@cbio.mskcc.org

Bioinformatics

Protein Domain Hotspots Reveal Functional Mutations across Genes in Cancer

In cancer genomics, frequent recurrence of mutations in independent tumor samples is a strong indication of functional impact. However, rare functional mutations can escape detection by recurrence analysis for lack of statistical power. We address this problem by extending the notion of recurrence of mutations from single genes to gene families that share homologous protein domains. In addition to lowering the threshold of detection, this sharpens the functional interpretation of the impact of mutations, as protein domains more succinctly embody function than entire genes. Mapping mutations in 22 different tumor types to equivalent positions in multiple sequence alignments of protein domains, we confirm well-known functional mutation hotspots and make two types of discoveries: 1) identification and functional interpretation of uncharacterized rare variants in one gene that are equivalent to well-characterized mutations in canonical cancer genes, such as uncharacterized ERBB4 (S303F) mutations that are analogous to canonical ERRB2 (S310F) mutations in the furin-like domain, and 2) detection of previously unknown mutation hotspots with novel functional implications. With the rapid expansion of cancer genomics projects, protein domain hotspot analysis is likely to provide many more leads linking mutations in proteins to the cancer phenotype.

Genomics

A multi-method approach for proteomic network inference in 11 human cancers

Protein expression and post-translational modification levels are tightly regulated in neoplastic cells to maintain cellular processes known as cancer hallmarks. The first Pan-Cancer initiative of The Cancer Genome Atlas (TCGA) Research Network has aggregated protein expression profiles for 3,467 patient samples from 11 tumor types using the antibody based reverse phase protein array (RPPA) technology. The resultant proteomic data can be utilized to computationally infer protein-protein interaction (PPI) networks and to study the commonalities and differences across tumor types. In this study, we compare the performance of 13 established network inference methods in their capacity to retrieve literature-curated pathway interactions from RPPA data. We observe that no single method has the best performance in all tumor types, but a group of six methods, including diverse techniques such as correlation, mutual information, and regression, consistently rank highly among the tested methods. A consensus network from this high-performing group reveals that signal transduction events involving receptor tyrosine kinases (RTKs), the RAS/MAPK pathway, and the PI3K/AKT/mTOR pathway, as well as innate and adaptive immunity signaling, are the most significant PPIs shared across all tumor types. Our results illustrate the utility of the RPPA platform as a tool to study proteomic networks in cancer.\n\nAvailabilityPPI networks from the TCGA or user-provided data can be visualized with the ProtNet web application at http://www.sanderlab.org/protnet/.

Bioinformatics

Systematic identification of cancer driving signaling pathways based on mutual exclusivity of genomic alterations

Recent cancer genome studies have identified numerous genomic alterations in cancer genomes. It is hypothesized that only a fraction of these genomic alterations drive the progression of cancer - often called driver mutations. Current sample sizes for cancer studies, often in the hundreds, are sufficient to detect pivotal drivers solely based on their high frequency of alterations. In cases where the alterations for a single function are distributed among multiple genes of a common pathway, however, single gene alteration frequencies might not be statistically significant. In such cases, we expect to observe that most samples are altered in only one of those alternative genes because additional alterations would not convey an additional selective advantage to the tumor. This leads to a mutual exclusion pattern of alterations, that can be exploited to identify these groups.\n\nWe developed a novel method for the identification of sets of mutually exclusive gene alterations in a signaling network. We scan the groups of genes with a common downstream effect, using a mutual exclusivity criterion that makes sure that each gene in the group significantly contributes to the mutual exclusivity pattern. We have tested the method on all available TCGA cancer genomics datasets, and detected multiple previously unreported alterations that show significant mutual exclusivity and are likely to be driver events.

Cancer Biology

Extensive Decoupling of Metabolic Genes in Cancer

Tumorigenesis involves, among other factors, the alteration of metabolic gene expression to support malignant, unrestrained proliferation. Here, we examine how the altered metabolism of cancer cells is reflected in changes in co-expression patterns of metabolic genes between normal and tumor tissues. Our emphasis on changes in the interactions of pairs of genes, rather than on the expression levels of individual genes, exposes changes in the activity of metabolic pathways which do not necessarily show clear patterns of over- or under-expression. We report the existence of key metabolic genes which act as hubs of differential co-expression, showing significantly different co-regulation patterns between normal and tumor states. Notably, we find that the extent of differential co-expression of a gene is only weakly correlated with its differential expression, suggesting that the two measures probe different features of metabolism. By leveraging our findings against existing pathway knowledge, we extract networks of functionally connected differentially co-expressed genes and the transcription factors which regulate them. Doing so, we identify a previously unreported network of dysregulated metabolic genes in clear cell renal cell carcinoma transcriptionally controlled by the transcription factor HNF4A. While HNF4A shows no significant differential expression, the co-expression HNF4A and several of its regulated target genes in normal tissue is completely abrogated in tumor tissue. Finally, we aggregate the results of differential co-expression analysis across seven distinct cancer types to identify pairs of metabolic genes which may be recurrently dysregulated. Among our results is a cluster of four genes, all located in the mitochondrial electron transport chain, which show significant loss of co-expression in tumor tissue, pointing to potential mitochondrial dysfunction in these tumor types.

Systems Biology

Perturbation biology models predict c-Myc as an effective co-target in RAF inhibitor resistant melanoma cells

Systematic prediction of cellular response to perturbations is a central challenge in biology, both for mechanistic explanations and for the design of effective therapeutic interventions. We addressed this challenge using a computational/experimental method, termed perturbation biology, which combines high-throughput (phospho)proteomic and phenotypic response profiles to targeted perturbations, prior information from signaling databases and network inference algorithms from statistical physics. The resulting network models are computationally executed to predict the effects of tens of thousands of untested perturbations. We report cell type-specific network models of signaling in RAF-inhibitor resistant melanoma cells based on data from 89 combinatorial perturbation conditions and 143 readouts per condition. Quantitative simulations predicted c-Myc as an effective co-target with BRAF or MEK. Experiments showed that co-targeting c-Myc, using the BET bromodomain inhibitor JQ1, and the RAF/MEK pathway, using kinase inhibitors is both effective and synergistic in this context. We propose these combinations as pre-clinical candidates to prevent or overcome RAF inhibitor resistance in melanoma.

Cancer Biology

Accurate prediction of transmembrane β-barrel proteins from sequences

Transmembrane {beta}-barrels are known to play major roles in substrate transport and protein biogenesis in gram-negative bacteria, chloroplasts and mitochondria. However, the exact number of transmembrane {beta}-barrel families is unknown and experimental structure determination is challenging. In theory, if one knows the number of strands in the {beta}-barrel, then the 3D structure of the barrel could be trivial, but current topology predictions do not predict accurate structures and are unable to give information beyond the {beta}-strands in the barrel. Recent work has shown successful prediction of globular and alpha-helical membrane proteins from sequence alignments, by using high ranked evolutionary couplings between residues as distance constraints to fold extended polypeptides. However, these methods, have not addressed the calculation of precise {beta}-sheet hydrogen bonding that defines transmembrane {beta}-barrels, and would be required to fold these proteins successfully. Hence we developed a method (EVFold_BB) that can successfully model transmembrane {beta}-barrels by combining evolutionary couplings together with topology predictions. EVFold_BB is validated by the accurate all-atom 3D modeling of 18 proteins, representing all known membrane {beta}-barrel families that have sufficient sequences available. To demonstrate the potential of our approach we predict the unknown 3D structure of the LptD protein, the plausibility of its accuracy is supported by the blindly predicted benchmarks, and is consistent with experimental observations. Our approach can naturally be extended to all unknown {beta}-barrel proteins with sufficient sequence information.

Bioinformatics

Cancer-associated recurrent mutations in RNase III domains of DICER1

Mutations in the RNase IIIb domain of DICER1 are known to disrupt processing of 5p-strand pre-miRNAs and these mutations have previously been associated with cancer. Using data from the Cancer Genome Atlas project, we show that these mutations are recurrent across four cancer types and that a previously uncharacterized recurrent mutation in the adjacent RNase IIIa domain also disrupts 5p-strand miRNA processing. Analysis of the downstream effects of the resulting imbalance 5p/3p shows a statistically significant effect on the expression of mRNAs targeted by major conserved miRNA families. In summary, these mutations in DICER1 lead to an imbalance in miRNA strands, which has an effect on mRNA transcript levels that appear to contribute to the oncogenesis.

Cancer Biology

Sequence co-evolution gives 3D contacts and structures of protein complexes

Protein-protein interactions are fundamental to many biological processes. Experimental screens have identified tens of thousands of interactions and structural biology has provided detailed functional insight for select 3D protein complexes. An alternative rich source of information about protein interactions is the evolutionary sequence record. Building on earlier work, we show that analysis of correlated evolutionary sequence changes across proteins identifies residues that are close in space with sufficient accuracy to determine the three-dimensional structure of the protein complexes. We evaluate prediction performance in blinded tests on 76 complexes of known 3D structure, predict protein-protein contacts in 32 complexes of unknown structure, and demonstrate how evolutionary couplings can be used to distinguish between interacting and non-interacting protein pairs in a large complex. With the current growth of sequence databases, we expect that the method can be generalized to genome-wide elucidation of protein-protein interaction networks and used for interaction predictions at residue resolution.

Bioinformatics