bioRxiv ScienceSearch

Biology subjects

Linial, M.

Publications and source records attributed to Linial, M..

4 recordsLinked to original sources

Substantial Batch Effects in TCGA Exome Sequences Undermine Pan-Cancer Analysis of Germline Variants

BackgroundIn recent years, research on cancer predisposition germline variants has emerged as a prominent field. The identity of somatic mutations is based on a reliable mapping of the patient germline variants. In addition, the statistics of germline variants frequencies in healthy individuals and cancer patients is the basis for seeking candidates for cancer predisposition genes. The Cancer Genome Atlas (TCGA) is one of the main sources of such data, providing a diverse collection of molecular data including deep sequencing for more than 30 types of cancer from >10,000 patients.\n\nMethodsOur hypothesis in this study is that whole exome sequences from healthy blood samples of cancer patients are not expected to show systematic differences among cancer types. To test this hypothesis, we analyzed common and rare germline variants across six cancer types, covering 2,241 samples from TCGA. In our analysis we accounted for inherent variables in the data including the different variant calling protocols, sequencing platforms, and ethnicity.\n\nResultsWe report on substantial batch effects in germline variants associated with cancer types. We attribute the effect to the specific sequencing centers that produced the data. Specifically, we measured 30% variability in the number of reported germline variants per sample across sequencing centers. The batch effect is further expressed in nucleotide composition and variant frequencies. Importantly, the batch effect causes substantial differences in germline variant distribution patterns across numerous genes, including prominent cancer predisposition genes such as BRCA1, RET, MAX, and KRAS. For most of known cancer predisposition genes, we found a distinct batch-dependent difference in germline variants.\n\nConclusionTCGA germline data is exposed to strong batch effects with substantial variabilities among TCGA sequencing centers. We claim that those batch effects are consequential for numerous TCGA pan-cancer studies. In particular, these effects may compromise the reliability and the potency to detect new cancer predisposition genes. Furthermore, interpretation of pan-cancer analyses should be revisited in view of the source of the genomic data after accounting for the reported batch effects.

cancer biology

The Translation Machinery Is Immune from miRNA Perturbations: A Cell-Based Probabilistic Approach

Mature microRNAs (miRNAs) regulate most human genes through direct base-pairing with mRNAs. We investigate some underlying principles of such regulation. To this end, we overexpressed miRNAs in different cell types and measured the mRNA decay rate under transcriptional arrest. Parameters extracted from these experiments were incorporated into a computational stochastic framework which was developed to simulate the cooperative action of miRNAs in living cells. We identified gene sets that exhibit coordinated behavior with respect to all major miRNAs, over a broad range of overexpression levels. While a small set of genes is highly sensitive to miRNA manipulations, about 180 genes are insensitive to miRNA manipulations as measured by their degree of mRNA retention. The insensitive genes are associated with the translation machinery. We conclude that the stochastic nature of miRNAs reveals an unexpected robustness of gene expression in living cells. Moreover, the use of a systematic probabilistic approach exposes design principles of cells states and in particular, the translational machinery.\n\nHighlightsO_LIA probabilistic-based simulator assesses the cellular response to thousands of miRNA overexpression manipulations\nC_LIO_LIThe translational machinery displays an exceptional resistance to manipulations of miRNAs.\nC_LIO_LIThe insensitivity of the translation machinery to miRNA manipulations is shared by different cell types\nC_LIO_LIThe composition of the most abundant miRNAs dominates cell identity\nC_LI

systems biology

Modeling Functional Genetic Alteration in Cancer Reveals New Candidate Driver Genes

Compiling the catalogue of genes actively involved in tumorigenesis (known as cancer drivers) is an ongoing endeavor, with profound implications to the understanding of tumorigenesis and treatment of the disease. An abundance of computational methods have been developed to screening the genome for candidate driver genes based on genomic data of somatic mutations in tumors. Most methods rely on detecting genes displaying excessive mutation rates compared to some background model. This approach is susceptible to false discoveries, due to its sensitivity to the assumptions of the background model, such as the need to account for hyper-mutated samples, cancer types and genomic loci. We present a fundamentally different approach. Instead of focusing on the number of mutations, we examine their content, and their expected effects on the functions of genes. We use a machine-learning model to predict functional effect scores of somatic mutations. For each gene, we compare the distribution of observed effect scores with the distribution expected at random, and report genes showing significant bias. By applying our framework on the ~20k protein-coding human genes, we detected 593 genes showing significant bias towards harmful mutations in the context of cancer. In contrast, we found only 6 significant genes biased in the opposite direction. The list of 593 genes, constructed without any prior knowledge of their role in cancer, shows an overwhelming overlap with known cancer driver genes, but also highlights many overlooked genes. These overlooked genes are promising candidates for novel cancer drivers. Our model is generic and is not restricted to the context of cancer. Applying the same framework to data of human-population genetic variation reveals the opposite trend. Unlike cancer, which is dominated by a bias towards harmful mutations, long-term evolution in healthy individuals results a bias towards less harmful mutations. The underlying assumptions of our framework are minimal, making it ideal for analyzing genetic data in search of genes subjected to positive or negative selection. It is fully open sourced and available for installation and use. Our framework presents a substantial development towards the application of state-of-the-art machine-learning algorithms in genetic studies.

bioinformatics

Lowest expressing microRNAs capture indispensable information - identifying cancer types

The primary function of microRNAs (miRNAs) is to maintain cell homeostasis. In cancerous tissues miRNAs expression undergo drastic alterations. In this study, we used miRNA expression profiles from The Cancer Genome Atlas (TCGA) of 24 cancer types and 3 healthy tissues, collected from >8500 samples. We seek to classify the cancers origin and tissue identification using the expression from 1046 reported miRNAs. Despite an apparent uniform appearance of miRNAs among cancerous samples, we recover indispensable information from lowly expressed miRNAs regarding the cancer/tissue types. Multiclass support vector machine classification yields an average recall of 58% in identifying the correct tissue and tumor types. Data discretization has led to substantial improvement reaching an average recall of 91% (95% median). We propose a straightforward protocol as a crucial step in classifying tumors of unknown primary origin. Our counter-intuitive conclusion is that in almost all cancer types, highly expressing miRNAs mask the significant signal that lower expressed miRNAs provide.

bioinformatics