bioRxiv ScienceSearch

Biology subjects

Lodewyk Wessels

Publications and source records attributed to Lodewyk Wessels.

3 recordsLinked to original sources

A novel independence test for somatic alterations in cancer shows that biology drives mutual exclusivity but chance explains co-occurrence

Just like recurrent somatic alterations characterize cancer genes, mutually exclusive or co-occurring alterations across genes suggest functional interactions. Identifying such patterns in large cancer studies thus helps the discovery of unknown interactions. Many studies use Fishers exact test or simple permutation procedures for this purpose. These tests assume identical gene alteration probabilities across tumors, which is not true for cancer. We show that violating this assumption yields many spurious co-occurrences and misses many mutual exclusivities. We present DISCOVER, a novel statistical test that addresses the limitations of existing tests. In a comparison with six published mutual exclusivity tests, DISCOVER is more sensitive while controlling its false positive rate. A pan-cancer analysis using DISCOVER finds no evidence for widespread co-occurrence. Most co-occurrences previously detected do not exceed expectation by chance. In contrast, many mutual exclusivities are identified. These cover well known genes involved in the cell cycle and growth factor signaling. Interestingly, also lesser known regulators of the cell cycle and Hedgehog signaling are identified.\n\nAvailabilityR and Python implementations of DISCOVER, as well as Jupyter notebooks for reproducing all results and figures from this paper can be found at http://ccb.nki.nl/software/discover.

Bioinformatics

OncoScape: Exploring the cancer aberration landscape by genomic data fusion

Although large-scale efforts for molecular profiling of cancer samples provide multiple data types for many samples, most approaches for finding candidate cancer genes rely on somatic mutations and DNA copy number only. We present a new method, OncoScape, which, for the first time, exploits five complementary data types across 11 cancer types to identify new candidate cancer genes. We find many rarely mutated genes that are strongly affected by other aberrations. We retrieve the majority of known cancer genes but also new candidates such as STK31 and MSRA with very high confidence. Several genes show a dual oncogene-and tumor suppressor-like behavior depending on the tumor type. Most notably, the well-known tumor suppressor RB1 shows strong oncogene-like signal in colon cancer. We applied OncoScape to cell lines representing ten cancer types, providing the most comprehensive comparison of aberrations in cell lines and tumor samples to date. This revealed that glioblastoma, breast and colon cancer show strong similarity between cell lines and tumors, while head and neck squamous cell carcinoma and bladder cancer, exhibit very little similarity between cell lines and tumors. To facilitate exploration of the cancer aberration landscape, we created a web portal enabling interactive analysis of OncoScape results.

Bioinformatics

Logic models to predict continuous outputs based on binary inputs with an application to personalized cancer therapy

Mining large datasets using machine learning approaches often leads to models that are hard to interpret and not amenable to the generation of hypotheses that can be experimentally tested. Finding actionable knowledge is becoming more important, but also more challenging as datasets grow in size and complexity. We present Logic Optimization for Binary Input to Continuous Output (LOBICO), a computational approach that infers small and easily interpretable logic models of binary input features that explain a binarized continuous output variable. Although the continuous output variable is binarized prior to optimization, the continuous information is retained to find the optimal logic model. Applying LOBICO to a large cancer cell line panel, we find that logic combinations of multiple mutations are more predictive of drug response than single gene predictors. Importantly, we show that the use of the continuous information leads to robust and more accurate logic models. LOBICO is formulated as an integer programming problem, which enables rapid computation on large datasets. Moreover, LOBICO implements the ability to uncover logic models around predefined operating points in terms of sensitivity and specificity. As such, it represents an important step towards practical application of interpretable logic models.

Bioinformatics