bioRxiv ScienceSearch

Biology subjects

Piccolo, S. R.

Publications and source records attributed to Piccolo, S. R..

4 recordsLinked to original sources

Combined analysis of genome sequencing and RNA-motifs reveals novel damaging non-coding mutations in human tumors

A major challenge in cancer research is to determine the biological and clinical significance of somatic mutations in non-coding regions. This has been studied in terms of recurrence, functional impact, and association to individual regulatory sites, but the combinatorial contribution of mutations to common RNA regulatory motifs has not been explored. We developed a new method, MIRA, to perform the first comprehensive study of significantly mutated regions (SMRs) affecting binding sites for RNA-binding proteins (RBPs) in cancer. Extracting signals related to RNA-related selection processes and using RNA sequencing data from the same samples we identified alterations in RNA expression and splicing linked to mutations on RBP binding sites. We found SRSF10 and MBNL1 motifs in introns, HNRPLL motifs at 5 UTRs, as well as 5 and 3 splice-site motifs, among others, with specific mutational patterns that disrupt the motif and impact RNA processing. MIRA facilitates the integrative analysis of multiple genome sites that operate collectively through common RBPs and can aid in the interpretation of non-coding variants in cancer. MIRA is available at https://github.com/comprna/mira.

cancer biology

Curated compendium of human transcriptional biomarker data

Genome-wide transcriptional profiles provide broad insights into cellular activity. One important use of such data isto identify relationships between transcription levels and patient outcomes. These translational insights can guide the development of biomarkers for predicting outcomes in clinical settings. Over the past decades, data from many translational-biomarker studies have been deposited in public repositories, enabling other scientists to reuse the data in follow-up studies. However, data-reuse efforts require considerable time and expertise because transcriptional data are generated using heterogeneous profiling technologies, preprocessed using diverse normalization procedures, and annotated in non-standard ways. To address this problem, we curated a compendium of 45 translational-biomarker datasets from the public domain. To increase the datas utility, we reprocessed the raw expression data using a standard computational pipeline and standardized the clinical annotations in a fully reproducible manner (see osf.io/ssk3t). We believe these data will be particularly useful to researchers seeking to validate gene-level findings or to perform benchmarking studies--for example, to compare and optimize machine-learning algorithms ability to predict biomedical outcomes.

cancer biology

Classifying cancer genome aberrations by their mutually exclusive effects on transcription

BackgroundMalignant tumors are typically caused by a conglomeration of genomic aberrations--including point mutations, small insertions, small deletions, and large copy-number variations. In some cases, specific chemotherapies and targeted drug treatments are effective against tumors that harbor certain genomic aberrations. However, predictive aberrations (biomarkers) have not been identified for many tumors. One way to address this problem is to examine the downstream, transcriptional effects of genomic aberrations and to identify characteristic patterns. Even though two tumors harbor different genomic aberrations, the transcriptional effects of those aberrations may be similar. These patterns could be used to inform treatment choices.\n\nResultsWe hypothesized that by grouping genomic aberrations according to their downstream transcriptional effects, we could identify drug-repurposing candidates that coincide with prior knowledge. To evaluate this hypothesis, we developed a computational pipeline that uses supervised machine learning to identify similarities among mutually exclusive genomic aberrations based on their transcriptional effects. We used data from 9300 tumors across 25 cancer types from The Cancer Genome Atlas to evaluate our approach. Our pipeline recapitulates known relationships in cancer pathways, and it identifies gene pairs known to predict responses to the same treatments.\n\nConclusionWe show that transcriptional data offer promise as a way to group genomic aberrations according to their downstream effects, and these groupings recapitulate known relationships. Our approach shows potential to help pharmacologists and clinical trialists narrow the search space for candidate gene/drug associations, including for rare mutations, and for identifying potential drug-repurposing opportunities.

genomics

A community-based collaboration to build prediction models for short-term discontinuation of docetaxel in metastatic castration-resistant prostate cancer patients

BackgroundDocetaxel has a demonstrated survival benefit for metastatic castration-resistant prostate cancer (mCRPC). However, 10-20% of patients discontinue docetaxel prematurely because of toxicity-induced adverse events, and managing risk factors for toxicity remains an ongoing challenge for health care providers and patients. Prospective identification of high-risk patients for early discontinuation has the potential to assist clinical decision-making and can improve the design of more efficient clinical trials. In partnership with Project Data Sphere (PDS), a non-profit initiative facilitating clinical trial data-sharing, we designed an open-data, crowdsourced DREAM (Dialogue for Reverse Engineering Assessments and Methods) Challenge for developing models to predict early discontinuation of docetaxel\n\nMethodsData from the comparator arms of four phase III clinical trials in first-line mCRPC were obtained from PDS, including 476 patients treated with docetaxel and prednisone from the ASCENT2 trial, 598 patients treated with docetaxel, prednisone/prednisolone, and placebo in the VENICE trial, 526 patients treated with docetaxel, prednisone, and placebo in the MAINSAIL trial, and 528 patients treated with docetaxel and placebo in the ENTHUSE 33 trial. Early discontinuation was defined as treatment stoppage within three months due to adverse treatment effects. Over 150 clinical features including laboratory values, medical history, lesion measures, prior treatment, and demographic variables were curated and made freely available for model building for all four trials. The ASCENT2, VENICE, and MAINSAIL trial data sets formed the training set that also included patient discontinuation status. The ENTHUSE 33 trial, with patient discontinuation status hidden, was used as an independent validation set to evaluate model performance. Prediction performance was assessed using area under the precision-recall curve (AUPRC) and the Bayes factor was used to compare the performance between prediction models.\n\nResultsThe frequency of early discontinuation was similar between training (ASCENT2, VENICE, and MAINSAIL) and validation (ENTHUSE 33) sets, 12.3% versus 10.4% of docetaxel-treated patients, respectively. In total, 34 independent teams submitted predictions from 61 different models. AUPRC ranged from 0.088 to 0.178 across submissions with a random model performance of 0.104. Seven models with comparable AUPRC scores (Bayes factor [&le;]; 3) were observed to outperform all other models. A post-challenge analysis of risk predictions generated by these seven models revealed three distinct patient subgroups: patients consistently predicted to be at high-risk or low-risk for early discontinuation and those with discordant risk predictions. Early discontinuation events were two-times higher in the high-versus low-risk subgroup and baseline clinical features such as presence/absence of metastatic liver lesions, and prior treatment with analgesics and ACE inhibitors exhibited statistically significant differences between the high- and low-risk subgroups (adjusted P < 0.05). An ensemble-based model constructed from a post-Challenge community collaboration resulted in the best overall prediction performance (AUPRC = 0.230) and represented a marked improvement over any individual Challenge submission. A\n\nFindingsOur results demonstrate that routinely collected clinical features can be used to prospectively inform clinicians of mCRPC patients risk to discontinue docetaxel treatment early due to adverse events and to the best of our knowledge is the first to establish performance benchmarks in this area. This work also underscores the \"wisdom of crowds\" approach by demonstrating that improved prediction of patient outcomes is obtainable by combining methods across an extended community. These findings were made possible because data from separate trials were made publicly available and centrally compiled through PDS.

bioinformatics