bioRxiv ScienceSearch

Biology subjects

Valentina Boeva

Publications and source records attributed to Valentina Boeva.

3 recordsLinked to original sources

Calculating biological module enrichment or depletion and visualizing data on large-scale molecular maps with ACSNMineR and RNaviCell R packages

Biological pathways or modules represent sets of interactions or functional relationships occurring at the molecular level in living cells. A large body of knowledge on pathways is organized in public databases such as the KEGG, Reactome, or in more specialized repositories, such as the Atlas of Cancer Signaling Network (ACSN). All these open biological databases facilitate analyses, improving our understanding of cellular systems. We hereby describe the R package ACSNMineR for calculation of enrichment or depletion of lists of genes of interest in biological pathways. ACSNMineR integrates ACSN molecular pathways, but can use any molecular pathway encoded as a GMT file, for instance sets of genes available in the Molecular Signatures Database (MSigDB). We also present the R package RNaviCell, that can be used in conjunction with ACSNMineR to visualize different data types on web-based, interactive ACSN maps. We illustrate the functionalities of the two packages with biological data taken from large-scale cancer datasets.

Bioinformatics

Clonal assessment of functional mutations in cancer based on a genotype-aware method for clonal reconstruction

In cancer, clonal evolution is characterized based on single nucleotide variants and copy number alterations. Nonetheless, previous methods failed to combine information from both sources to accurately reconstruct clonal populations in a given tumor sample or in a set of tumor samples coming from the same patient. Moreover, previous methods accepted as input all variants predicted by variant-callers, regardless of differences in dispersion of variant allele frequencies (VAFs) due to uneven depth of coverage and possible presence of strand bias, prohibiting accurate inference of clonal architecture. We present a general framework for assignment of functional mutations to specific cancer clones, which is based on distinction between passenger variants with expected low dispersion of VAF versus putative functional variants, which may not be used for the reconstruction of cancer clonal architecture but can be assigned to inferred clones at the final stage. The key element of our framework is QuantumClone, a method to cluster variants into clones, which we have thoroughly tested on simulated data. QuantumClone takes into account VAFs and genotypes of corresponding regions together with information about normal cell contamination. We applied our framework to whole genome sequencing data for 19 neuroblastoma trios each including constitutional, diagnosis and relapse samples. We discovered specific pathways recurrently altered by deleterious mutations in different clonal populations. Some such pathways were previously reported (e.g., MAPK and neuritogenesis) while some were novel (e.g., epithelial-mesenchymal transition, cell survival and DNA repair). Most pathways and their modules had more mutations at relapse compared to diagnosis.

Cancer Biology

Computational Pan-Genomics: Status, Promises and Challenges

Many disciplines, from human genetics and oncology to plant breeding, microbiology and virology, commonly face the challenge of analyzing rapidly increasing numbers of genomes. In case of Homo sapiens, the number of sequenced genomes will approach hundreds of thousands in the next few years. Simply scaling up established bioinformatics pipelines will not be sufficient for leveraging the full potential of such rich genomic datasets. Instead, novel, qualitatively different computational methods and paradigms are needed. We will witness the rapid extension of computational pan-genomics, a new sub-area of research in computational biology. In this paper, we generalize existing definitions and understand a pan-genome as any collection of genomic sequences to be analyzed jointly or to be used as a reference. We examine already available approaches to construct and use pan-genomes, discuss the potential benefits of future technologies and methodologies, and review open challenges from the vantage point of the above-mentioned biological disciplines. As a prominent example for a computational paradigm shift, we particularly highlight the transition from the representation of reference genomes as strings to representations as graphs. We outline how this and other challenges from different application domains translate into common computational problems, point out relevant bioinformatics techniques and identify open problems in computer science. With this review, we aim to increase awareness that a joint approach to computational pan-genomics can help address many of the problems currently faced in various domains.

Genomics