bioRxiv ScienceSearch

Biology subjects

Koester, J.

Publications and source records attributed to Koester, J..

2 recordsLinked to original sources

Full-length de novo viral quasispecies assembly through variation graph construction

MotivationViruses populate their hosts as a viral quasispecies: a collection of genetically related mutant strains. Viral quasispecies assembly refers to reconstructing the strain-specific haplotypes from read data, and predicting their relative abundances within the mix of strains, an important step for various treatment-related reasons. Reference-genome-independent (\"de novo\") approaches have yielded benefits over reference-guided approaches, because reference-induced biases can become overwhelming when dealing with divergent strains. While being very accurate, extant de novo methods only yield rather short contigs. It remains to reconstruct full-length haplotypes together with their abundances from such contigs.\n\nMethodWe first construct a variation graph, a recently popular, suitable structure for arranging and integrating several related genomes, from the short input contigs, without making use of a reference genome. To obtain paths through the variation graph that reflect the original haplotypes, we solve a minimization problem that yields a selection of maximal-length paths that is optimal in terms of being compatible with the read coverages computed for the nodes of the variation graph. We output the resulting selection of maximal length paths as the haplotypes, together with their abundances.\n\nResultsBenchmarking experiments on challenging simulated data sets show significant improvements in assembly contiguity compared to the input contigs, while preserving low error rates. As a consequence, our method outperforms all state-of-the-art viral quasispecies assemblers that aim at the construction of full-length haplotypes, in terms of various relevant assembly measures. Our tool, Virus-VG, is publicly available at https://bitbucket.org/jbaaijens/virus-vg.

bioinformatics

Modeling and simulating networks of interdependent protein interactions

Protein interactions are fundamental building blocks of biochemical reaction systems underlying cellular functions. The complexity and functionality of these systems emerge not only from the protein interactions themselves but also from the dependencies between these interactions, e.g., allosteric effects, mutual exclusion or steric hindrance. Therefore, formal models for integrating and using information about such dependencies are of high interest. We present an approach for endowing protein networks with interaction dependencies using propositional logic, thereby obtaining constrained protein interaction networks (\"constrained networks\"). The construction of these networks is based on public interaction databases and known as well as text-mined interaction dependencies. We present an efficient data structure and algorithm to simulate protein complex formation in constrained networks. The efficiency of the model allows a fast simulation and enables the analysis of many proteins in large networks. Therefore, we are able to simulate perturbation effects (knockout and overexpression of single or multiple proteins, changes of protein concentrations). We illustrate how our model can be used to analyze a partially constrained human adhesome network. Comparing complex formation under known dependencies against without dependencies, we find that interaction dependencies limit the resulting complex sizes. Further we demonstrate that our model enables us to investigate how the interplay of network topology and interaction dependencies influences the propagation of perturbation effects. Our simulation software CPINSim (for Constrained Protein Interaction Network Simulator) is available under the MIT license at http://github.com/BiancaStoecker/cpinsimandviaBioconda (https://bioconda.github.io).\n\nAuthor summaryProteins are the main molecular tools of cells. They do not act individually, but rather collectively in order to peform complex cellular actions. Recent years have led to a relatively good understanding about which proteins may interact, both in general and in specific conditions, leading to the definition of protein interaction networks. However, the reality is more complex, and protein interactions are not independent of each other. Instead, several potential interaction partners of a specific protein may compete for the same binding domain, making all of these interactions mutually exclusive. Additionally, a binding of a protein to another one can enable or prevent their interactions with other proteins, even if those interactions are mediated by different domains. Hence, understanding how the dependencies (or constraints) of protein interactions affect the behaviour of the system is an important and timely goal, as data is now becoming available. Here we present a mathematical framework to formalize such interaction constraints and incorporate them into the simulation of protein complex formation. With our framework, we are able to better understand how perturbations of single proteins (knockout or overexpression) impact other proteins in the network.

systems biology