bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Genetic heterogeneity in autism: from single gene to a pathway perspective

AbstractThe extreme genetic heterogeneity of autism spectrum disorder (ASD) represents a major challenge. Recent advances in genetic screening and systems biology approaches have extended our knowledge of the genetic etiology of ASD. In this review, we discuss the paradigm shift from a single gene causation model to pathway perturbation model as a guide to better understand the pathophysiology of ASD. We discuss recent genetic findings obtained through next-generation sequencing (NGS) and examine various integrative analyses using systems biology and complex networks approaches that identify convergent patterns of genetic elements associated with ASD. This review provides a summary of the genetic findings of family-based genome screening studies.

Genetics

Selection on network dynamics drives differential rates of protein domain evolution

The long-held principle that functionally important proteins evolve slowly has recently been challenged by studies in mice and yeast showing that the severity of a protein knockout only weakly predicts that protein's rate of evolution. However, the relevance of these studies to evolutionary changes within proteins is unknown, because amino acid substitutions, unlike knockouts, often only slightly perturb protein activity. To quantify the phenotypic effect of small biochemical perturbations, we developed an approach to use computational systems biology models to measure the influence of individual reaction rate constants on network dynamics. We show that this dynamical influence is predictive of protein domain evolutionary rate in vertebrates and yeast, even after controlling for expression level and breadth, network topology, and knockout effect. Thus, our results not only demonstrate the importance of protein domain function in determining evolutionary rate, but also the power of systems biology modeling to uncover unanticipated evolutionary forces.

Evolutionary Biology

Elucidating interplay of speed and accuracy in biological error correction

One of the most fascinating features of biological systems is the ability to sustain high accuracy of all major cellular processes despite the stochastic nature of underlying chemical processes. It is widely believed that such low errors are the result of the error correcting mechanism known as kinetic proofreading. However, it is usually argued that enhancing the accuracy should result in slowing down the process leading to so-called speed-accuracy trade-off. We developed a discrete-state stochastic framework that allowed us to investigate the mechanisms of the proofreading using the method of first-passage processes. With this framework, we simultaneously analyzed speed and accuracy of the two fundamental biological processes, DNA replication and tRNA selection during the translation. The results indicate that speed-accuracy trade-off is not always observed. However, when the trade-off is present, the biological systems tend to optimize the speed rather than the accuracy of the processes, as long as the error level is tolerable. Additional constraints due to the energetic cost of proofreading also play a role in the error correcting process. Our theoretical findings provide a new microscopic picture of how complex biological processes are able to function so fast with a high accuracy.

biophysics

Interacting networks of resistance, virulence and core machinery genes identified by genome-wide epistasis analysis

Recent advances in the scale and diversity of population genomic datasets for bacteria now provide the potential for genome-wide patterns of co-evolution to be studied at the resolution of individual bases. The major human pathogen Streptococcus pneumoniae represents the first bacterial organism for which densely enough sampled population data became available for such an analysis. Here we describe a new statistical method, genomeDCA, which uses recent advances in computational structural biology to identify the polymorphic loci under the strongest co-evolutionary pressures. Genome data from over three thousand pneumococcal isolates identified 5,199 putative epistatic interactions between 1,936 sites. Over three-quarters of the links were between sites within the pbp2x, pbp1a and pbp2b genes, the sequences of which are critical in determining non-susceptibility to beta-lactam antibiotics. A network-based analysis found these genes were also coupled to that encoding dihydrofolate reductase, changes to which underlie trimethoprim resistance. Distinct from these resistance genes, a large network component of 384 protein coding sequences encompassed many genes critical in basic cellular functions, while another distinct component included genes associated with virulence. These results have the potential both to identify previously unsuspected protein-protein interactions, as well as genes making independent contributions to the same phenotype. This approach greatly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for experimental work.\n\nAuthor SummaryEpistatic interactions between polymorphisms in DNA are recognized as important drivers of evolution in numerous organisms. Study of epistasis in bacteria has been hampered by the lack of both densely sampled population genomic data, suitable statistical models and powerful inference algorithms for extremely high-dimensional parameter spaces. We introduce the first model-based method for genome-wide epistasis analysis and use the largest available bacterial population genome data set on Streptococcus pneumoniae (the pneumococcus) to demonstrate its potential for biological discovery. Our approach reveals interacting networks of resistance, virulence and core machinery genes in the pneumococcus, which highlights putative candidates for novel drug targets. Our method significantly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for experimental work.

Genetics

Exploring community structure in biological networks with random graphs

BackgroundCommunity structure is ubiquitous in biological networks. There has been an increased interest in unraveling the community structure of biological systems as it may provide important insights into a systems functional components and the impact of local structures on dynamics at a global scale. Choosing an appropriate community detection algorithm to identify the community structure in an empirical network can be difficult, however, as the many algorithms available are based on a variety of cost functions and are difficult to validate. Even when community structure is identified in an empirical system, disentangling the effect of community structure from other network properties such as clustering coefficient and assortativity can be a challenge.\n\nResultsHere, we develop a generative model to produce undirected, simple, connected graphs with a specified degrees and pattern of communities, while maintaining a graph structure that is as random as possible. Additionally, we demonstrate two important applications of our model: (a) to generate networks that can be used to benchmark existing and new algorithms for detecting communities in biological networks; and (b) to generate null models to serve as random controls when investigating the impact of complex network features beyond the byproduct of degree and modularity in empirical biological networks.\n\nConclusionOur model allows for the systematic study of the presence of community structure and its impact on network function and dynamics. This process is a crucial step in unraveling the functional consequences of the structural properties of biological systems and uncovering the mechanisms that drive these systems.

Bioinformatics

Survival of the simplest: the cost of complexity in microbial evolution

The evolution of microbial and viral organisms often generates clonal interference, a mode of competition between genetic clades within a population. In this paper, we show that interference strongly constrains the genetic and phenotypic complexity of evolving systems. Our analysis uses biophysically grounded evolutionary models for an organisms quantitative molecular phenotypes, such as fold stability and enzymatic activity of genes. We find a generic mode of asexual evolution called phenotypic interference with strong implications for systems biology: it couples the stability and function of individual genes to the populations global speed of evolution. This mode occurs over a wide range of evolutionary parameters appropriate for microbial populations. It generates selection against genome complexity, because the fitness cost of mutations increases faster than linearly with the number of genes. Recombination can generate a distinct mode of sexual evolution that eliminates the superlinear cost. We show that positive selection can drive a transition from asexual to facultative sexual evolution, providing a specific, biophysically grounded scenario for the evolution of sex. In a broader context, our analysis suggests that the systems biology of microbial organisms is strongly intertwined with their mode of evolution.

evolutionary biology

Matching models across abstraction levels with Gaussian Processes

Biological systems are often modelled at different levels of abstraction depending on the particular aims/resources of a study. Such different models often provide qualitatively concordant predictions over specific parametrisations, but it is generally unclear whether model predictions are quantitatively in agreement, and whether such agreement holds for different parametrisations. Here we present a generally applicable statistical machine learning methodology to automatically reconcile the predictions of different models across abstraction levels. Our approach is based on defining a correction map, a random function which modifies the output of a model in order to match the statistics of the output of a different model of the same system. We use two biological examples to give a proof-of-principle demonstration of the methodology, and discuss its advantages and potential further applications.

Biophysics

Convection - Diffusion model of talking bacteria

Quorum sensing is cell to cell communication process through chemical signals formally known as autoinducers. When the concentration of quorum sensing molecules reached threshold concentration bacteria are in active state or quorum state. In this article, we propose a mathematical model of quorum sensing systems and study this biological system numerically. Moreover, we compare the different numerical scheme with the batch culture of P.aeruginosa. We observed a negative diffusion coefficient which plays an important role in the quorum sensing mechanism.

biophysics

Boosting Gene Expression Clustering with System-Wide Biological Information: A Robust Autoencoder Approach

Gene expression analysis provides genome-wide insights into the transcriptional activity of a cell. One of the first computational steps in exploration and analysis of the gene expression data is clustering. With a number of standard clustering methods routinely used, most of the methods do not take prior biological information into account. In this paper, we propose a new approach for gene expression clustering analysis. The approach benefits from a new deep learning architecture, Robust Autoencoder, which provides a more accurate high-level representation of the feature sets, and from incorporating prior biological information into the clustering process. We tested our approach on two distinct gene expression datasets and compared the performance with two widely used clustering methods, hierarchical clustering and k-means, as well as with a recent deep learning clustering approach. As a result, our approach outperformed all other clustering methods on the labeled yeast gene expression dataset. Furthermore we showed that it is better in identifying the functionally common clusters than k-means on the unlabeled human gene expression dataset. The results demonstrate that our new deep learning architecture could generalize well the specific properties of gene expression profiles. Furthermore, the results confirm our hypothesis that the prior biological network knowledge could be helpful in the gene expression clustering task.

bioinformatics

New inhibitors of Mycobacterium tuberculosis identified using systems chemical biology

New antibiotics are needed to combat rising resistance, with new Mycobacterium tuberculosis (Mtb) drugs of highest priority. Conventional whole-cell and biochemical antibiotic screens have failed. We developed a novel strategy termed PROSPECT (PRimary screening Of Strains to Prioritize Expanded Chemistry and Targets) in which we screen compounds against pools of strains depleted for essential bacterial targets. We engineered strains targeting 474 Mtb essential genes and screened pools of 100-150 strains against activity-enriched and unbiased compounds libraries, measuring > 8.5-million chemical-genetic interactions. Primary screens identified >10-fold more hits than screening wild-type Mtb alone, with chemical-genetic interactions providing immediate, direct target insight. We identified > 40 novel compounds targeting DNA gyrase, cell wall, tryptophan, folate biosynthesis, and RNA polymerase, as well as inhibitors of a novel target EfpA. Chemical optimization yielded EfpA inhibitors with potent wild-type activity, thus demonstrating PROSPECTs ability to yield inhibitors against novel targets which would have eluded conventional drug discovery.

microbiology

Discovery of the role of a SLOG superfamily biological conflict systems associated protein IodA (YpsA) in oxidative stress protection and cell division inhibition in Gram-positive bacteria

Bacteria adapt to different environments by regulating cell division and several conditions that modulate cell division have been documented. Understanding how bacteria transduce environmental signals to control cell division is critical to comprehend the global network of cell division regulation. In this article we describe a role for Bacillus subtilis YpsA, an uncharacterized protein of the SLOG superfamily of nucleotide and ligand-binding proteins, in cell division. We observed that YpsA provides protection against oxidative stress as cells lacking ypsA show increased susceptibility to hydrogen peroxide treatment. We found that increased expression of ypsA leads to cell division inhibition due to defective assembly of FtsZ, the tubulin-like essential protein that marks the sites of cell division. We showed that cell division inhibition by YpsA is linked to glucose availability. We generated YpsA mutants that are no longer able to inhibit cell division. Finally, we show that the role of YpsA is possibly conserved in Firmicutes, as overproduction of YpsA in Staphylococcus aureus also impairs cell division. Therefore, we propose ypsA to be renamed as iodA for inhibitor of division.\n\nIMPORTANCEAlthough key players of cell division in bacteria have been largely characterized, the factors that regulate these division proteins are still being discovered and evidence for the presence of yet-to-be discovered factors has been accumulating. How bacteria sense the availability of nutrients and how that information is used to regulate cell division positively or negatively is less well-understood even though some examples exist in the literature. We discovered that a protein of hitherto unknown function belonging to the SLOG superfamily of nucleotide/ligand-binding proteins, YpsA, influences cell division in Bacillus subtilis by integrating metabolic status such as the availability of glucose. We showed that YpsA is important for oxidative stress response in B. subtilis. Furthermore, we provide evidence that cell division inhibition function of YpsA is also conserved in another Firmicute Staphylococcus aureus. This first report on the role of YpsA (IodA) brings us a step closer in understanding the complete tool set that bacteria have at their disposal to regulate cell division precisely to adapt to varying environmental conditions.

microbiology

Differences in protein dosage underlie nongenetic differences in traits

Phenotypic expression of many traits varies among isogenic individuals in homogeneous environments. Intrinsic variation in the protein chaperone system affects a wide variety of traits in diverse biological systems. In C. elegans, expression of hsp-16.2 chaperone biomarkers predicts the penetrance of mutations and lifespan after heat shock. But the physiological mechanisms by which cells express different amounts of the biomarker were unknown. Here, we used an in vivo microscopy approach to dissect the mechanisms of cell-to-cell variation in hsp-16.2 biomarker expression, focusing on the intestines, which generate most signal. We found both intrinsic noise and signaling noise are low. The major axis of cell-to-cell variation in gene expression is composed of general differences in protein dosage. Thus, hsp-16.2 biomarkers reveal states of high or low effective dosages for many genes. It is possible that natural variation in protein dosage or chaperone activity may account for missing heritability of some traits.

systems biology

The geometry of heterosis

Heterosis, the superiority of hybrids over their parents for quantitative traits, represents a crucial issue in plant and animal breeding. Heterosis has given rise to countless genetic, genomic and molecular studies, but has rarely been investigated from the point of view of systems biology. We hypothesized that heterosis is an emergent property of living systems resulting from frequent concave relationships between genotypic variables and phenotypes, or between different phenotypic levels. We chose the enzyme-flux relationship as a model of the concave genotype-phenotype (GP) relationship, and showed that heterosis can be easily created in the laboratory. First, we reconstituted in vitro the upper part of glycolysis. We simulated genetic variability of enzyme activity by varying enzyme concentrations in test tubes. Mixing the content of \"parental\" tubes resulted in \"hybrids\", whose fluxes were compared to the parental fluxes. Frequent heterotic fluxes were observed, under conditions that were determined analytically and confirmed by computer simulation. Second, to test this model in a more realistic situation, we modeled the glycolysis/fermentation network in yeast by considering one input flux, glucose, and two output fluxes, glycerol and acetaldehyde. We simulated genetic variability by randomly drawing parental enzyme concentrations under various conditions, and computed the parental and hybrid fluxes using a system of differential equations. Again we found that a majority of hybrids exhibited positive heterosis for metabolic fluxes. Cases of negative heterosis were due to local convexity between certain enzyme concentrations and fluxes. In both approaches, heterosis was maximized when the parents were phenotypically close and when the distributions of parental enzyme concentrations were contrasted and constrained. These conclusions are not restricted to metabolic systems: they only depend on the concavity of the GP relationship, which is commonly observed at various levels of the phenotypic hierarchy, and could account for the pervasiveness of heterosis.

systems biology

Ordinary Differential Equations in Cancer Biology

Ordinary differential equations (ODEs) provide a classical framework to model the dynamics of biological systems, given temporal experimental data. Qualitative analysis of the ODE model can lead to further biological insight and deeper understanding compared to traditional experiments alone. Simulation of the model under various perturbations can generate novel hypotheses and motivate the design of new experiments. This short paper will provide an overview of the ODE modeling framework, and present examples of how ODEs can be used to address problems in cancer biology.

Systems Biology

Network controllability: viruses are driver agents in dynamic molecular systems

In recent years control theory has been applied to biological systems with the aim of identifying the minimum set of molecular interactions that can drive the network to a required state. However in an intra-cellular network it is unclear what control means. To address this limitation we use viral infection, specifically HIV-1 and HCV, as a paradigm to model control of an infected cell. Using a large human signalling network comprised of over 6000 human proteins and more than 34000 directed interactions, we compared two dynamic states: normal/uninfected and infected. Our network controllability analysis demonstrates how a virus efficiently brings the dynamic host system into its control by mostly targeting existing critical control nodes, requiring fewer nodes than in the uninfected network. The driver nodes used by the virus are distributed throughout the pathways in specific locations enabling effective control of the cell via the high control centrality of the viral and targeted host nodes. Furthermore, this viral infection of the human system permits discrimination between available network-control models, and demonstrates the minimum-dominating set (MDS) method better accounts for how biological information and signals are transferred than the maximum matching (MM) method as it identified most of the HIV-1 proteins as critical driver nodes and goes beyond identifying receptors as the only critical driver nodes. This is because MDS, unlike MM, accounts for the inherent non-linearity of signalling pathways. Our results demonstrate control-theory gives a more complete and dynamic understanding of the viral hijack mechanism when compared with previous analyses limited to static single-state networks.

systems biology

The evolution of central dogma of molecular biology: a logic-based dynamic approach

It is nearly half a century past the age of the introduction of the Central Dogma (CD) of molecular biology. This biological axiom has been developed and currently appears to be all the more complex. In this study, we modified CD by adding further species to the CD information flow and mathematically expressed CD within a dynamic framework by using Boolean network based on its present-day and 1965 editions. We show that the enhancement of the Dogma not only now entails a higher level of complexity, but it also shows a higher level of robustness, thus far more consistent with the nature of biological systems. Using this mathematical modeling approach, we put forward a logic-based expression of our conceptual view of molecular biology. Finally, we show that such biological concepts can be converted into dynamic mathematical models using a logic-based approach and thus may be useful as a framework for improving static conceptual models in biology.

systems biology

Improved LC-MS chromatographic alignment increases the accuracy of label-free quantitative proteomics: Comparison of spectral counting versus ion intensity-based proteomic quantification strategies.

The ability to provide an unbiased qualitative and quantitative description of the global changes to proteins in a cell or an organism would permit the systems-wide study of complex biological systems. Label-free quantitative shotgun proteomic strategies (including LC-MS ion intensity quantification and spectral counting) are attractive because of their relatively low cost, ease of implementation, and the lack of multiplexing restrictions when comparing multiple samples. Owing to improvements in the resolution and sensitivity of mass spectrometers, and the availability of analytical software packages, protein quantification by LC-MS ion intensity has increased in popularity. Here, we have addressed the importance of chromatographic alignment on protein quantification, and then assessed how spectral counting compares to ion intensity-based proteomic quantification. Using a spiked-in protein strategy, we analysed two situations that commonly arise in the application of proteomics to cell biology: (i) samples with a small number of proteins of differential abundance in a larger non-changing background, and (ii) samples with a larger number of proteins of differential abundance. To perform these assessments on biologically relevant samples, we used isolated integrin adhesion complexes (IACs). Technical replicate analysis of isolated IACs resulted in a range of alignment scores using the Progenesis QI software package and demonstrated that higher LC-MS chromatographic alignment scores increased the precision of protein quantification. Furthermore, implementation of a simple sample batch-running strategy enabled good chromatographic alignment for hundreds of samples over multiple batches. Finally, we applied the sample batch-running strategy and compared quantification by LC-MS ion intensity to spectral counting and found that quantification by LC-MS ion intensity was more accurate and precise. In summary, these results demonstrate that chromatographic alignment is important for precise and accurate protein quantification based on LC-MS ion intensity and accordingly we present a simple sample re-ordering strategy to facilitate improved alignment. These findings are not only relevant to label-free quantification using Progenesis QI but may be useful to the wide range of MS-based quantification strategies that rely on chromatographic alignment.

bioinformatics

Managing Uncertainty in Metabolic Network Structure and Improving Predictions Using EnsembleFBA

Genome-scale metabolic network reconstructions (GENREs) are repositories of knowledge about the metabolic processes that occur in an organism. GENREs have been used to discover and interpret metabolic functions, and to engineer novel network structures. A major barrier preventing more widespread use of GENREs, particularly to study non-model organisms, is the extensive time required to produce a high-quality GENRE. Many automated approaches have been developed which reduce this time requirement, but automatically-reconstructed draft GENREs still require curation before useful predictions can be made. We present a novel ensemble approach to the analysis of GENREs which improves the predictive capabilities of draft GENREs and is compatible with many automated reconstruction approaches. We refer to this new approach as Ensemble Flux Balance Analysis (EnsembleFBA). We validate EnsembleFBA by predicting growth and gene essentiality in the model organism Pseudomonas aeruginosa UCBPP-PA14. We demonstrate how EnsembleFBA can be included in a systems biology workflow by predicting essential genes in six Streptococcus species and mapping the essential genes to small molecule ligands from DrugBank. We found that some metabolic subsystems contribute disproportionately to the set of predicted essential reactions in a way that is unique to each Streptococcus species. These species-specific network structures lead to species-specific outcomes from small molecule interactions. Through these analyses of P. aeruginosa and six Streptococci, we show that ensembles increase the quality of predictions without drastically increasing reconstruction time, thus making GENRE approaches more practical for applications which require predictions for many non-model organisms. All of our functions and accompanying example code are available in an open online repository.\n\nAuthor SummaryMetabolism is the driving force behind all biological activity. Genome-scale metabolic network reconstructions (GENREs) are representations of metabolic systems that can be analyzed mathematically to make predictions about how a biochemical system will behave as well as to design biochemical systems with new properties. GENREs have traditionally been reconstructed manually, which can require extensive time and effort. Recent software solutions automate the process (drastically reducing the required effort) but the resulting GENREs are of lower quality and produce less reliable predictions than the manually-curated versions. We present a novel method (\"EnsembleFBA\") which overcomes uncertainties involved in automated reconstruction by pooling many different draft GENREs together into an ensemble. We tested EnsembleFBA by predicting the growth and essential genes of the common pathogen Pseudomonas aeruginosa. We found that when predicting growth or essential genes, ensembles of GENREs achieved much better precision or captured many more essential genes than any of the individual GENREs within the ensemble. By improving the predictions that can be made with automatically-generated GENREs, we open the door to studying systems which would otherwise be infeasible.

Systems Biology