bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Predicting growth conditions from internal metabolic fluxes in an in-silico model of E. coli

A widely studied problem in systems biology is to predict bacterial phenotype from growth conditions, using mechanistic models such as flux balance analysis (FBA). However, the inverse prediction of growth conditions from phenotype is rarely considered. Here we develop a computational framework to carry out this inverse prediction on a computational model of bacterial metabolism. We use FBA to calculate bacterial phenotypes from growth conditions in E. coli, and then we assess how accurately we can predict the original growth conditions from the phenotypes. Prediction is carried out via regularized multinomial regression. Our analysis provides several important physiological and statistical insights. First, we show that by analyzing metabolic end products we can consistently predict growth conditions. Second, prediction is reliable even in the presence of small amounts of impurities. Third, flux through a relatively small number of reactions per growth source (~10) is sufficient for accurate prediction. Fourth, combining the predictions from two separate models, one trained only on carbon sources and one only on nitrogen sources, performs better than models trained to perform joint prediction. Finally, that separate predictions perform better than a more sophisticated joint prediction scheme suggests that carbon and nitrogen utilization pathways, despite jointly affecting cellular growth, may be fairly decoupled in terms of their dependence on specific assortments of molecular precursors.

Systems Biology

RNA-Seq and Protein Mass Spectrometry in Microdissected Kidney Tubules Reveal Signaling Processes that Initiate Lithium-Induced Diabetes Insipidus

ABSTRACT1Lithium salts, used for treatment of bipolar disorder, frequently induce nephrogenic diabetes insipidus (NDI), limiting therapeutic success. NDI is associated with loss of expression of the molecular water channel, aquaporin-2, in the renal collecting duct (CD). Here, we use the methods of systems biology in a well-established rat model of lithium-induced NDI to identify signaling pathways activated at the onset of polyuria. Using single-tubule RNA-Seq, full transcriptomes were determined in microdissected cortical CDs of rats 72 hrs after initiation of lithium chloride (LiCl) administration (vs. time-controls without LiCl). Transcriptome-wide changes in mRNA abundances were mapped to gene sets associated with curated canonical signaling pathways, showing evidence for activation of NF-{kappa}B signaling with induction of genes coding for multiple chemokines as well as most components of the Major Histocompatibility Complex (MHC) Class I antigen-presenting complex. Administration of antiinflammatory doses of dexamethasone to LiCl-treated rats countered the loss of aquaporin-2 protein. RNA-Seq also confirmed prior evidence of a shift from quiescence into the cell cycle with arrest. Time course studies demonstrated an early (12 hrs) increase in multiple immediate early genes including several transcription factors. Protein mass spectrometry in microdissected cortical CDs provided corroborative evidence but also identified decreased abundance of several anti-oxidant proteins. Integration of new data with prior data about lithium effects at a molecular level leads to a signaling model in which lithium increases ERK activation leading to induction of NF-{kappa}B signaling and an inflammatory-like response that represses Aqp2 gene transcription.

systems biology

Reproducible model development in the Cardiac Electrophysiology Web Lab

The modelling of the electrophysiology of cardiac cells is one of the most mature areas of systems biology. This extended concentration of research effort brings with it new challenges, foremost among which is that of choosing which of these models is most suitable for addressing a particular scientific question. In a previous paper, we presented our initial work in developing an online resource for the characterisation and comparison of electrophysiological cell models in a wide range of experimental scenarios. In that work, we described how we had developed a novel protocol language that allowed us to separate the details of the mathematical model (the majority of cardiac cell models take the form of ordinary differential equations) from the experimental protocol being simulated. We developed a fully-open online repository (which we termed the Cardiac Electrophysiology Web Lab) which allows users to store and compare the results of applying the same experimental protocol to competing models. In the current paper we describe the most recent and planned extensions of this work, focused on supporting the process of model building from experimental data. We outline the necessary work to develop a machine-readable language to describe the process of inferring parameters from wet lab datasets, and illustrate our approach through a detailed example of fitting a model of the hERG channel using experimental data. We conclude by discussing the future challenges in making further progress in this domain towards our goal of facilitating a fully reproducible approach to the development of cardiac cell models.

systems biology

SBpipe: a collection of pipelines for automating repetitive simulation and analysis tasks

Background: The rapid growth of the number of mathematical models in Systems Biology fostered the development of many tools to simulate and analyse them. The reliability and precision of these tasks often depend on multiple repetitions and they can be optimised if executed as pipelines. In addition, new formal analyses can be performed on these repeat sequences, revealing important insights about the accuracy of model predictions.\n\nResults: Here we introduce SBpipe, an open source software tool for automating repetitive tasks in model building and simulation. Using basic configuration files, SBpipe builds a sequence of repeated model simulations or parameter estimations, performs analyses from this generated sequence, and finally generates a LaTeX/PDF report. The parameter estimation pipeline offers analyses of parameter profile likelihood and parameter correlation using samples from the computed estimates. Specific pipelines for scanning of one or two model parameters at the same time are also provided. Pipelines can run on multicore computers, Sun Grid Engine (SGE), or Load Sharing Facility (LSF) clusters, speeding up the processes of model building and simulation. SBpipe can execute models implemented in Copasi, Python or coded in any other programming language using Python as a wrapper module. Future support for other software simulators can be dynamically added without affecting the current implementation.\n\nConclusions: SBpipe allows users to automatically repeat the tasks of model simulation and parameter estimation, and extract robustness information from these repeat sequences in a solid and consistent manner, facilitating model development and analysis. The source code and documentation of this project are freely available at the web site: https://pdp10.github.io/sbpipe/.

systems biology

WASABI: a dynamic iterative framework for gene regulatory network inference

Inference of gene regulatory networks from gene expression data has been a long-standing and notoriously difficult task in systems biology. Recently, single-cell transcriptomic data have been massively used for gene regulatory network inference, with both successes and limitations. In the present work we propose an iterative algorithm called WASABI, dedicated to inferring a causal dynamical network from time-stamped single-cell data, which tackles some of the limitations associated with current approaches. We first introduce the concept of waves, which posits that the information provided by an external stimulus will affect genes one-by-one through a cascade, like waves spreading through a network. This concept allows us to infer the network one gene at a time, after genes have been ordered regarding their time of regulation. We then demonstrate the ability of WASABI to correctly infer small networks, which have been simulated in silico using a mechanistic model consisting of coupled piecewise-deterministic Markov processes for the proper description of gene expression at the single-cell level. We finally apply WASABI on in vitro generated data on an avian model of erythroid differentiation. The structure of the resulting gene regulatory network sheds a fascinating new light on the molecular mechanisms controlling this process. In particular, we find no evidence for hub genes and a much more distributed network structure than expected. Interestingly, we find that a majority of genes are under the direct control of the differentiation-inducing stimulus. In conclusion, WASABI is a versatile algorithm which should help biologists to fully exploit the power of time-stamped single-cell data.

systems biology

Network Architecture and Mutational Sensitivity of the C. elegans Metabolome

A fundamental issue in evolutionary systems biology is understanding the relationship between the topological architecture of a biological network, such as a metabolic network, and the evolution of the network. The rate at which an element in a metabolic network accumulates genetic variation via new mutations depends on both the size of the mutational target it presents and its robustness to mutational perturbation. Quantifying the relationship between topological properties of network elements and the mutability of those elements will facilitate understanding the variation in and evolution of networks at the level of populations and higher taxa.\n\nWe report an investigation into the relationship between two topological properties of 29 metabolites in the C. elegans metabolic network and the sensitivity of those metabolites to the cumulative effects of spontaneous mutation. The relationship between several measures of network centrality and sensitivity to mutation is weak, but point estimates of the correlation between network centrality and mutational variance are positive, with only one exception. There is a marginally significant correlation between core number and mutational heritability. There is a small but significant negative correlation between the shortest path length between a pair of metabolites and the mutational correlation between those metabolites.\n\nPositive association between the centrality of a metabolite and its mutational heritability is consistent with centrally-positioned metabolites presenting a larger mutational target than peripheral ones, and is inconsistent with centrality conferring mutational robustness, at least in toto. The weakness of the correlation between shortest path length and the mutational correlation between pairs of metabolites suggests that network locality is an important but not overwhelming factor governing mutational pleiotropy. These findings provide necessary background against which the effects of other evolutionary forces, most importantly natural selection, can be interpreted.

systems biology

Supervised learning on synthetic data for reverse engineering gene regulatory networks from experimental time-series

The reconstruction of gene regulatory networks from time resolved gene expression measurements is a key challenge in systems biology with applications in health and disease. While the most popular network inference methods are based on unsupervised learning approaches, supervised learning methods have proven their potential for superior reconstruction performance. However, obtaining the appropriate volume of informative training data constitutes a key limitation for the success of such methods.\n\nHere, we introduce a supervised learning approach to detect gene-gene regulation based on exclusively synthetic training data, termed surrogate learning, and show its performance for synthetic and experimental time-series. We systematically investigate different simulation configurations of biologically representative time-series of transcripts and augmentation of the data with a measurement model. We compare the resulting synthetic datasets to experimental data, and evaluate classifiers trained on them for detection of gene-gene regulation from experimental time-series. For classifiers, we consider hybrid convolutional recurrent neural networks, random forests and logistic regression, and evaluate the reconstruction performance of different simulation settings, data pre-processing and classifiers.\n\nWhen training and test time-courses are generated from the same distribution, we find that the largest tested neural network architecture achieves the best performance of 0.448 {+/-} 0.047 (mean {+/-} std) in maximally achievable F1 score over all datasets outperforming random forests by 32.4 % {+/-} 14 % (mean {+/-} std). Reconstruction performance is sensitive to discrepancies between synthetic training and test data, highlighting the importance of matching training and test data domains. For an experimental gene expression dataset from E.coli, we find that training data generated with measurement model, multi-gene perturbations, but without data standardization is best suited for training classifiers for network reconstruction from the experimental test data. We further demonstrate superiority to multiple unsupervised, state-of-the-art methods for networks comprising 20 genes of the experimental data from E.coli (average AUPR best supervised = 0.22 vs best unsupervised = 0.07).\n\nWe expect the proposed surrogate learning approach to be broadly applicable. It alleviates the requirement for large, difficult to attain volumes of experimental training data and instead relies on easily accessible synthetic data. Successful application for new experimental conditions and other data types is only limited by the automatable and scalable process of designing simulations which generate suitable synthetic data.

systems biology

Qualitative modeling of signaling networks in replicative senescence by selecting optimal node and arc sets

Signaling networks are an important tool of modern systems biology and drug development. Here, we present a new methodology to qualitatively model signaling networks by combining experimental data and prior knowledge about protein connectivity. Unlike other methods, our approach does not focus solely on selecting which reactions are involved but also on whether a protein is present. This allows the user to model more complicated experiments and incorporate more knowledge into the model. To demonstrate the capabilities of our method we compared the signaling networks of young and replicative senescent human primary HFL-1 fibroblasts, whose differences are expected to be due mainly to changes in the expression of the proteins rather than the reactions involved. The resulting networks indicate that, compared to young cells, aged cells are not as responsive to insulin stimulation and activate pathways that establish and maintain senescence.\n\nAuthor summaryCells have developed a complex network of biochemical reactions to monitor their environment and react to changes. Although multiple pathways, tuned to identify specific stimuli, have been discovered, it is generally understood that the signaling process typically involves multiple pathways and is context depended. Consequently, reconstructing the signaling network utilized by cells at any given moment is not a trivial task. In this article, we report on a novel logic-based method for identifying signaling network by combining experimental data with prior knowledge about the connectivity of the involved proteins. Unlike other methods proposed so far, our method uses data to evaluate the presence or absence of reactions and proteins alike. We reconstructed and compared the signaling network of human primary HFL-1 fibroblasts as they undergo replicative senescence in the presence of 6 different stimuli. The resulting networks indicate that, compared to young cells, senescent cells are not responsive to insulin stimulation and activate pathways that are known to establish and maintain senescence.

systems biology

Coal-Miner: A Coalescent-Based Method For GWA Studies Of Quantitative Traits With Complex Evolutionary Origins

Association mapping (AM) methods are used in genome-wide association (GWA) studies to test for statistically significant associations between genotypic and phenotypic data. The genotypic and phenotypic data share common evolutionary origins - namely, the evolutionary history of sampled organisms - introducing covariance which must be distinguished from the covariance due to biological function that is of primary interest in GWA studies. A variety of methods have been introduced to perform AM while accounting for sample relatedness. However, the state of the art predominantly utilizes the simplifying assumption that sample relatedness is effectively fixed across the genome. In contrast, population genetic theory and empirical studies have shown that sample relatedness can vary greatly across different loci within a genome; this phenomena - referred to as local genealogical variation - is commonly encountered in many genomic datasets. New AM methods are needed to better account for local variation in sample relatedness within genomes.\n\nWe address this gap by introducing Coal-Miner, a new statistical AM method. The Coal-Miner algorithm takes the form of a methodological pipeline. The initial stages of Coal-Miner seek to detect candidate loci, or loci which contain putatively causal markers. Subsequent stages of Coal-Miner perform test for association using a linear mixed model with multiple effects which account for sample relatedness locally within candidate loci and globally across the entire genome.\n\nUsing synthetic and empirical datasets, we compare the statistical power and type I error control of Coal-Miner against state-of-theart AM methods. The simulation conditions reflect a variety of genomic architectures for complex traits and incorporate a range of evolutionary scenarios, each with different evolutionary processes that can generate local genealogical variation. The empirical benchmarks include a large-scale dataset that appeared in a recent high-profile publication. Across the datasets in our study, we find that Coal-Miner consistently offers comparable or typically better statistical power and type I error control compared to the state-of-art methods.\n\nCCS CONCEPTSApplied computing [->] Computational genomics; Computational biology; Molecular sequence analysis; Molecular evolution; Computational genomics; Systems biology; Bioinformatics; Population genetics;\n\nACM Reference formatHussein A. Hejase, Natalie Vande Pol, Gregory M. Bonito, Patrick P. Edger, and Kevin J. Liu. 2017. Coal-Miner: a coalescent-based method for GWA studies of quantitative traits with complex evolutionary origins. In Proceedings of ACM BCB, Boston, MA, 2017 (BCB), 10 pages. DOI: 10.475/123 4

bioinformatics

Enter the matrix: Interpreting unsupervised feature learning with matrix decomposition to discover hidden knowledge in high-throughput omics data

Omics data contains signal from the molecular, physical, and kinetic inter- and intra-cellular interactions that control biological systems. Matrix factorization techniques can reveal low-dimensional structure from high-dimensional data that reflect these interactions. These techniques can uncover new biological knowledge from diverse high-throughput omics data in topics ranging from pathway discovery to time course analysis. We review exemplary applications of matrix factorization for systems-level analyses. We discuss appropriate application of these methods, their limitations, and focus on analysis of results to facilitate optimal biological interpretation. The inference of biologically relevant features with matrix factorization enables discovery from high-throughput data beyond the limits of current biological knowledge--answering questions from high-dimensional data that we have not yet thought to ask.

systems biology

In Silico Processing of the Complete CRISPR-Cas Spacer Space for Identification of PAM Sequences

Despite extensive exploration of the diversity of CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) systems, biological applications have been mostly confined to Class 2 systems, specifically the Cas9 and Cas12 (formerly Cpf1) single effector proteins. A key limitation of exploring and utilizing other CRISPR-Cas systems with unique functionalities, particularly Class I types and their multi-protein effector complex, is the knowledge of the systems protospacer adjacent motif (PAM) sequence identity. In this work, we developed a systematic pipeline, named CASPERpam, that enables us to comprehensively assess the PAM sequences of all the available CRISPR-Cas systems in the NCBI database of bacterial genomes. The CASPERpam analysis revealed that within the 30,389 assemblies previously screen for CRISPR arrays, there exists 26,364 spacers that match somewhere in the viral, bacterial, and plasmid databases of NCBI, using the constraints of 95% sequence identity and 95% sequence coverage for blast hits. When grouping these results by species, we were able to identify putative PAM sequences for 1,049 among 1,493 unique species. The remaining species either have insufficient data or an undetermined result from the analysis. Finally, we were able to infer certain design principles that are relevant for understanding PAM diversity and a baseline for further experimental studies including PAM assays. We envision CASPERpam is a useful bioinformatic tool for understanding and harnessing the diversity of CRISPR systems.

bioinformatics

Rapid quantitative pharmacodynamic imaging with Bayesian estimation

We recently described rapid quantitative pharmacodynamic imaging, a novel method for estimating sensitivity of a biological system to a drug. We tested its accuracy in simulated biological signals with varying receptor sensitivity and varying levels of random noise, and presented initial proof-of-concept data from functional MRI (fMRI) studies in primate brain. However, the initial simulation testing used a simple iterative approach to estimate pharmacokinetic-pharmacodynamic (PKPD) parameters, an approach that was computationally efficient but returned parameters only from a small, discrete set of values chosen a priori.\n\nHere we revisit the simulation testing using a Bayesian method to estimate the PKPD parameters. This improved accuracy compared to our previous method, and noise without intentional signal was never interpreted as signal. We also reanalyze the fMRI proof-of-concept data. The success with the simulated data, and with the limited fMRI data, is a necessary first step toward further testing of rapid quantitative pharmacodynamic imaging.

Pharmacology and Toxicology

MAMMOTh: a new database for curated MAthematical Models of bioMOlecular sysTems

Living systems have a complex hierarchical organization that can be viewed as a set of dynamically interacting subsystems. Thus, to simulate the internal nature and dynamics of the whole biological system we should use the iterative way for a model reconstruction, which is a consistent composition and combination of its elementary subsystems. In accordance with this bottom-up approach, we have developed MAMMOTh (MAthematical Models of bioMOlecular sysTems) database that allows integrating manually curated mathematical models of biomolecular systems, which are fit to the experimental data. The database entries are organized as building blocks in a way that the model parts can be used in different combinations to describe systems with higher organizational level (metabolic pathways and/or transcription regulatory networks). The database supports export of single model or their combinations in SBML or Mathematica standards. The database currently contains more than 100 mathematical models for Escherichia coli elementary subsystems (enzymatic reactions and gene expression regulatory processes) that can be combined in at least 5100 complex/sophisticated models concerning such biological processes as: de novo nucleotide biosynthesis, aerobic/anaerobic respiration, and nitrate/nitrite utilization in E. coli. All current models are functionally interconnected and sufficiently complement public model resources.\n\nDatabase URL: http://mammoth.biomodelsgroup.ru

Systems Biology

Towards physical principles of biological evolution

Biological systems reach organizational complexity that far exceeds the complexity of any known inanimate objects. Biological entities undoubtedly obey the laws of quantum physics and statistical mechanics. However, is modern physics sufficient to adequately describe, model and explain the evolution of biological complexity? Detailed parallels have been drawn between statistical thermodynamics and the population-genetic theory of biological evolution. Based on these parallels, we outline new perspectives on biological innovation and major transitions in evolution, and introduce a biological equivalent of thermodynamic potential that reflects the innovation propensity of an evolving population. Deep analogies have been suggested to also exist between the properties of biological entities and processes, and those of frustrated states in physics, such as glasses. Such systems are characterized by frustration whereby local state with minimal free energy conflict with the global minimum, resulting in \"emergent phenomena\". We extend such analogies by examining frustration-type phenomena, such as conflicts between different levels of selection, in biological evolution. These frustration effects appear to drive the evolution of biological complexity. We further address evolution in multidimensional fitness landscapes from the point of view of percolation theory and suggest that percolation at level above the critical threshold dictates the tree-like evolution of complex organisms. Taken together, these multiple connections between fundamental processes in physics and biology imply that construction of a meaningful physical theory of biological evolution might not be a futile effort. However, it is unrealistic to expect that such a theory can be created in one scoop; if it ever comes to being, this can only happen through integration of multiple physical models of evolutionary processes. Furthermore, the existing framework of theoretical physics is unlikely to suffice for adequate modeling of the biological level of complexity, and new developments within physics itself are likely to be required.

evolutionary biology

A dynamical model of TCRβ gene recombination: Coupling the initiation of Dβ-Jβ rearrangement to TCRβ allelic exclusion

One paradigm of random monoallelic gene expression is that of T-cell receptor (TCR){beta} allelic exclusion in T lymphocytes. However, the dynamics that sustain asymmetric choice in TCR{beta} dual allele usage and the production of TCR{beta} monoallelic expressing T-cells remain poorly understood. Here, we develop a computational model to explore a scheme of TCR{beta} allelic exclusion based on the stochastic initiation of DNA rearrangement [V(D)J recombination] at homologous alleles in T-cell progenitors, and thus account for the genotypic profiles typically associated with allelic exclusion in differentiated T-cells. Disturbances in these dynamics at the level of an individual allele have limited consequences on these pro1les, robust feature of the system that is underscored by our simulations. Our study predicts a biological system in which locus-specific, prime epigenetic allelic activation effects set the stage to both optimize the production of TCR{beta} allelically excluded T-cells and curtail the emergence of their allelically included counterparts.

immunology

PorSignDB: a database of in vivo perturbation signatures for dissecting clinical outcome of PCV2 infection

Porcine Circovirus Type 2 (PCV2) is a pathogen that has the ability to cause often devastating disease manifestations in pig populations with major economic implications. How PCV2 establishes subclinical persistence and why certain individuals progress to lethal lymphoid depletion remain to be elucidated. Here we present PorSignDB, a gene signature database describing in vivo porcine tissue physiology that we generated from a large compendium of in vivo transcriptional profiles and that we subsequently leveraged for deciphering the distinct physiological states underlying PCV2-affected lymph nodes. This systems biology approach indicated that subclinical PCV2 infections shut down the immune system. A robust signature of PCV2 disease emphasized that immune activation is dysfunctional in subclinical infections, however, in contrast it is promoted in PCV2 patients with clinical manifestations. Functional genomics further uncovered IL-2 as a driver of PCV2-mediated disease and we identified STAT3 as a druggable PCV2 host factor candidate. Our systematic dissection of the mechanistic basis of PCV2 reveals that subclinical and clinical PCV2 display two diametrically opposed immunotranscriptomic recalibrations that represent distinct physiological states in vivo, which suggests a paradigm shift in this field. Finally, our PorSignDB signature database is publicly available as a community resource (http://www.vetvirology.ugent.be/PorSignDB/, included in Gene Sets from Community Contributors http://software.broadinstitute.org/gsea/msigdb/contributed_genesets.jsp) and provides systems biologists with a valuable tool for catalyzing studies of human and veterinary disease.\n\nAuthor SummaryPorcine Circovirus Type 2 (PCV2) is a small but economically important pathogen circulating endemically in pig populations. Although PCV2 causes mostly chronic subclinical infections, many individuals develop a lethal form of circoviral disease consisting of a collapse of lymphoid tissue. In order to provide a fresh look at how PCV2 reprograms host tissue, we created PorSignDB, a compendium of hundreds of transcriptomic gene-expression signatures derived from primary porcine tissue specimens of well over 1500 patients or lab animals. By leveraging PorSignDB on transcriptomic data of PCV2 patients, we uncover that subclinical PCV2 reprograms the host into a striking state of non-infection, which explains its failure to respond to an initial phase of circoviral presence. A PCV2 disease signature further demonstrates that the silenced immune system associated with subclinical PCV2 becomes fully activate in PCV2 patients, triggering severe circoviral disease. Further genomic and functional analysis demonstrate STAT3 as a druggable host factor and IL-2 as a disease driver. Together, this study demonstrates the mechanistic underpinnings of clinical outcome of PCV2 infections: subclinical and clinical PCV2 display two entirely opposing transcriptomic recalibrations of lymphoid tissue.

microbiology

A Multi-Stage Representation Of Cell Proliferation As A Markov Process

The stochastic simulation algorithm commonly known as Gillespies algorithm (originally derived for modelling well-mixed systems of chemical reactions) is now used ubiquitously in the modelling of biological processes in which stochastic effects play an important role. In well mixed scenarios at the sub-cellular level it is often reasonable to assume that times between successive reaction/interaction events are exponentially distributed and can be appropriately modelled as a Markov process and hence simulated by the Gillespie algorithm. However, Gillespies algorithm is routinely applied to model biological systems for which it was never intended. In particular, processes in which cell proliferation is important (e.g. embryonic development, cancer formation) should not be simulated naively using the Gillespie algorithm since the history-dependent nature of the cell cycle breaks the Markov process. The variance in experimentally measured cell cycle times is far less than in an exponential cell cycle time distribution with the same mean.\n\nHere we suggest a method of modelling the cell cycle that restores the memoryless property to the system and is therefore consistent with simulation via the Gillespie algorithm. By breaking the cell cycle into a number of independent exponentially distributed stages we can restore the Markov property at the same time as more accurately approximating the appropriate cell cycle time distributions. The consequences of our revised mathematical model are explored analytically as far as possible. We demonstrate the importance of employing the correct cell cycle time distribution by recapitulating the results from two models incorporating cellular proliferation (one spatial and one non-spatial) and demonstrating that changing the cell cycle time distribution makes quantitative and qualitative differences to the outcome of the models. Our adaptation will allow modellers and experimentalists alike to appropriately represent cellular proliferation - vital to the accurate modelling of many biological processes - whilst still being able to take advantage of the power and efficiency of the popular Gillespie algorithm.

cell biology

CyanoGate: A Golden Gate modular cloning suite for engineering cyanobacteria based on the plant MoClo syntax

Recent advances in synthetic biology research have been underpinned by an exponential increase in available genomic information and a proliferation of advanced DNA assembly tools. The adoption of plasmid vector assembly standards and parts libraries has greatly enhanced the reproducibility of research and exchange of parts between different labs and biological systems. However, a standardised Modular Cloning (MoClo) system is not yet available for cyanobacteria, which lag behind other prokaryotes in synthetic biology despite their huge potential in biotechnological applications. By building on the assembly library and syntax of the Plant Golden Gate MoClo kit, we have developed a versatile system called CyanoGate that unites cyanobacteria with plant and algal systems. We have generated a suite of parts and acceptor vectors for making i) marked/unmarked knock-outs or integrations using an integrative acceptor vector, and ii) transient multigene expression and repression systems using known and novel replicative vectors. We have tested and compared the CyanoGate system in the established model cyanobacterium Synechocystis sp. PCC 6803 and the more recently described fast-growing strain Synechococcus elongatus UTEX 2973. The system is publicly available and can be readily expanded to accommodate other standardised MoClo parts.

synthetic biology