bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Widespread Historical Contingency in Influenza Viruses

In systems biology and genomics, epistasis characterizes the impact that a substitution at a particular location in a genome can have on a substitution at another location. This phenomenon is often implicated in the evolution of drug resistance or to explain why particular disease-causing mutations do not have the same outcome in all individuals. Hence, uncovering these mutations and their locations in a genome is a central question in biology. However, epistasis is notoriously difficult to uncover, especially in fast-evolving organisms. Here, we present a novel statistical approach that replies on a model developed in ecology and that we adapt to analyze genetic data in fast-evolving systems such as the influenza A virus. We validate the approach using a two-pronged strategy: extensive simulations demonstrate a low-to-moderate sensitivity with excellent specificity and precision, while analyses of experimentally-validated data recover known interactions, including in a eukaryotic system. We further evaluate the ability of our approach to detect correlated evolution during antigenic shifts or at the emergence of drug resistance. We show that in all cases, correlated evolution is prevalent in influenza A viruses, involving many pairs of sites linked together in chains, a hallmark of historical contingency. Strikingly, interacting sites are separated by large physical distances, which entails either long-range conformational changes or functional tradeoffs, for which we find support with the emergence of drug resistance. Our work paves a new way for the unbiased detection of epistasis in a wide range of organisms by performing whole-genome scans.

Evolutionary Biology

DynOmics to identify delays and co-expression patterns across time course experiments

Dynamic changes in biological systems can be captured by measuring molecular expression from different levels (e.g., genes and proteins) across time. Integration of such data aims to identify molecules that show similar expression changes over time; such molecules may be co-regulated and thus involved in similar biological processes. Combining data sources presents a systematic approach to study molecular behaviour. It can compensate for missing data in one source, and can reduce false positives when multiple sources highlight the same pathways. However, integrative approaches must accommodate the challenges inherent in omics data, including high-dimensionality, noise, and timing differences in expression. As current methods for identification of co-expression cannot cope with this level of complexity, we developed a novel algorithm called DynOmics. DynOmics is based on the fast Fourier transform, from which the difference in expression initiation between trajectories can be estimated. This delay can then be used to realign the trajectories and identify those which show a high degree of correlation. Through extensive simulations, we demonstrate that DynOmics is efficient and accurate compared to existing approaches. We consider two case studies highlighting its application, identifying regulatory relationships across omics data within an organism and for comparative gene expression analysis across organisms.

Bioinformatics

KymoButler: A deep learning software for automated kymograph tracing and analysis

Kymographs are graphical representations of spatial position over time, which are often used in biology to visualise the motion of fluorescent particles, molecules, vesicles, or organelles moving along a predictable path. Although in kymographs tracks of individual particles are qualitatively easily distinguished, their automated quantitative analysis is much more challenging. Kymographs often exhibit low signal-to-noise-ratios (SNRs), and available tools that automate their analysis usually require manual supervision. Here we developed KymoButler, a Deep Learning-based software to automatically track dynamic processes in kymographs. We demonstrate that KymoButler performs as well as expert manual data analysis on kymographs with complex particle trajectories from a variety of different biological systems. The software was packaged in a web-based \"one-click\" application for use by the wider scientific community. Our approach significantly speeds up data analysis, avoids unconscious bias, and represents another step towards the widespread adaptation of Machine Learning techniques in biological data analysis.

cell biology

The application of zeta diversity as a continuous measure of compositional change in ecology

Zeta diversity provides the average number of shared species across n sites (or shared operational taxonomic units (OTUs) across n cases). It quantifies the variation in species composition of multiple assemblages in space and time to capture the contribution of the full suite of narrow, intermediate and wide-ranging species to biotic heterogeneity. Zeta diversity was proposed for measuring compositional turnover in plant and animal assemblages, but is equally relevant for application to any biological system that can be characterised by a row by column incidence matrix. Here we illustrate the application of zeta diversity to explore compositional change in empirical data, and how observed patterns may be interpreted. We use 10 datasets from a broad range of scales and levels of biological organisation - from DNA molecules to microbes, plants and birds - including one of the original data sets used by R.H. Whittaker in the 1960s to express compositional change and distance decay using beta diversity. The applications show (i) how different sampling schemes used during the calculation of zeta diversity may be appropriate for different data types and ecological questions, (ii) how higher orders of zeta may in some cases better detect shifts, transitions or periodicity, and importantly (iii) the relative roles of rare versus common species in driving patterns of compositional change. By exploring the application of zeta diversity across this broad range of contexts, our goal is to demonstrate its value as a tool for understanding continuous biodiversity turnover and as a metric for filling the empirical gap that exists on spatial or temporal change in compositional diversity.

ecology

Harmonizing semantic annotations for computational models in biology

Life science researchers use computational models to articulate and test hypotheses about the behavior of biological systems. Semantic annotation is a critical component for enhancing the interoperability and reusability of such models as well as for the integration of the data needed for model parameterization and validation. Encoded as machine-readable links to knowledge resource terms, semantic annotations describe the computational or biological meaning of what models and data represent. These annotations help researchers find and repurpose models, accelerate model composition, and enable knowledge integration across model repositories and experimental data stores. However, realizing the potential benefits of semantic annotation requires the development of model annotation standards that adhere to a community-based annotation protocol. Without such standards, tool developers must account for a variety of annotation formats and approaches, a situation that can become prohibitively cumbersome and which can defeat the purpose of linking model elements to controlled knowledge resource terms. Currently, no consensus protocol for semantic annotation exists among the larger biological modeling community. Here, we report on the landscape of current semantic annotation practices among the COmputational Modeling in BIology NEtwork (COMBINE) community and provide a set of recommendations for building a consensus approach to semantic annotation.

scientific communication and education

Identifying Mechanisms of Regulation to Model Carbon Flux During Heat Stress And Generate Testable Hypotheses

Understanding biological response to stimuli requires identifying mechanisms that coordinate changes across pathways. One of the promises of multi-omics studies is achieving this level of insight by simultaneously identifying different levels of regulation. However, computational approaches to integrate multiple types of data are lacking. An effective systems biology approach would be one that uses statistical methods to detect signatures of relevant network motifs and then builds metabolic circuits from these components to model shifting regulatory dynamics. For example, transcriptome and metabolome data complement one another in terms of their ability to describe shifts in physiology. Here, we extend a previously described method used to identify single nucleotide polymorphism (SNPs) associated with metabolic changes (Gieger et al., 2008). We apply this strategy to link changes in sulfur, amino acid and lipid production under heat stress by relating ratios of compounds to potential precursors and regulators. This approach provides integration of multi-omics data to link previously described, discrete units of regulation into functional pathways and hypothesizes novel biology relevant to the heat stress response.

genomics

Inferring protein-protein interaction networks from inter-protein sequence co-evolution

Interaction between proteins is a fundamental mechanism that underlies virtually all biological processes. Many important interactions are conserved across a large variety of species. The need to maintain interaction leads to a high degree of co-evolution between residues in the interface between partner proteins. The inference of protein-protein interaction networks from the rapidly growing sequence databases is one of the most formidable tasks in systems biology today. We propose here a novel approach based on the Direct-Coupling Analysis of the co-evolution between inter-protein residue pairs. We use ribosomal and trp operon proteins as test cases: For the small resp. large ribosomal subunit our approach predicts protein-interaction partners at a true-positive rate of 70% resp. 90% within the first 10 predictions, with areas of 0.69 resp. 0.81 under the ROC curves for all predictions. In the trp operon, it assigns the two largest interaction scores to the only two interactions experimentally known. On the level of residue interactions we show that for both the small and the large ribosomal subunit our approach predicts interacting residues in the system with a true positive rate of 60% and 85% in the first 20 predictions. We use artificial data to show that the performance of our approach depends crucially on the size of the joint multiple sequence alignments and analyze how many sequences would be necessary for a perfect prediction if the sequences were sampled from the same model that we use for prediction. Given the performance of our approach on the test data we speculate that it can be used to detect new interactions, especially in the light of the rapid growth of available sequence data.

Bioinformatics

HIVs Feedback Circuit Breaks The Fundamental Limit On Noise Suppression To Stabilize Fate

Diverse biological systems utilize gene-expression fluctuations ( noise) to drive lineage-commitment decisions1-5. However, once a commitment is made, noise becomes detrimental to reliable function6,7 and the mechanisms enabling post-commitment noise suppression are unclear. We used time-lapse imaging and mathematical modeling, and found that, after a noise-driven event, human immunodeficiency virus (HIV) strongly attenuated expression noise through a non-transcriptional negative-feedback circuit. Feedback is established by serial generation of RNAs from post-transcriptional splicing, creating a precursor-product relationship where proteins generated from spliced mRNAs auto-deplete their own precursor un-spliced mRNAs. Strikingly, precursor auto-depletion overcomes the theoretical limits on conventional noise suppression--minimizing noise far better than transcriptional auto-repression--and dramatically stabilizes commitment to the active-replication state. This auto-depletion feedback motif may efficiently suppress noise in other systems ranging from detained introns to non-sense mediated decay.

cell biology

Bioty: A cloud-based development toolkit for programming experiments and interactive applications with living cells

Recent advancements in life-science instrumentation and automation enable entirely new modes of human interaction with microbiological processes and corresponding applications for science and education through biology cloud labs. A critical barrier for remote life-science experimentation is the absence of suitable abstractions and interfaces for programming living matter. To this end we conceptualize a programming paradigm that provides stimulus control functions and sensor control functions for realtime manipulation of biological (physical) matter. Additionally, a simulation mode facilitates higher user throughput, program debugging, and biophysical modeling. To evaluate this paradigm, we implemented a JavaScript-based web toolkit, Bioty, that supports realtime interaction with swarms of phototactic Euglena cells hosted on a cloud lab. Studies with remote users demonstrate that individuals with little to no biology knowledge and intermediate programming knowledge were able to successfully create and use scientific applications and games. This work informs the design of programming environments for controlling living matter in general and lowers the access barriers to biology experimentation for professional and citizen scientists, learners, and the lay public.\n\nSignificance StatementBiology cloud labs are an emerging approach to lower access barriers to life-science experimentation. However, suitable programming approaches and user interfaces are lacking, especially ones that enable the interaction with the living matter itself - not just the control of equipment. Here we present and implement a corresponding programming paradigm for realtime interactive applications with remotely housed biological systems, and which is accessible and useful for scientists, programmers and lay people alike. Our user studies show that scientists and non-scientists are able to rapidly develop a variety of applications, such as interactive biophysics experiments and games. This paradigm has the potential to make first-hand experiences with biology accessible to all of society and to accelerate the rate of scientific discovery.

bioengineering

LIPEA: Lipid Pathway Enrichment Analysis

MotivationAnalyzing associations among multiple omic variables to infer mechanisms that meaningfully link them is a crucial step in systems biology. Gene Set Enrichment Analysis (GSEA) was conceived to pursue this aim in computational genomics, unveiling significant pathways associated to certain gene signatures under investigation. Lipidomics is a rapidly growing omic field, and absolute quantification of lipid abundance by shotgun mass spectrometry is generating high-throughput datasets that depict lipid metabolism in a plethora of conditions and organisms. In addition, high-throughput lipidomics represents a new important ally to develop personalized medicine approaches, investigate the causes and predict effective biomarkers in metabolic diseases, and not only.\n\nResultsHere, we present Lipid Pathway Enrichment Analysis (LIPEA), a web-tool for over-representation analysis of lipid signatures and detection of the biological pathways in which they are enriched. LIPEA is a new valid resource for biologists and physicians to mine pathways significantly associated to a set of lipids, helping them to discover whether common and collective mechanisms are hidden behind those lipids. LIPEA was extensively tested and we provide two examples where our system gave successfully results related with Major Depression Disease (MDD) and insulin re-sistance.\n\nAvailabilityThe tool is available as web platform at https://lipea.biotec.tu-dresden.de.

bioinformatics

Interaction pathways promote spliceosome module integration and network-level robustness to cascading effects

The functionality of distinct types of protein networks depends on the patterns of protein-protein interactions. A problem to solve is understanding the fragility of protein networks to predict system malfunctioning due to mutations and other errors. Spectral graph theory provides tools to understand the structural and dynamical properties of a system based on the mathematical properties of matrices associated with the networks. We combined two of such tools to explore the fragility to cascading effects of the network describing protein interactions within a key macromolecular complex, the spliceosome. Using S. cerevisiae as a model system we show that the spliceosome network has more indirect paths connecting proteins than random networks. Such multiplicity of paths may promote routes to cascading effects to propagate across the network. However, the modular network structure concentrates paths within modules, thus constraining the propagation of such cascading effects, as indicated by analytical results from the spectral graph theory and by numerical simulations of a minimal mathematical model parameterized with the spliceosome network. We hypothesize that the concentration of paths within modules favors robustness of the spliceosome against failure, but may lead to a higher vulnerability of functional subunits which may affect the temporal assembly of the spliceosome. Our results illustrate the utility of spectral graph theory for identifying fragile spots in biological systems and predicting their implications.

molecular biology

Tuning gene expression variability and multi-gene regulation by dynamic transcription factor control

Many natural transcription factors are regulated in a pulsatile fashion, but it remains unknown whether synthetic gene expression systems can benefit from such dynamic regulation. Using a fast-acting, light-responsive transcription factor in Saccharomyces cerevisiae, we show that dynamic pulsatile signals reduce cell-to-cell variability in gene expression. We then show that by encoding such signals into a single input, expression mean and variability can be precisely and independently tuned. Further, we construct a light-responsive promoter library and demonstrate how pulsatile signaling also enables graded multi-gene regulation at fixed expression ratios, despite differences in promoter dose-response characteristics. Pulsatile regulation can thus lead to highly beneficial functional behaviors in synthetic biological systems, which previously required laborious optimization of genetic parts or complex construction of synthetic gene networks.

synthetic biology

A tunable dual-input system for on-demand dynamic geneexpression regulation.

Cellular systems have evolved numerous mechanisms to finely control signalling pathway activation and properly respond to changing environmental stimuli. This is underpinned by dynamic spatiotemporal patterns of gene expression. Indeed, in addition to gene transcription and translation regulation, modulation of protein levels, dynamics and localization are also essential checkpoints that govern cell functions. The introduction of tetracycline-inducible promoters has allowed gene expression control using orthogonal small molecules, facilitating rapid and reversible manipulation to study gene function in biological systems. However, differing protein stabilities means this solely transcriptional regulation is insufficient to allow precise ON-OFF dynamics, thus hindering generation of temporal profiles of protein levels seen in vivo. We developed an improved Tet-On based system augmented with conditional destabilising elements at the post-translational level that permits simultaneous control of gene expression and protein stability. Integrating these properties to control expression of a fluorescent protein in mouse Embryonic Stem Cells (mESCs), we found that adding protein stability control allows faster response times to changes in small molecules, fully tunable and enhanced dynamic range, and vastly improved microfluidic-based in-silico feedback control of gene expression. Finally, we highlight the effectiveness of our dual-input system to finely modulate levels of signalling pathway components in stem cells.

synthetic biology

A Multi-layer, Self-aligning Hydrogel Micro-molding Process Offering a Fabrication Route to Perfusable 3D In-Vitro Microvasculature

The in-vitro fabrication of hierarchical biological systems such as human vasculature, which are made up of two or more cell types with intricate co-culture architectures, is by far one of the most complicated challenges that tissue engineers have faced. Here, we introduce a versatile method to create multi-layered, cell-laden hydrogel microstructures with coaxial geometries and heterogeneous mechanical and biological properties. The technique can be used to build in-vitro vascular networks that are fully embedded in hydrogels of physiologically realistic mechanical stiffness. Our technique produces free-standing 3D structures, eliminating rigid polymeric surfaces from the vicinity of cells and allowing layers of multiple cell types to be defined with tailored extracellular matrix (ECM) composition and stiffness, and in direct contact with each other. We demonstrate co-axial geometries with diameters ranging from 200-2000 m and layer thicknesses as small as 50-200 {micro}m in agarose- collagen (AC) composite hydrogels. Coaxial geometries with such fine feature sizes are beyond the capabilities of most bioprinting techniques. A potential application of such a structure is to simulate vascular networks in the brain with endothelial cells surrounded by multiple layers of pericytes and other glial cells. For this purpose, the composition and mechanical properties of the composite AC hydrogels have been optimized for cell viability and biological performance of endothelial and glial cell types in both 2D and 3D culture modes. Multi-layered vascular constructs with an endothelial layer surrounded by layers of glial cells have been fabricated. This prototype in-vitro model resembles vascular geometries and opens the way for complex multi-luminal blood vessels to be fabricated.

bioengineering

An Information Geometry approach to model topological trajectories and Gene Expression Radius from UMAP geometry.

Understanding the relationship between gene expression dynamics and cellular identity remains a central challenge in single cell biology. Here, we introduce a novel computational and mathematical framework that integrates information geometry, fuzzy topology, and UMAP analysis to model gene expression landscapes derived from single cell RNA sequencing data. We formalize gene expression data as a fuzzy topological space, where interactions between expression points are governed by probabilistic distributions inspired by manifold learning approaches such as UMAP. Within this framework, we define an information geometric structure through a Fisher metric induced by these distributions, enabling the computation of geodesic trajectories that capture cellular differentiation processes. A key contribution of this work is the derivation of analytical conditions, expressed as expression radius formulas, that characterize local neighborhoods in gene expression space. These conditions allow for the identification of genes associated with stem cell states and predictions in transitional cell types in future work. Application of the proposed framework to single cell datasets reveals biologically meaningful gene sets enriched in key regulatory pathways and transcription factors, demonstrating the capacity of our approach to uncover latent structure in complex gene expression data. Our results suggest that integrating differential geometry with statistical learning theory offers a powerful paradigm for modeling genotype and phenotype relationships and cellular state transitions, with potential implications for precision medicine and systems biology.

bioinformatics

Genetic networks of the oxytocin system in the human brain: A gene expression and large-scale fMRI meta-analysis study

Oxytocin is a neuropeptide involved in animal and human reproductive and social behaviour, with potential implications for a range of psychiatric disorders. However, the therapeutic potential of oxytocin in mental health care suggested by animal research has not been successfully translated into clinical practice, partly due to a poor understanding of the expression and distribution of the oxytocin signaling pathway in the human brain, and its complex interactions with other biological systems. Among the genes involved in the oxytocin signaling pathway, three genes have been frequently implicated in human social behavior: OXT (structural gene for oxytocin), OXTR (oxytocin receptor), and CD38 (central oxytocin secretion). We characterized the distribution of the OXT, OXTR, and CD38 mRNA across the brain, identified putative gene pathway interactions by comparing gene expression patterns across 29131 genes, and assessed associations between gene expression patterns and cognitive states via large-scale fMRI meta-analysis. In line with the animal literature, oxytocin pathway gene expression was enriched in central, temporal, and olfactory regions. Across the brain, there was high co-expression of the oxytocin pathway genes with both dopaminergic (DRD2) and muscarinic acetylcholine (CHRM4) genes, reflecting an anatomical basis for critical gene pathway interactions. Finally, fMRI meta-analysis revealed that oxytocin pathway maps correspond with motivation and emotion processing, demonstrating the value of probing gene expression maps to identify brain functional targets for future pharmacological trials.

neuroscience

CONFESS: Fluorescence-based single-cell ordering in R

Modern high-throughput single-cell technologies facilitate the efficient processing of hundreds of individual cells to comprehensively study their morphological and genomic heterogeneity. Fluidigms C1 Auto Prep system isolates fluorescence-stained cells into specially designed capture sites, generates high-resolution image data and prepares the associated cDNA libraries for mRNA sequencing. Current statistical methods focus on the analysis of the gene expression profiles and ignore the important information carried by the images. Here we propose a new direction for single-cell data analysis and develop CONFESS, a customized cell detection and fluorescence signal estimation model for images coming from the Fluidigm C1 system. Applied to a set of HeLa cells expressing fluorescence cell cycle reporters, the method predicted the progression state of hundreds of samples and enabled us to study the spatio-temporal dynamics of the HeLa cell cycle. The output can be easily integrated with the associated single-cell RNA-seq expression profiles for deeper understanding of a given biological system. CONFESS R package is available at Bioconductor (http://bioconductor.org/packages/release/bioc/html/CONFESS.html).

bioinformatics

Extracting a Biologically Relevant Latent Space from Cancer Transcriptomes with Variational Autoencoders

The Cancer Genome Atlas (TCGA) has profiled over 10,000 tumors across 33 different cancer-types for many genomic features, including gene expression levels. Gene expression measurements capture substantial information about the state of each tumor. Certain classes of deep neural network models are capable of learning a meaningful latent space. Such a latent space could be used to explore and generate hypothetical gene expression profiles under various types of molecular and genetic perturbation. For example, one might wish to use such a model to predict a tumors response to specific therapies or to characterize complex gene expression activations existing in differential proportions in different tumors. Variational autoencoders (VAEs) are a deep neural network approach capable of generating meaningful latent spaces for image and text data. In this work, we sought to determine the extent to which a VAE can be trained to model cancer gene expression, and whether or not such a VAE would capture biologically-relevant features. In the following report, we introduce a VAE trained on TCGA pan-cancer RNA-seq data, identify specific patterns in the VAE encoded features, and discuss potential merits of the approach. We name our method \"Tybalt\" after an instigative, cat-like character who sets a cascading chain of events in motion in Shakespeares \"Romeo and Juliet\". From a systems biology perspective, Tybalt could one day aid in cancer stratification or predict specific activated expression patterns that would result from genetic changes or treatment effects.

bioinformatics