bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Integrated aqueous humor ceRNA and miRNA-TF-mRNA network analysis reveals potential molecular mechanisms governing primary open-angle glaucoma pathogenesis

Primary open-angle glaucoma (POAG) is the leading cause of blindness globally, which develops through complex and poorly understood biological mechanisms. Herein, we conducted an integrated bioinformatics analysis of extant aqueous humor (AH) gene expression datasets in order to identify key genes and regulatory mechanisms governing POAG progression. We downloaded AH gene expression datasets (GSE101727 and GSE105269) corresponding to healthy controls and POAG patients from the Gene Expression Omnibus. We then identified mRNAs, microRNAs (miRNAs), and long non-coding RNAs (lncRNAs) that were differentially expressed (DE) between control and POAG patients. DEmRNAs and DElncRNAs were then subjected to pathway enrichment analyses, after which a protein-protein interaction (PPI) network was generated. This network was then expanded to establish lncRNA-miRNA-mRNA and miRNA-transcription factor(TF)-mRNA networks. In total, the GSE101727 dataset was used to identify 2746 DElncRNAs and 2208 DEmRNAs, while the GSE105269 dataset was used to identify 45 DEmiRNAs. We ultimately constructed a competing endogenous RNA (ceRNA) network incorporating 37, 5, and 14 of these lncRNAs, miRNAs and mRNAs, respectively. The proteins encoded by these 14 hub mRNAs were found to be significantly enriched for activities that may be linked to POAG pathogenesis. In addition, we generated a miRNA-TF-mRNA regulatory network containing 2 miRNAs (miR-135a-5p and miR-139-5p), 5 TFs (TGIF2, TBX5, HNF1A, TCF3, and FOS) and 5 mRNAs (SHISA7, ST6GAC2, TXNIP, FOS, and DCBLD2). The SHISA7, ST6GAC2, TXNIP, FOS, and DCBLD2 genes that may be viable therapeutic targets for the prevention or treatment of POAG, and regulated by the TFs (TGIF2, HNF1A, TCF3, and FOS).

molecular biology

Effect of shear and tensile loading on fibrin molecular structure revealed by coherent Raman microscopy

Blood clots are essential biomaterials that prevent blood loss and provide a temporary scaffold for tissue repair. In their function, these materials must be capable of resisting mechanical forces from hemodynamic shear and contractile tension without rupture. Fibrin networks, the primary load-bearing element in blood clots, have unique nonlinear mechanical properties resulting from their hierarchical structure, which provides multiscale load bearing from fiber deformation to protein unfolding. Here, we study the fiber and molecular scale response of fibrin under shear and tensile loads in situ using a combination of fluorescence and vibrational (molecular) microscopy. Imaging protein fiber orientation and molecular vibrations, we find that fiber orientation and molecular changes in fibrin appear at much larger strains under shear compared to uniaxial tension. Orientation levels reached at 150% shear strain were reached already at 60% tensile strain, and molecular unfolding of fibrin was only seen at shear strains above 300%, whereas fibrin unfolding began already at 20% tensile strain. Moreover, shear deformation caused progressive changes in vibrational modes consistent with increased protofibril and fiber packing that were already present even at very low tensile deformation. Together with a bioinformatic analysis of the fibrinogen primary structure, we propose a scheme for the molecular response of fibrin from low to high deformation, which may relate to the teleological origin of its resistance to shear and tensile forces. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=71 SRC="FIGDIR/small/205005v1_ufig1.gif" ALT="Figure 1"> View larger version (21K): org.highwire.dtl.DTLVardef@d72dfborg.highwire.dtl.DTLVardef@10bed75org.highwire.dtl.DTLVardef@12d33aorg.highwire.dtl.DTLVardef@1e9b40f_HPS_FORMAT_FIGEXP M_FIG C_FIG

biophysics

Protein Structure Refinement Guided by Atomic Packing Frustration Analysis

1Recent advances in machine learning, bioinformatics and the understanding of the folding problem have enabled efficient predictions of protein structures with moderate accuracy, even for targets when there is little information from templates. All-atom molecular dynamics simulations provide a route to refine such predicted structures, but unguided atomistic simulations, even when lengthy in time, often fail to eliminate incorrect structural features that would allow the structure to become more energetically favorable owing to the necessity of making large scale motions and overcoming energy barriers for side chain repacking. In this study, we show that localizing packing frustration at atomic resolution by examining the statistics of the energetic changes that occur when the local environment of a site is changed allows one to identify the most likely locations of incorrect contacts. The global statistics of atomic resolution frustration in structures that have been predicted using various algorithms provide strong indicators of structural quality when tested over a database of 20 targets from previous CASP experiments. Residues that are more correctly located turn out to be more minimally frustrated than more poorly positioned sites. These observations provide a diagnosis of both global and local quality of predicted structures, and thus can be used as guidance in all-atom refinement simulations of the 20 targets. Refinement simulations guided by atomic packing frustration turn out to be quite efficient and significantly improve the quality of the structures.

biophysics

Evolution and genomic signatures of spontaneous somatic mutation in Drosophila intestinal stem cells

Spontaneous mutations can alter tissue dynamics and lead to cancer initiation. While large-scale sequencing projects have illustrated processes that influence somatic mutation and subsequent tumour evolution, the mutational dynamics operating in the very early stages of cancer development are currently not well understood. In order to explore mutational dynamics in the early stages of cancer evolution we exploited neoplasia arising spontaneously in the Drosophila intestine. We analysed whole-genome sequencing data through the development of a dedicated bioinformatic pipeline to detect structural variants, single nucleotide variants, and indels. We found neoplasia formation to be driven largely through the inactivation of Notch by structural variants, many of which involve highly complex genomic rearrangements. Strikingly, the genome-wide mutational burden of neoplasia - at six weeks of age - was found to be similar to that of several human cancers. Finally, we identified genomic features associated with spontaneous mutation and defined the evolutionary dynamics and mutational landscape operating within intestinal neoplasia over the short lifespan of the adult fly. Our findings provide unique insight into mutational dynamics operating over a short time scale in the genetic model system, Drosophila melanogaster.

genomics

A biaryl-linked tripeptide from Planomonospora leads to widespread class of minimal RiPP gene clusters

Microbial natural products impress by their bioactivity, structural diversity and ingenious biosynthesis. While screening the rare actinobacterial genus Planomonospora, cyclopeptides 1A and 1B were discovered, featuring an unusual Tyr-His biaryl-bridging across a tripeptide scaffold, with the sequences N-acetyl-Tyr-Tyr-His (1A) and N-acetyl-Tyr-Phe-His (1B). Genome analysis of the 1A producing strain pointed to-wards a ribosomal synthesis of 1A, from a pentapeptide precursor encoded by the tiny 18-nucleotide gene bycA, to our knowledge the smallest gene ever reported. Further, biaryl instalment is performed by the closely linked gene bycB, encoding a cytochrome P450 monooxygenase. Biosynthesis of 1A was confirmed by heterologous production in Streptomyces, yielding the mature product. Bioinformatic analysis of related cytochrome P450 monooxygenases indicated that they constitute a widespread family of pathways, associated to 5-aa coding sequences in approximately 200 (actino)bacterial genomes, all with potential for a biaryl linkage between amino acids 1 and 3. We propose the name biarylicins for this newly discovered family of RiPPs.

microbiology

Multivariate genome-wide association study of rapid automatized naming and rapid alternating stimulus in Hispanic and African American youth.

Reading disability is a complex neurodevelopmental disorder that is characterized by difficulties in reading despite educational opportunity and normal intelligence. Performance on rapid automatized naming (RAN) and rapid alternating stimulus (RAS) tests gives a reliable predictor of reading outcome. These tasks involve the integration of different neural and cognitive processes required in a mature reading brain. Most studies examining the genetic factors that contribute to RAN and RAS performance have focused on pedigree-based analyses in samples of European descent, with limited representation of groups with Hispanic or African ancestry. In the present study, we conducted a multivariate genome-wide association analysis to identify shared genetic factors that contribute to performance across RAN Objects, RAN Letters, and RAS Letters/Numbers in a sample of Hispanic and African American youth (n=1,331). We then tested whether these factors also contribute to variance in reading fluency and word reading. Genome-wide significant, pleiotropic, effects across RAN Objects, RAN Letters, and RAS Letters/Numbers were observed for SNPs located on chromosome 10q23.31 (rs1555839, multivariate association, p=2.23 x 10-8), which also showed significant association with reading fluency and word reading performance (p <0.001). Bioinformatic analysis of this region using epigenetic data from the NIH Roadmap Epigenomics Mapping Consortium indicates active transcription of the gene RNLS in the brain. Neuroimaging genetic analysis of fourteen cortical regions in an independent sample of typically developing children across multiple ethnicities (n=690) showed that rs1555839 was associated with variation in volume of the right inferior parietal cortex--a region of the brain that processes numerical information and has been implicated in reading disability. This study provides support for a novel locus on chromosome 10q23.31 associated with RAN, RAS, and reading-related performance.\n\nAUTHOR SUMMARYReading disability has a strong genetic component that is explained by multiple genes and genetic factors. The complex genetic architecture along with diverse cognitive impairments associated with reading disability, poses challenges in identifying novel genes and variants that confer risk. One method to begin parsing genetic and neurobiological mechanisms that contribute to reading disability is to take advantage of the high correlation among reading-related cognitive traits like rapid automatized naming (RAN) and rapid alternating stimulus (RAS) to identify shared genetic factors that contribute to common biological mechanisms. In the present study, we used a multivariate genome-wide analysis approach that identified a region of chromosome 10q23.31 associated with variation in RAN Objects, RAN Letters, and RAS Letters/Numbers performance in a sample of 1,331 Hispanic and African American youth in the Genes, Reading, and Dyslexia (GRaD) Study. Genetic variants in this region were also associated with reading fluency in GRaD, and differences in brain structures implicated in reading disability in a separate sample of 690 children. The gene, RNLS, is located within the implicated region of chromosome 10q23.31 and plays a role in breaking down a class of chemical messengers known to affect attention, learning, and memory in the brain. These findings provide a basis to inform our understanding of the biological basis of reading disability.

genetics

A most wanted list of conserved protein families with no known domains

The number and proportion of genes with no known function are growing rapidly. To quantify this phenomenon and provide criteria for prioritizing genes for functional characterization, we developed a bioinformatics pipeline that identifies robustly defined protein families with no annotated domains, ranks these with respect to phylogenetic breadth, and identifies them in metagenomics data. We applied this approach to 271 965 protein families from the SFams database and discovered many with no functional annotation, including >118 000 families lacking any known protein domain. From these, we prioritized 6 668 conserved protein families with at least three sequences from organisms in at least two distinct classes. These Function Unknown Families (FUnkFams) are present in Tara Oceans Expedition and Human Microbiome Project metagenomes, with distributions associated with sampling environment. Our findings highlight the extent of functional novelty in sequence databases and establish an approach for creating a \"most wanted\" list of genes to characterize.

genomics

Two C++ Libraries for Counting Trees on a Phylogenetic Terrace

MotivationThe presence of terraces in phylogenetic tree space, that is, a potentially large number of distinct tree topologies that have exactly the same analytical likelihood score, was first described by Sanderson et al, (2011). However, popular software tools for maximum likelihood and Bayesian phylogenetic inference do not yet routinely report, if inferred phylogenies reside on a terrace, or not. We believe, this is due to the unavailability of an efficient library implementation to (i) determine if a tree resides on a terrace, (ii) calculate how many trees reside on a terrace, and (iii) enumerate all trees on a terrace.\n\nResultsIn our bioinformatics programming practical we developed two efficient and independent C++ implementations of the SUPERB algorithm by Constantinescu and Sankoff (1995) for counting and enumerating the trees on a terrace. Both implementations yield exactly the same results and are more than one order of magnitude faster and require one order of magnitude less memory than a previous 3rd party python implementation.\n\nAvailabilityThe source codes are available under GNU GPL at https://github.com/terraphast\n\nContactAlexandros.Stamatakis@h-its.org

evolutionary biology

Mmp10 is required for post-translational methylation of arginine at the active site of methyl-coenzyme M reductase

Catalyzing the key step for anaerobic methane production and oxidation, methyl-coenzyme M reductase or Mcr plays a key role in the global methane cycle. The McrA subunit possesses up to five post-translational modifications (PTM) at its active site. Bioinformatic analyses had previously suggested that methanogenesis marker protein 10 (Mmp10) could play an important role in methanogenesis. To examine its role, MMP1554, the gene encoding Mmp10 in Methanococcus maripaludis, was deleted with a new genetic tool, resulting in the specific loss of the 5-(S)-methylarginine PTM of residue 275 in the McrA subunit and a 40~60 % reduction in the maximal rates of methane formation by whole cells. Methylation was restored by complementations with the wild-type gene. However, the rates of methane formation of the complemented strains were not always restored to the wild type level. This study demonstrates the importance of Mmp10 and the methyl-Arg PTM on Mcr activity.

microbiology

Impact of sequence variant detection and bacterial DNA extraction methods on the measurement of microbial community composition in human stool

BackgroundThe human gut microbiome has been widely studied in the context of human health and metabolism, however the question of how to analyze this community remains contentious. This study compares new and previously well established methods aimed at reducing bias in bioinformatics analysis (QIIME 1 and DADA2) and bacterial DNA extraction of human fecal samples in 16S rRNA marker gene surveys.\n\nResultsAnalysis of a mock DNA community using DADA2 identified more chimeras (QIIME 1: 0.70% of total reads vs DADA2: 1.96%), fewer sequence variants, (QIIME 1: 1297.4 + 98.88 vs. DADA2: 136.27 + 11.35, mean + SD) and correct taxa at a higher resolution of classification (i.e. genus-level) than open reference OTU picking in QIIME 1. Additionally, the extraction of whole cell mock community bacterial DNA using four commercially available kits resulted in varying DNA yield, quality and bacterial community composition. Of the four kits compared, ZymoBIOMICS DNA Miniprep Kit provided the greatest yield, with a slight enrichment of Enterococcus. However, QIAamp Fast DNA Stool Mini Kit resulted in the highest DNA quality. Mo Bio PowerFecal DNA Kit had the most dramatic effect on the mock community composition, resulting in an increased proportion of members of the family Enterobacteriaceae and genus Eshcerichia as well as members of genera Lactobacillus and Pseudomonas. The presence of a sterile fecal matrix had a slight, but inconsistent effect on the yield, quality and taxa identified after extraction with all four DNA extraction kits. Extraction of bacterial DNA from native stool samples revealed a distinct effect of the DNA stabilization reagent DNA/RNA Shield on community composition, causing an increase in the detected abundance of members of orders Bifidobacteriales, Bacteroidales, Turicibacterales, Clostridiales and Enterobacteriales.\n\nConclusionThese results confirm that the DADA2 algorithm is superior to sequence clustering by similarity to determine microbial community structure. Additionally, commercially available kits used for bacterial DNA extraction from fecal samples have some effect on the proportion of high abundance members detected in a microbial community, but it is less significant than the effect of using DNA stabilization reagent, DNA/RNA Shield.

molecular biology

OMGene: Mutual improvement of gene models through optimisation of evolutionary conservation

BackgroundThe accurate determination of the genomic coordinates for a given gene - its gene model - is of vital importance to the utility of its annotation, and the accuracy of bioinformatic analyses derived from it. Currently-available methods of computational gene prediction, while on the whole successful, often disagree on the model for a given predicted gene, with some or all of the variant gene models failing to match the biologically observed structure. Many prediction methods can be bolstered by using experimental data such as RNA-seq and mass spectrometry. However, these resources are not always available, and rarely give a comprehensive portrait of an organisms transcriptome due to temporal and tissue-specific expression profiles.\n\nResultsOrthology between genes provides evolutionary evidence to guide the construction of gene models. OMGene (Optimise My Gene) aims to optimise gene models in the absence of experimental data by optimising the derived amino acid alignments for gene models within orthogroups. Using RNA-seq data sets from plants and fungi, considering intron/exon junction representation and exon coverage, and assessing the intra-orthogroup consistency of subcellular localisation predictions, we demonstrate the utility of OMGene for improving gene models in annotated genomes.\n\nConclusionsWe show that significant improvements in the accuracy of gene model annotations can be made in both established and de novo annotated genomes by leveraging information from multiple species.

genomics

Genomic Locus Modulating Corneal Thickness in the Mouse Identifies POU6F2 as a Potential Risk of Developing Glaucoma

Purpose: Central corneal thickness (CCT) is one of the most heritable ocular traits and it is also a phenotypic risk factor for primary open angle glaucoma (POAG). The present study uses the BXD Recombinant Inbred (RI) strains to identify novel quantitative trait loci (QTLs) modulating CCT in the mouse with the potential of identifying a molecular link between CCT and risk of developing POAG.\n\nMethods: The BXD RI strain set was used to define mammalian genomic loci modulating CCT, with a total of 818 corneas measured from 61 BXD RI strains (between 60-100 days of age). The mice were anesthetized and the eyes were positioned in front of the lens of the Phoenix Micron IV Image-Guided OCT system or the Bioptigen OCT system. CCT data for each strain was averaged and used to identify quantitative trait loci (QTLs) modulating this phenotype using the bioinformatics tools on GeneNetwork (www.genenetwork.org). The candidate genes and genomic loci identified in the mouse were then directly compared with the summary data from a human primary open-angle glaucoma (POGA) genome wide association study (NEIGHBORHOOD) to determine if any genomic elements modulating mouse CCT are also risk factors for POAG.\n\nResults: This analysis revealed one significant QTL on Chr 13 and a suggestive QTL on Chr 7. The significant locus on Chr 13 (13 to 19 Mb) was examined further to define candidate genes modulating this eye phenotype. For the Chr 13 QTL in the mouse, only one gene in the region (Pou6f2) contained nonsynonymous SNPs. Of these five nonsynonymous SNPs in Pou6f2, two resulted in changes in the amino acid proline which could result in altered secondary structure affecting protein function. The 7 Mb region under the mouse Chr 13 peak distributes over 2 chromosomes in the human: Chr 1 and Chr 7. These genomic loci were examined in the NEIGHBORHOOD database to determine if they are potential risk factors for human glaucoma identified using meta-data from human GWAS. The top 50 hits all resided within one gene (POU6F2), with the highest significance level of p = 10-6 for SNP rs76319873. POU6F2 is found in retinal ganglion cells and in corneal limbal stem cells. To test the effect of POU6F2 on CCT we examined the corneas of a Pou6f2-null mice and the corneas were thinner than those of wild-type littermates. In addition, these POU6F2 RGCs die early in the DBA/2J model of glaucoma than most RGCs.\n\nConclusions: Using a mouse genetic reference panel, we identified a transcription factor, Pou6f2, that modulates CCT in the mouse. POU6F2 is also found in a subset of retinal ganglion cells and these RGCs are sensitive to injury.\n\nAuthors SummaryGlaucoma is a complex group of diseases with several known causal mutations and many known risk factors. One well-known risk factor for developing primary open angle glaucoma is the thickness of the central cornea. The present study leverages a unique blend of systems biology methods using BXD recombinant inbred mice and genome-wide association studies from humans to define a putative molecular link between a phenotypic risk factor (central corneal thickness) and glaucoma. We identified a transcription factor, POU6F2, that is found in the developing retinal ganglion cells and cornea. POU6F2 is also present in a subpopulation of retinal ganglion cells and in stem cells of the cornea. Functional studies reveal that POU6F2 is associated the central corneal thickness and with susceptibility of retinal ganglion cells to injury.

genetics

Mitochondrial genomes of the regionally extinct Nittany Lion (Puma concolor from Pennsylvania)

Mountain lions (Puma concolor) were once endemic across the United States. The Northeastern population of mountain lions has been largely nonexistent since the early 1800s and was officially declared extinct in 2011. This regionally extinct mountain lion is Pennsylvania State Universitys official mascot, where it is referred to as the Nittany Lion. Our goal in this study was to use recent methodological advances in ancient DNA and massively parallel sequencing to reconstruct complete mitochondrial DNA (mtDNA) genomes of multiple Nittany Lions by sampling from preserved skins. This effort is part of a broader Nittany Lion Genome project intended to involve undergraduates in ancient DNA and bioinformatics research and to engage the broader Penn State community in discussions about conservation biology and extinction. Complete mtDNA genome sequences were obtained from five individuals. When compared to previously published sequences, Nittany Lions are not more similar to each other than to individuals from the Western U.S. and Florida. Supporting previous findings, North American mountain lions overall were more closely related to each other than to those from South America and had lower genetic diversity. This result emphasizes the importance of continued conservation in the Western U.S. and Florida to prevent further regional extinctions.

ecology

Transcriptional Profiling of Somatostatin Interneurons in the Spinal Dorsal Horn

The spinal dorsal horn (SDH) is comprised of distinct neuronal populations that process different somatosensory modalities. Somatostatin (SST)-expressing interneurons in the SDH have been implicated specifically in mediating mechanical pain. Identifying the transcriptomic profile of SST neurons could elucidate the unique genetic features of this population and enable selective analgesic targeting. To that end, we combined the Isolation of Nuclei Tagged in Specific Cell Types (INTACT) method and Fluorescence Activated Nuclei Sorting (FANS) to capture tagged SST nuclei in the SDH of adult male mice. Using RNA-sequencing (RNA-seq), we uncovered more than 13,000 genes. Differential gene expression analysis revealed more than 900 genes with at least 2-fold enrichment. In addition to many known dorsal horn genes, we identified and validated several novel transcripts from pharmacologically tractable functional classes: Carbonic Anhydrase 12 (Car12), Phosphodiesterase 11A (Pde11a), Protease-Activated Receptor 3 (F2rl2) and G-protein Coupled Receptor 26 (Gpr26). In situ hybridization of these novel genes revealed differential expression patterns in the SDH, demonstrating the presence of transcriptionally distinct subpopulations within the SST population. Pathway analysis revealed several enriched signaling pathways including cyclic AMP-mediated signaling, Nitric Oxide Synthase signaling, and voltage-gated calcium channels, highlighting the importance of these pathways to SST neuron function. Overall, our findings provide new insights into the gene repertoire of SST dorsal horn neurons and reveal several candidate targets for pharmacological modulation of this pain-mediating population.\n\nSignificance StatementSomatostatin(SST)-expressing interneurons in the spinal dorsal horn (SDH) are required for the perception of mechanical pain. Identifying the distinctive genes expressed by SST neurons could facilitate the development of novel, circuit-targeting analgesics. Thus, we applied cell type-specific RNA-sequencing (RNA-seq) to provide the first transcriptional profile of SST neurons in the SDH. Bioinformatic analysis revealed hundreds of genes enriched in SST neurons, including several previously undescribed genes from druggable classes (Car12, Pde11a, F2rl2 and Gpr26). Taken together, our study unveils a comprehensive transcriptional signature for SST neurons, highlights promising candidate genes for future analgesic development, and establishes a flexible method for transcriptional profiling of any spinal cord cell type.

neuroscience

Mechanical changes in proteins with large-scale motions highlight the formation of structural locks

Protein function depends just as much on flexibility as on structure, and in numerous cases, a proteins biological activity involves transitions that will impact both its conformation and its mechanical properties. Here, we use a coarse-grain approach to investigate the impact of structural changes on protein flexibility. More particularly, we focus our study on proteins presenting large-scale motions. We show how calculating directional force constants within residue pairs, and investigating their variation upon protein closure, can lead to the detection of a limited set of residues that form a structural lock in the proteins closed conformation. This lock, which is composed of residues whose side-chains are tightly interacting, highlights a new class of residues that are important for protein function by stabilizing the closed structure, and that cannot be detected using earlier tools like local rigidity profiles or distance variations maps, or alternative bioinformatics approaches, such as coevolution scores.

biophysics

Tn-Core: context-specific reconstruction of core metabolic models using Tn-seq data

MotivationTn-seq (transposon mutagenesis and sequencing) and constraint-based metabolic modelling represent highly complementary approaches. They can be used to probe the core genetic and metabolic networks underlying a biological process, revealing invaluable information for synthetic biology engineering of microbial cell factories. However, while algorithms exist for integration of -omics data sets with metabolic models, no method has been explicitly developed for integration of Tn-seq data with metabolic reconstructions.\n\nResultsWe report the development of Tn-Core, a Matlab toolbox designed to generate gene-centric, context-specific core reconstructions consistent with experimental Tn-seq data. Extensions of this algorithm allow: i) the generation of context-specific functional models through integration of both Tn-seq and RNA-seq data; ii) to visualize redundancy in core metabolic processes; and iii) to assist in curation of de novo draft metabolic models. The utility of Tn-Core is demonstrated primarily using a Sinorhizobium meliloti model as a case study.\n\nAvailability and implementationThe software can be downloaded from https://github.com/diCenzo-GC/Tn-Core. All results presented in this work have been obtained with Tn-Core v. 1.0.\n\nContactgeorgecolin.dicenzo@unifi.it, marco.fondi@unifi.it\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Characterizing Cancer Drug Response andBiological Correlates: A Geometric NetworkApproach

In the present work, we consider a geometric network approach to study common biological features of anticancer drug response. We use for this purpose the panel of 60 human cell lines (NCI-60) provided by the National Cancer Institute. Our study suggests that utilization of mathematical tools for network-based analysis can provide novel insights into drug response and cancer biology. We adopted a discrete notion of Ricci curvature to measure the robustness of biological networks constructed with a pre-treatment gene expression dataset and coupled the results with the GI50 response of the cell lines to the drugs. The link between network robustness and Ricci curvature was implemented using the theory of optimal mass transport. Our hypothesis behind this idea is that robustness in the biological network contributes to tumor drug resistance, thereby enabling us to predict the effectiveness and sensitivity of drugs in the cell lines. Based on the resulting drug response ranking, we assessed the impact of genes that are likely associated with individual drug response. For important genes identified, we performed a gene ontology enrichment analysis using a curated bioinformatics database which resulted in very plausible biological processes associated with drug response across cell lines and cell types from the biological and literature viewpoint. These results demonstrate the potential of using the mathematical network analysis in assessing drug response and in identifying relevant genomic biomarkers and biological processes for precision medicine.

cancer biology

Nextstrain: real-time tracking of pathogen evolution

SummaryUnderstanding the spread and evolution of pathogens is important for effective public health measures and surveillance. Nextstrain consists of a database of viral genomes, a bioinformatics pipeline for phylodynamics analysis, and an interactive visualisation platform. Together these present a real-time view into the evolution and spread of a range of viral pathogens of high public health importance. The visualization integrates sequence data with other data types such as geographic information, serology, or host species. Nextstrain compiles our current understanding into a single accessible location, publicly available for use by health professionals, epidemiologists, virologists and the public alike.\n\nAvailability and implementationAll code (predominantly JavaScript and Python) is freely available from github.com/nextstrain and the web-application is available at nextstrain.org.

evolutionary biology