bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Developmental chromatin restriction of pro-growth gene networks acts as an epigenetic barrier to axon regeneration in cortical neurons

Axon regeneration in the central nervous system is prevented in part by a developmental decline in the intrinsic regenerative ability of maturing neurons. This loss of axon growth ability likely reflects widespread changes in gene expression, but the mechanisms that drive this shift remain unclear. Chromatin accessibility has emerged as a key regulatory mechanism in other cellular contexts, raising the possibility that chromatin structure may contribute to the age-dependent loss of regenerative potential. Here we establish an integrated bioinformatic pipeline that combines analysis of developmentally dynamic gene networks with transcription factor regulation and genome-wide maps of chromatin accessibility. When applied to the developing cortex, this pipeline detected overall closure of chromatin in sub-networks of genes associated with axon growth. We next analyzed mature CNS neurons that were supplied with various pro-regenerative transcription factors. Unlike prior results with SOX11 and KLF7, here we found that neither JUN nor an activated form of STAT3 promoted substantial corticospinal tract regeneration. Correspondingly, chromatin accessibility in JUN or STAT3 target genes was substantially lower than in predicted targets of SOX11 and KLF7. Finally, we used the pipeline to predict pioneer factors that could potentially relieve chromatin constraints at growth-associated loci. Overall this integrated analysis substantiates the hypothesis that dynamic chromatin accessibility contributes to the developmental decline in axon growth ability and influences the efficacy of pro-regenerative interventions in the adult, while also pointing toward selected pioneer factors as high-priority candidates for future combinatorial experiments.

neuroscience

Template switching causes artificial junction formation and false identification of circular RNAs

Hundreds of thousands of putative circular RNAs have been identified through deep sequencing and bioinformatic analyses. However, the circularity of these putative RNA circles has not been experimentally validated due to limited methodologies currently available. We reported here that the template-switching capability of commonly used reverse transcriptases (e.g., SuperScript II) leads to the formation of artificial junction sequences, and consequently misclassification of large linear RNAs as RNA circles. Use of reverse transcriptases without terminal transferase activity (e.g., MonsterScript) for cDNA synthesis is critical for the identification of physiological circular RNAs. We also report two methods, MonsterScript junction PCR and high-resolution melting curve analyses, which can reliably distinguish circular RNAs from their linear forms and thus, can be used to discover and validate true circular RNAs.\n\nSignificance StatementThe vast majority of circular RNAs were identified through computational detection of junction sequences in the deep sequencing reads because these unique fusion sequences represent back-splicing events. We found that artificial junction sequences could be formed through template switching (TS) when MMLV-derived reverse transcriptases, e.g., SuperScript II, are used to synthesize cDNAs. Thus, many of the reported circular RNAs may not be RNA circles, but rather experimental artifacts. Fake circular RNAs can be avoided by using reverse transcriptases without terminal transferase activity (e.g., MonsterScript) for cDNA synthesis. We developed two novel methods, MonsterScript junction PCR and high-resolution melting curve analyses, for distinguishing circular RNAs from their linear form.

molecular biology

Gene Coregulation and Coexpression in the Aryl Hydrocarbon Receptor-mediated Transcriptional Regulatory Network in the Mouse Liver

Tissue-specific network models of chemical-induced gene perturbation can improve our mechanistic understanding of the intracellular events leading to adverse health effects resulting from chemical exposure. The aryl hydrocarbon receptor (AHR) is a ligand-inducible transcription factor (TF) that activates a battery of genes and produces a variety of species-specific adverse effects in response to the potent and persistent environmental contaminant 2,3,7,8-tetrachlorodibenzo-p-dioxin (TCDD). Here we assemble a global map of the AHR gene regulatory network in TCDD-treated mouse liver from a combination of previously published gene expression and genome-wide TF binding data sets. Using Kohonen selforganizing maps and subspace clustering, we show that genes co-regulated by common upstream TFs in the AHR network exhibit a pattern of co-expression. Specifically, directly-bound, indirectly-bound and non-genomic AHR target genes exhibit distinct patterns of gene expression, with the directly bound targets generally associated with highest median expression. Further, among the directly bound AHR target genes, the expression level increases with the number of AHR binding sites in the proximal promoter regions. Finally, we show that co-regulated genes in the AHR network activate distinct groups of downstream biological processes, with AHR-bound target genes enriched for metabolic processes and enrichment of immune responses among AHR-unbound target genes, likely reflecting infiltration of immune cells into the mouse liver upon TCDD treatment. This work describes an approach to the reconstruction and analysis of transcriptional regulatory cascades underlying cellular stress response using bioinformatic and statistical tools.

genomics

Long-read sequencing reveals the splicing profile of the calcium channel gene CACNA1C in human brain

RNA splicing is a key mechanism linking genetic variation with psychiatric disorders. Splicing profiles are particularly diverse in brain and difficult to accurately identify and quantify. We developed a new approach to address this challenge, combining long-range PCR and nanopore sequencing with a novel bioinformatics pipeline. We identify the full-length coding transcripts of CACNA1C in human brain. CACNA1C is a psychiatric risk gene that encodes the voltage-gated calcium channel CaV1.2. We show that CACNA1Cs transcript profile is substantially more complex than appreciated, identifying 38 novel exons and 241 novel transcripts. Importantly, many of the novel variants are abundant, and predicted to encode channels with altered function. The splicing profile varies between brain regions, especially in cerebellum. We demonstrate that human transcript diversity (and thereby protein isoform diversity) remains under-characterised, and provide a feasible and cost-effective methodology to address this. A detailed understanding of isoform diversity will be essential for the translation of psychiatric genomic findings into pathophysiological insights and novel psychopharmacological targets.

neuroscience

Genome-wide study identifies 611 loci associated with risk tolerance and risky behaviors

Humans vary substantially in their willingness to take risks. In a combined sample of over one million individuals, we conducted genome-wide association studies (GWAS) of general risk tolerance, adventurousness, and risky behaviors in the driving, drinking, smoking, and sexual domains. We identified 611 approximately independent genetic loci associated with at least one of our phenotypes, including 124 with general risk tolerance. We report evidence of substantial shared genetic influences across general risk tolerance and risky behaviors: 72 of the 124 general risk tolerance loci contain a lead SNP for at least one of our other GWAS, and general risk tolerance is moderately to strongly genetically correlated ([Formula] to 0.50) with a range of risky behaviors. Bioinformatics analyses imply that genes near general-risk-tolerance-associated SNPs are highly expressed in brain tissues and point to a role for glutamatergic and GABAergic neurotransmission. We find no evidence of enrichment for genes previously hypothesized to relate to risk tolerance.

genetics

NeoDTI: Neural integration of neighborinformation from a heterogeneous network fordiscovering new drug-target interactions

MotivationAccurately predicting drug-target interactions (DTIs) in silico can guide the drug discovery process and thus facilitate drug development. Computational approaches for DTI prediction that adopt the systems biology perspective generally exploit the rationale that the properties of drugs and targets can be characterized by their functional roles in biological networks.\n\nResultsInspired by recent advance of information passing and aggregation techniques that generalize the convolution neural networks (CNNs) to mine large-scale graph data and greatly improve the performance of many network-related prediction tasks, we develop a new nonlinear end-to-end learning model, called NeoDTI, that integrates diverse information from heterogeneous network data and automatically learns topology-preserving representations of drugs and targets to facilitate DTI prediction. The substantial prediction performance improvement over other state-of-the-art DTI prediction methods as well as several novel predicted DTIs with evidence supports from previous studies have demonstrated the superior predictive power of NeoDTI. In addition, NeoDTI is robust against a wide range of choices of hyperparameters and is ready to integrate more drug and target related information (e.g., compound-protein binding affinity data). All these results suggest that NeoDTI can offer a powerful and robust tool for drug development and drug repositioning.\n\nAvailability and implementationThe source code and data used in NeoDTI are available at: https://github.com/FangpingWan/NeoDTI.\n\nContactzengjy321@tsinghua.edu.cn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Functional characterization of sensory neuron membrane proteins (SNMPs)

Sensory neuron membrane proteins (SNMPs) play a critical role in the insect olfactory system but there is a deficit of functional studies beyond Drosophila. Here, we provide functional characterisation of insect SNMPs through the use of bioinformatics, genome curation, transcriptome data analysis, phylogeny, expression profiling, and RNAi gene knockdown techniques. We curated 81 genes from 35 insect species and identified a novel lepidopteran SNMP gene family, SNMP3. Phylogenetic analysis shows that lepidopteran SNMP3, but not the previously annotated lepidopteran SNMP2, is the true homologue of the dipteran SNMP2. Digital expression, microarray and qPCR analyses show that the lepidopteran SNMP1 is specifically expressed in adult antennae. SNMP2 is widely expressed in multiple tissues while SNMP3 is specifically expressed in the larval midgut. Microarray analysis suggest SNMP3 may be involved in the silkworm immunity response to virus and bacterial infections. We functionally characterised SNMP1 in the silkworm using RNAi and behavioural assays. Our results suggested that Bombyx mori SNMP1 is a functional orthologue of the Drosophila melanogaster SNMP1 and plays a critical role in pheromone detection. Split-ubiquitin yeast hybridization study shows that BmorSNMP1 has a protein-protein interaction with the BmorOR1 pheromone receptor, and the BmorOrco co-receptor. Concluding, we propose a novel molecular model in which BmorOrco, BmorSNMP1 and BmorOR1 form a heteromer in the detection of the silkworm sex pheromone bombykol.

molecular biology

Structural model of Cyc2, the primary electron acceptor of Acidithiobacillus ferrooxidans respiratory chain, as a modular cytochrome - β-barrel fusion protein, and mechanistic proposals based on this model

Acidithiobacillus ferrooxidans oxidizes Fe(II) to Fe(III) to feed electrons into its respiratory chain. The primary electron acceptor of this complex system is Cyc2, an outer membrane protein of unknown structure. This work proposes a feasible model of Cyc2s global structure, based on homology modeling, residue-residue coevolution data, bioinformatics predictions and limited knowledge about Cyc2s function. The proposal is that the sequence segment spanning residues ~30 to ~90 folds as a cytochrome-like domain that contains a heme group which would presumably bind and oxidize external Fe(II), whereas the remaining segment from residue ~90 until the end adopts a {beta}-barrel fold similar to that of most outer membrane proteins. Such model differs strongly from a published model, but is backed up by more data and is more compatible with the known topology of outer membrane proteins and with Cyc2s function of internalizing reducing equivalents. The small size of the cytochrome-like domain would allow it to reside inside, and/or slide through, the {beta}-barrel domain, thus communicating in a controlled fashion the extracellular medium with the periplasm to import electrons through the outer membrane. All the models discussed are provided as PyMOL session files in the Supporting Information and can be visualized online at http://lucianoabriata.altervista.org/modelshome.html

biophysics

Primate MHC class I from Genomes

The major histocompatibility complex (MHC) molecule plays a central role in the adaptive immunity of jawed vertebrates. Allelic variations have been studied extensively in some primate species, however a comprehensive description of the number of genes remains incomplete. Here, a bioinformatics program was developed to identify three MHC Class I exons (EX2, EX3 and EX4) from Whole Genome Sequencing (WGS) datasets. With this algorithm, MHC Class I exons sequences were extracted from 30 WGS datasets of primates, representatives of Apes, Old World and New World monkeys and prosimians. There is a high variability in the number of genes between species. From human WGS, six viable genes (HLA-A, -B, -C, -E, -F, and -G) and four pseudogene sequences (HLA-H, -J, -L, -V) are obtained. These genes serve to identify the phylogenetic clades of MHC-I in primates. The results indicate that human clades of HLA-A -B and -C were generated shortly after the separation of Old World monkeys. The clades pertaining to HLA-E, -H and -F are found in all primate families, except in Prosimians. In the clades defined by HLA-G, -L and -J, there are sequences from Old world monkeys. Specific clades are found in the four primate families. The evolution of these genes is consistent with birth and death processes having a high turnover rates.

immunology

Cytokinin perception in potato: New features of canonic players

Potato is the most economically important non-cereal food crop. Tuber formation in potato is regulated by phytohormones, cytokinins (CKs) in particular. The present work was aimed to study CK signal perception in potato. The sequenced potato genome of doubled monoploid Phureja was used for bioinformatic analysis and as a tool for identification of putative CK receptors from autotetraploid potato cv. Desiree. All basic elements of multistep phosphorelay (MSP) required for CK signal transduction were identified in Phureja genome, including three genes orthologous to three CK receptor genes (AHK 2-4) of Arabidopsis. As distinct from Phureja, autotetraploid potato contains at least two allelic isoforms of each receptor type. Putative receptor genes from Desiree plants were cloned, sequenced and expressed, and main characteristics of encoded proteins, firstly their consensus motifs, structure models, ligand-binding properties, and the ability to transmit CK signal, were determined. In all studied aspects the predicted sensor histidine kinases met the requirements for genuine CK receptors. Expression of potato CK receptors was found to be organ-specific and sensitive to growth conditions, particularly to sucrose content. Our results provide a solid basis for further in-depth study of CK signaling system and biotechnological improvement of potato.

plant biology

A pseudogene of caffeic acid-o-methyltransferase (COMT) in Acacia mangium: Comparative analysis with other COMT plant promoters

Acacia mangium is a prominent tree species in the forest plantation industry of Southeast Asia, grown mainly to produce pulp and paper, and to a lesser extent wood chips and solid wood products. Lignin, a natural complex polymer used by plants for structural support and defence, has to be chemically removed during the production of quality paper. Delignification is very expensive and moreover, is an environmental pollutant. Understanding the complex mechanisms that underlie the regulation of lignin biosynthetic genes requires in-depth knowledge of not only the genes involved but also their regulatory elements. Using Thermal Asymmetric Interlaced PCR, a 770 bp promoter sequence with 93% identity with COMT1 gene from Acacia auriculiformis x A. mangium hybrid was isolated from A. mangium. Bioinformatics analysis revealed the presence of cis acting elements commonly found in other lignin biosynthesis genes such as TATA box, CAAT box, W box, AC-I and AC-11 elements. However, a nonsense mutation that created a premature stop codon was found on the first exon. Modelling of MYB transcription factor binding site on this newly isolated pseudogene shows it has binding sites for important transcription factors involved in lignin biosynthesis both in Arabidopsis thaliana and Eucalyptus grandis. Given the remarkable structures of its regulatory region, the possible structure of its transcript was detected using Mfold. Results show the transcript are capable of forming stem loop structures, a characteristic commonly attributed to presence of miRNA. Possible functions of pseudoAmCOMT1 were discussed.

genetics

Optimization and uncertainty analysis of ODE models using second order adjoint sensitivity analysis

MotivationParameter estimation methods for ordinary differential equation (ODE) models of biological processes can exploit gradients and Hessians of objective functions to achieve convergence and computational efficiency. However, the computational complexity of established methods to evaluate the Hessian scales linearly with the number of state variables and quadratically with the number of parameters. This limits their application to low-dimensional problems.\n\nResultsWe introduce second order adjoint sensitivity analysis for the computation of Hessians and a hybrid optimization-integration based approach for profile likelihood computation. Second order adjoint sensitivity analysis scales linearly with the number of parameters and state variables. The Hessians are effectively exploited by the proposed profile likelihood computation approach. We evaluate our approaches on published biological models with real measurement data. Our study reveals an improved computational efficiency and robustness of optimization compared to established approaches, when using Hessians computed with adjoint sensitivity analysis. The hybrid computation method was more than two-fold faster than the best competitor. Thus, the proposed methods and implemented algorithms allow for the improvement of parameter estimation for medium and large scale ODE models.\n\nAvailabilityThe algorithms for second order adjoint sensitivity analysis are implemented in the Advance MATLAB Interface CVODES and IDAS (AMICI, https://github.com/ICB-DCM/AMICI/). The algorithm for hybrid profile likelihood computation is implemented in the parameter estimation toolbox (PESTO, https://github.com/ICB-DCM/PESTO/). Both toolboxes are freely available under the BSD license.\n\nContactjan.hasenauer@helmholtz-muenchen.de\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Rare Variant Pathogenicity Triage and Inclusion of Synonymous Variants Improves Analysis of Disease Associations

Many G protein-coupled receptors (GPCRs) lack common variants that lead to reproducible genome-wide disease associations. Here we used rare variant approaches to assess the disease associations of 85 orphan or understudied GPCRs in an unselected cohort of 51,289 individuals. Rare loss-of-function variants, missense variants predicted to be pathogenic or likely pathogenic, and a subset of rare synonymous variants were used as independent data sets for sequence kernel association testing (SKAT). Strong, phenome-wide disease associations shared by two or more variant categories were found for 39% of the GPCRs. Validating the bioinformatics and SKAT analyses, functional characterization of rare missense and synonymous variants of GPR39, a Family A GPCR, showed altered expression and/or Zn2+-mediated signaling for members of both variant classes. Results support the utility of rare variant analyses for identifying disease associations for genes that lack common variants, while also highlighting the functional importance of rare synonymous variants.\n\nAuthor summaryRare variant approaches have emerged as a viable way to identify disease associations for genes without clinically important common variants. Rare synonymous variants are generally considered benign. We demonstrate that rare synonymous variants represent a potentially important dataset for deriving disease associations, here applied to analysis of a set of orphan or understudied GPCRs. Synonymous variants yielded disease associations in common with loss-of-function or missense variants in the same gene. We rationalize their associations with disease by confirming their impact on expression and agonist activation of a representative example, GPR39. This study highlights the importance of rare synonymous variants in human physiology, and argues for their routine inclusion in any comprehensive analysis of genomic variants as potential causes of disease.

genetics

The Staphylococcus aureus Two-Component System AgrAC Displays Four Distinct Genomic Arrangements That Delineate Genomic Virulence Factor Signatures

Two-component systems (TCSs) consist of a histidine kinase and a response regulator. Here, we evaluated the conservation of the AgrAC TCS among 149 completely sequenced S. aureus strains. It is composed of four genes: agrBDCA. We found that: i) AgrAC system (agr) was found in all but one of the 149 strains; ii) The agr positive strains were further classified into four agr types based on AgrD protein sequences, iii) the four agr types not only specified the chromosomal arrangement of the agr genes but also the sequence divergence of AgrC histidine kinase protein, which confers signal specificity, iv) the sequence divergence was reflected in distinct structural properties especially in the transmembrane region and second extracellular binding domain, and v) there was a strong correlation between the agr type and the virulence genomic profile of the organism. Taken together, these results demonstrate that bioinformatic analysis of the agr locus leads to a classification system that correlates with the presence of virulence factors and protein structural properties.

microbiology

A novel GATA-binding protein 4 gene variation associated with familial atrial septal defect

Atrial septal defect (ASD) is the most common congenital heart defect. Part of ASD exhibits familial predisposition, but the genetic mechanism remains largely unknown. In the current study, we use multiple methods to identify and confirm the gene associated with a familial ASD. Chromosomal microarray analyses, whole exome sequencing, Sanger sequencing, multiple bioinformatics programs, in silico protein structure modeling and molecular dynamics simulation were performed to predict the pathogenic of the variant gene. Dual-Luciferase reporter gene assay was performed to evaluate the influence of downstream target gene of the target variation. A novel, heterozygous, missense variant GATA-binding protein 4 (GATA4):c.958C>T, p.R320W was identified. An autosomal dominant inheritance pattern with incomplete penetrance was observed in the family. Multiple prediction indicate the variant in GATA4 to be deleterious. Molecular dynamics simulation further revealed that the variation of p.R320W could prevent the zinc finger of GATA4 from interacting with the DNA. Dual-Luciferase reporter assay demonstrated a significant decrease in transcriptional activity (0.90{+/-}0.099 vs 1.50{+/-}0.079, p = 0.001) of the variant GATA4 compared with the wild type. We believe the novel variation of GATA4 (c.958C>T, p.R320W) with a pattern of incomplete inheritance that may be highly associated with this familial ASD. The finding enriched our knowledge of variations that may associated with ASD.

genetics

25 Years of Molecular Biology Databases: A Study of Proliferation, Impact, and Maintenance

Online resources enable unfettered access to and analysis of scientific data and are considered crucial for the advancement of modern science. Despite the clear power of online data resources, including web-available databases, proliferation can be problematic due to challenges in sustainability and long-term persistence. As areas of research become increasingly dependent on access to collections of data, an understanding of the scientific communitys capacity to develop and maintain such resources is needed.\n\nThe advent of the Internet coincided with expanding adoption of database technologies in the early 1990s, and the molecular biology community was at the forefront of using online databases to broadly disseminate data. The journal Nucleic Acids Research has long published articles dedicated to the description of online databases, as either debut or update articles. Snapshots throughout the entire history of online databases can be found in the pages of Nucleic Acids Research s \"Database Issue.\" Given the prominence of the Database Issue in the molecular biology and bioinformatics communities and the relative rarity of consistent historical documentation, database articles published in Database Issues provide a particularly unique opportunity for longitudinal analysis.\n\nTo take advantage of this opportunity, the study presented here first identifies each unique database described in 3055 Nucleic Acids Research Database Issue articles published between 1991-2016 to gather a rich dataset of databases debuted during this time frame, regardless of current availability. In total, 1727 unique databases were identified and associated descriptive statistics were gathered for each, including year debuted in a Database Issue and the number of all associated Database Issue publications and accompanying citation counts. Additionally, each database identified was assessed for current availability through testing of all associated URLs published. Finally, to assess maintenance, database websites were inspected to determine the last recorded update. The resulting work allows for an examination of the overall historical trends, such as the rate of database proliferation and attrition as well as an evaluation of citation metrics and on-going database maintenance.

molecular biology

Evaluation of Whole Exome Sequencing as an Alternative of BeadChip and Whole Genome Sequencing in Human Population Genetic Analysis

Understanding the underlying genetic structure of human populations is of fundamental interest to both biological and social sciences. Advances in high-throughput genotyping technology have markedly improved our understanding of global patterns of human genetic variation. The most widely used methods for collecting variant information at the DNA-level include whole genome sequencing, which continues to remain costly, and the more economical solution of array-based techniques, as these are capable of simultaneously genotyping a pre-selected set of variable DNA sites in the human genome. The largest publicly accessible set of human genomic sequence data available today originates from exome sequencing that comprises around 1.2% of the whole genome (approximately 30 million base pairs). In this study, we compared the application of the exome dataset to the array-based dataset and to the gold standard whole genome dataset using the same population genetic analysis methods. Our results draw attention to some of the inherent problems that arise from using pre-selected SNP sets for population genetic analysis. Additionally, we demonstrate that exome sequencing provides a better alternative to the array-based methods for population genetic analysis. In this study, we propose a strategy for unbiased variant collection from exome data and offer a bioinformatics protocol for proper data processing.

genomics

Reproducible integration of multiple sequencing datasets to form high-confidence SNP, indel, and reference calls for five human genome reference materials

Benchmark small variant calls from the Genome in a Bottle Consortium (GIAB) for the CEPH/HapMap genome NA12878 (HG001) have been used extensively for developing, optimizing, and demonstrating performance of sequencing and bioinformatics methods. Here, we develop a reproducible, cloud-based pipeline to integrate multiple sequencing datasets and form benchmark calls, enabling application to arbitrary human genomes. We use these reproducible methods to form high-confidence calls with respect to GRCh37 and GRCh38 for HG001 and 4 additional broadly-consented genomes from the Personal Genome Project that are available as NIST Reference Materials. These new genomes broad, open consent with few restrictions on availability of samples and data is enabling a uniquely diverse array of applications. Our new methods produce 17% more high-confidence SNPs, 176% more indels, and 12% larger regions than our previously published calls. To demonstrate that these calls can be used for accurate benchmarking, we compare other high-quality callsets to ours (e.g., Illumina Platinum Genomes), and we demonstrate that the majority of discordant calls are errors in the other callsets, We also highlight challenges in interpreting performance metrics when benchmarking against imperfect high-confidence calls. We show that benchmarking tools from the Global Alliance for Genomics and Health can be used with our calls to stratify performance metrics by variant type and genome context and elucidate strengths and weaknesses of a method.

genomics