bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Proximity Measures as Graph Convolution Matrices for Link Prediction in Biological Networks

MotivationLink prediction is an important and well-studied problem in computational biology, with a broad range of applications including disease gene prioritization, drug-disease associations, and drug response in cancer. The general principle in link prediction is to use the topological characteristics and the attributes-if available- of the nodes in the network to predict new links that are likely to emerge/disappear. Recently, graph representation learning methods, which aim to learn a low-dimensional representation of topological characteristics and the attributes of the nodes, have drawn increasing attention to solve the link prediction problem via learnt low-dimensional features. Most prominently, Graph Convolution Network (GCN)-based network embedding methods have demonstrated great promise in link prediction due to their ability of capturing non-linear information of the network. To date, GCN-based network embedding algorithms utilize a Laplacian matrix in their convolution layers as the convolution matrix and the effect of the convolution matrix on algorithm performance has not been comprehensively characterized in the context of link prediction in biomedical networks. On the other hand, for a variety of biomedical link prediction tasks, traditional node similarity measures such as Common Neighbor, Ademic-Adar, and other have shown promising results, and hence there is a need to systematically evaluate the node similarity measures as convolution matrices in terms of their usability and potential to further the state-of-the-art. ResultsWe select 8 representative node similarity measures as convolution matrices within the single-layered GCN graph embedding method and conduct a systematic comparison on 3 important biomedical link prediction tasks: drug-disease association (DDA) prediction, drug-drug interaction (DDI) prediction, protein-protein interaction (PPI) prediction. Our experimental results demonstrate that the node similarity-based convolution matrices significantly improves GCN-based embedding algorithms and deserve more attention in the future biomedical link prediction AvailabilityOur method is implemented as a python library and is available at githublink Contactmustafa.coskun@agu.edu.tr Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Set-Min sketch: a probabilistic map for power-law distributions with application to k-mer annotation

AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSMotivationC_ST_ABSIn many bioinformatics pipelines, k-mer counting is often a required step, with existing methods focusing on optimizing time or memory usage. These methods usually produce very large count tables explicitly representing k-mers themselves. Solutions avoiding explicit representation of k-mers include Minimal Perfect Hash Functions (MPHFs) or Count-Min sketches. The former is only applicable to static maps not subject to updates, while the latter suffers from potentially very large point-query errors, making it unsuitable when counters are required to be highly accurate. ResultsWe introduce Set-Min sketch - a sketching technique for representing associative maps inspired by Count-Min sketch - and apply it to the problem of representing k-mer count tables. Set-Min is provably more accurate than both Count-Min and Max-Min - an improved variant of Count-Min for static datasets that we define here. We show that Set-Min sketch provides a very low error rate, both in terms of the probability and the size of errors, at the expense of a very moderate memory increase. On the other hand, Set-Min sketches are shown to take up to an order of magnitude less space than MPHF-based solutions, especially for large values of k. Space-efficiency of Set-Min takes advantage of the power-law distribution of k-mer counts in genomic datasets. Availabilityhttps://github.com/yhhshb/fress

bioinformatics↗

Adding software to package management systems can increase their citation by 280%

A growing number of biomedical methods and protocols are being disseminated as open-source software packages. When put in concert with other packages, they can execute in-depth and comprehensive computational pipelines. Therefore, their integration with other software packages plays a prominent role in their adoption in addition to their availability. Accordingly, package management systems are developed to standardize the discovery and integration of software packages. Here we study the impact of package management systems on software dissemination and their scholarly recognition. We study the citation pattern of more than 18,000 scholarly papers referenced by more than 23,000 software packages hosted by Bioconda, Bioconductor, BioTools, and ToolShed--the package management systems primarily used by the Bioinformatics community. Our results suggest that there is significant evidence that the scholarly papers citation count increases after their respective software was published to package management systems. Additionally, our results show that the impact of different package management systems on the scholarly papers recognition is of the same magnitude. These results may motivate scientists to distribute their software via package management systems, facilitating the composition of computational pipelines and helping reduce redundancy in package development. Significance StatementSoftware packages are the building blocks of computational pipelines. A myriad of packages are developed; however, the lack of integration and discovery standards hinders their adoption, leaving most scientists scholarly contributions unrecognized. Package management systems are developed to facilitate software dissemination and integration. However, developing software to meet their code and packaging standards is an involved process. Therefore, our study results on the significant impact of the package management systems on scholarly papers recognition can motivate scientists to invest in disseminating their software via package management systems. Dissemination of more software via package management systems will lead to a more straightforward composition of computational pipelines and less redundancy in software packages.

bioinformatics↗

Tysserand - Fast reconstruction of spatial networks from bioimages

SummaryNetworks provide a powerful framework to analyze spatial omics experiments. However, we lack tools that integrate several methods to easily reconstruct networks for further analyses with dedicated libraries. In addition, choosing the appropriate method and parameters can be challenging. We propose tysserand, a Python library to reconstruct spatial networks from spatially resolved omics experiments. It is intended as a common tool to which the bioinformatics community can add new methods to reconstruct networks, choose appropriate parameters, clean resulting networks and pipe data to other libraries. Availability and implementationtysserand software and tutorials with a Jupyter notebook to reproduce the results are available at https://github.com/VeraPancaldiLab/tysserand Supplementary informationSupplementary data are available at Bioarxiv online.

bioinformatics↗

BOSO: a novel feature selection algorithm for linear regression with high-dimensional data

MotivationWith the frenetic growth of high-dimensional datasets in different biomedical domains, there is an urgent need to develop predictive methods able to deal with this complexity. Feature selection is a relevant strategy in machine learning to address this challenge. ResultsWe introduce a novel feature selection algorithm for linear regression called BOSO (Bilevel Optimization Selector Operator). We conducted a benchmark of BOSO with key algorithms in the literature, finding a superior performance in highdimensional datasets. Proof-of-concept of BOSO for predicting drug sensitivity in cancer is presented. A detailed analysis is carried out for methotrexate, a well-studied drug targeting cancer metabolism. AvailabilityA Matlab implementation of BOSO is available as a Supplementary Material. Contactfplanes@tecnun.es Supplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Frameshift and frame-preserving mutations in zebrafish presenilin 2 affect different cellular functions in young adult brains

BackgroundMutations in PRESENILIN 2 (PSEN2) cause early disease onset familial Alzheimers disease (EOfAD) but their mode of action remains elusive. One consistent observation for all PRESENILIN gene mutations causing EOfAD is that a transcript is produced with a reading frame terminated by the normal stop codon - the "reading frame preservation rule". Mutations that do not obey this rule do not cause the disease. The reasons for this are debated. MethodsA frameshift mutation (psen2N140fs) and a reading frame-preserving mutation (psen2T141_L142delinsMISLISV) were previously isolated during genome editing directed at the N140 codon of zebrafish psen2 (equivalent to N141 of human PSEN2). We mated a pair of fish heterozygous for each mutation to generate a family of siblings including wild type and heterozygous mutant genotypes. Transcriptomes from young adult (6 months) brains of these genotypes were analysed. Bioinformatics techniques were used to predict cellular functions affected by heterozygosity for each mutation. ResultsThe reading frame preserving mutation uniquely caused subtle, but statistically significant, changes to expression of genes involved in oxidative phosphorylation, long term potentiation and the cell cycle. The frameshift mutation uniquely affected genes involved in Notch and MAPK signalling, extracellular matrix receptor interactions and focal adhesion. Both mutations affected ribosomal protein gene expression but in opposite directions. ConclusionA frameshift and frame-preserving mutation at the same position in zebrafish psen2 cause discrete effects. Changes in oxidative phosphorylation, long term potentiation and the cell cycle may promote EOfAD pathogenesis in humans.

bioinformatics↗

ProCbA: Protein Function Prediction based on Clique Analysis

Protein function prediction based on protein-protein interactions (PPI) is one of the most important challenges of the Post-Genomic era. Due to the fact that determining protein function by experimental techniques can be costly, function prediction has become an important challenge for computational biology and bioinformatics. Some researchers utilize graph- (or network-) based methods using PPI networks for un-annotated proteins. The aim of this study is to increase the accuracy of the protein function prediction using two proposed methods. To predict protein functions, we propose a Protein Function Prediction based on Clique Analysis (ProCbA) and Protein Function Prediction on Neighborhood Counting using functional aggregation (ProNC-FA). Both ProCbA and ProNC-FA can predict the functions of unknown proteins. In addition, in ProNC-FA which is not including new algorithm; we try to address the essence of incomplete and noisy data of PPI era in order to achieving a network with complete functional aggregation. The experimental results on MIPS data and the 17 different explained datasets validate the encouraging performance and the strength of both ProCbA and ProNC-FA on function prediction. Experimental result analysis as can be seen in Section IV, the both ProCbA and ProNC-FA are generally able to outperform all the other methods.

bioinformatics↗

HierCC: A multi-level clustering scheme for population assignments based on core genome MLST

MotivationRoutine infectious disease surveillance is increasingly based on large-scale whole genome sequencing databases. Real-time surveillance would benefit from immediate assignments of each genome assembly to hierarchical population structures. Here we present HierCC, a scalable clustering scheme based on core genome multi-locus typing that allows incremental, static, multi-level cluster assignments of genomes. We also present HCCeval, which identifies optimal thresholds for assigning genomes to cohesive HierCC clusters. HierCC was implemented in EnteroBase in 2018, and has since genotyped >400,000 genomes from Salmonella, Escherichia, Yersinia and Clostridioides. AvailabilityImplementation: http://enterobase.warwick.ac.uk/ and Source codes: https://github.com/zheminzhou/HierCC Contactzhemin.zhou@warwick.ac.uk Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

EpitopeVec: Linear Epitope Prediction Using DeepProtein Sequence Embeddings

MotivationB-cell epitopes (BCEs) play a pivotal role in the development of peptide vaccines, immunodiagnostic reagents, and antibody production, and thus generally in infectious disease prevention and diagnosis. Experimental methods used to determine BCEs are costly and time-consuming. It thus becomes essential to develop computational methods for the rapid identification of BCEs. Though several computational methods have been developed for this task, cross-testing of classifiers trained and tested on different datasets revealed their limitations, with accuracies of 51 to 53%. ResultsWe describe a new method called EpitopeVec, which utilizes residue properties, modified antigenicity scales, and a Protvec representation of peptides for linear BCE prediction with machine learning techniques. Evaluating on several large and small data sets, as well as cross-testing demonstrated an improvement of the state-of-the-art performances in terms of accuracy and AUC. Predictive performance depended on the type of antigen (viral, bacterial, eukaryote, etc.). In view of that, we also trained our method on a large viral dataset to create a linear viral BCE predictor. AvailablityThe software is available at https://github.com/hzi-bifo/epitope-prediction under the GPL3.0 license. Contactalice.mchardy@helmholtz-hzi.de Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Unheeded SARS-CoV-2 protein? Look deep into negative-sense RNA

SARS-CoV-2 is a novel positive-sense single-stranded RNA virus from the Coronaviridae family (genus Betacoronavirus), which has been established as causing the COVID-19 pandemic. The genome of SARS-CoV-2 is one of the largest among known RNA viruses, comprising of at least 26 known protein-coding loci. Studies thus far have outlined the coding capacity of the positive-sense strand of the SARS-CoV-2 genome, which can be used directly for protein translation. However, it has been recently shown that transcribed negative-sense viral RNA intermediates that arise during viral genome replication from positive-sense viruses can also code for proteins. No studies have yet explored the potential for negative-sense SARS-CoV-2 RNA intermediates to contain protein coding-loci. Thus, using sequence and structure-based bioinformatics methodologies, we have investigated the presence and validity of putative negative-sense ORFs (nsORFs) in the SARS-CoV-2 genome. Nine nsORFs were discovered to contain strong eukaryotic translation initiation signals and high codon adaptability scores, and several of the nsORFs were predicted to interact with RNA-binding proteins. Evolutionary conservation analyses indicated that some of the nsORFs are deeply conserved among related coronaviruses. Three-dimensional protein modelling revealed the presence of higher order folding among all putative SARS-CoV-2 nsORFs, and subsequent structural mimicry analyses suggest similarity of the nsORFs to DNA/RNA-binding proteins and proteins involved in immune signaling pathways. Altogether, these results suggest the potential existence of still undescribed SARS-CoV-2 proteins, which may play an important role in the viral lifecycle and COVID-19 pathogenesis. Contactpetr.pecinka@osu.cz; tlb20@cam.ac.uk

bioinformatics↗

REALDIST: Real-valued protein distance prediction

Protein structure prediction continues to stand as an unsolved problem in bioinformatics and biomedicine. Deep learning algorithms and the availability of metagenomic sequences have led to the development of new approaches to predict inter-residue distances--the key intermediate step. Different from the recently successful methods which frame the problem as a multi-class classification problem, this article introduces a real-valued distance prediction method REALDIST. Using a representative set of 43 thousand protein chains, a variant of deep ResNet is trained to predict real-valued distance maps. The contacts derived from the real-valued distance maps predicted by this method, on the most difficult CASP13 free-modeling protein datasets, demonstrate a long-range top-L precision of 52%, which is 17% higher than the top CASP13 predictor Raptor-X and slightly higher than the more recent trRosetta method. Similar improvements are observed on the CAMEO hard and very hard datasets. Three-dimensional (3D) structure prediction guided by real-valued distances reveals that for short proteins the mean accuracy of the 3D models is slightly higher than the top human predictor AlphaFold and server predictor Quark in the CASP13 competition.

bioinformatics↗

Predicting and validating protein degradation in proteomes using deep learning

Age, disease, and exposure to environmental factors can induce tissue remodelling and alterations in protein structure and abundance. In the case of human skin, ultraviolet radiation (UVR)-induced photo-ageing has a profound effect on dermal extracellular matrix (ECM) proteins. We have previously shown that ECM proteins rich in UV-chromophore amino acids are differentially susceptible to UVR. However, this UVR-mediated mechanism alone does not explain the loss of UV-chromophore-poor assemblies such as collagen. Here, we aim to develop novel bioinformatics tools to predict the relative susceptibility of human skin proteins to not only UVR and photodynamically produced ROS but also to endogenous proteases. We test the validity of these protease cleavage site predictions against experimental datasets (both previously published and our own, derived by exposure of either purified ECM proteins or a complex cell-derived proteome, to matrix metalloproteinase [MMP]-9). Our deep Bidirectional Recurrent Neural Network (BRNN) models for cleavage site prediction in nine MMPs, four cathepsins, elastase-2, and granzyme-B perform better than existing models when validated against both simple and complex protein mixtures. We have combined our new BRNN protease cleavage prediction models with predictions of relative UVR/ROS susceptibility (based on amino acid composition) into the Manchester Proteome Susceptibility Calculator (MPSC) webapp http://www.manchesterproteome.manchester.ac.uk/#/MPSC (or http://130.88.96.141/#/MPSC). Application of the MPSC to the dermal proteome suggests that fibrillar collagens and elastic fibres will be preferentially degraded by proteases alone and by UVR/ROS and protease in combination, respectively. We also identify novel targets of oxidative damage and protease activity including dermatopontin (DPT), fibulins (EFEMP-1,-2, FBLN-1,-2,-5), defensins (DEFB1, DEFA3, DEFA1B, DEFB4B), proteases and protease inhibitors themselves (CTSA, CTSB, CTSZ, CTSD, TIMPs-1,-2,-3, SPINK6, CST6, PI3, SERPINF1, SERPINA-1,-3,-12). The MPSC webapp has the potential to identify novel protein biomarkers of tissue damage and to aid the characterisation of protease degradomics leading to improved identification of novel therapeutic targets.

bioinformatics↗

ramr: an R package for detection of rare aberrantly methylated regions

MotivationWith recent advances in the field of epigenetics, the focus is widening from large and frequent disease- or phenotype-related methylation signatures to rare alterations transmitted mitotically or transgenerationally (constitutional epimutations). Merging evidence indicate that such constitutional alterations, albeit occurring at a low mosaic level, may confer risk of disease later in life. Given their inherently low incidence rate and mosaic nature, there is a need for bioinformatic tools specifically designed to analyse such events. ResultsWe have developed a method (ramr) to identify aberrantly methylated DNA regions (AMRs). ramr can be applied to methylation data obtained by array or next-generation sequencing techniques to discover AMRs being associated with elevated risk of cancer as well as other diseases. We assessed accuracy and performance metrics of ramr and confirmed its applicability for analysis of large public data sets. Using ramr we identified aberrantly methylated regions that are known or may potentially be associated with development of colorectal cancer and provided functional annotation of AMRs that arise at early developmental stages. Availability and implementationThe R package is freely available at https://github.com/BBCG/ramr

bioinformatics↗

Phylogenetic Novelty Scores: a New Approach for Weighting Genetic Sequences

BackgroundMany important applications in bioinformatics, including sequence alignment and protein family profiling, employ sequence weighting schemes to mitigate the effects of non-independence of homologous sequences and under- or over-representation of certain taxa in a dataset. These schemes aim to assign high weights to sequences that are novel compared to the others in the same dataset, and low weights to sequences that are over-represented. ResultsWe formalise this principle by rigorously defining the evolutionary novelty of a sequence within an alignment. This results in new sequence weights that we call phylogenetic novelty scores. These scores have various desirable properties, and we showcase their use by considering, as an example application, the inference of character frequencies at an alignment column -- important, for example, in protein family profiling. We give computationally efficient algorithms for calculating our scores and, using simulations, show that they improve the accuracy of character frequency estimation compared to existing sequence weighting schemes. ConclusionsOur phylogenetic novelty scores can be useful when an evolutionarily meaningful system for adjusting for uneven taxon sampling is desired. They have numerous possible applications, including estimation of evolutionary conservation scores and sequence logos, identification of targets in conservation biology, and improving and measuring sequence alignment accuracy.

bioinformatics↗

Easyreporting simplifies the implementation of Reproducible Research Layers in R software

During last years "irreproducibility" became a general problem in omics data analysis due to the use of sophisticated and poorly described computational procedures. For avoiding misleading results, it is necessary to inspect and reproduce the entire data analysis as a unified product. Reproducible Research (RR) provides general guidelines for public access to the analytic data and related analysis code combined with natural language documentation, allowing third-parties to reproduce the findings. We developed easyreporting, a novel R/Bioconductor package, to facilitate the implementation of an RR layer inside reports/tools without requiring any knowledge of the R Markdown language. We describe the main functionalities and illustrate how to create an analysis report using a typical case study concerning the analysis of RNA-seq data. Then, we also show how to trace R functions automatically. Thanks to this latter feature, easyreporting results beneficial for developers to implement procedures that automatically keep track of the analysis steps within Graphical User Interfaces (GUIs). Easyreporting can be useful in supporting the reproducibility of any data analysis project and the implementation of GUIs. It turns out to be very helpful in bioinformatics, where the complexity of the analyses makes it extremely difficult to trace all the steps and parameters used in the study.

bioinformatics↗

Uncovering novel mutational signatures by de novo extraction with SigProfilerExtractor

Mutational signature analysis is commonly performed in genomic studies surveying cancer and normal somatic tissues. Here we present SigProfilerExtractor, an automated tool for accurate de novo extraction of mutational signatures for all types of somatic mutations. Benchmarking with a total of 34 distinct scenarios encompassing 2,500 simulated signatures operative in more than 60,000 unique synthetic genomes and 20,000 synthetic exomes demonstrates that SigProfilerExtractor outperforms thirteen other tools across all datasets with and without noise. For genome simulations with 5% noise, reflecting high-quality genomic datasets, SigProfilerExtractor outperforms other approaches by elucidating between 20% and 50% more true positive signatures while yielding more than 5-fold less false positive signatures. Applying SigProfilerExtractor to 4,643 whole-genome sequenced and 19,184 whole-exome sequenced cancers reveals four previously missed mutational signatures. Two of the signatures are confirmed in independent cohorts with one of these signatures associating with tobacco smoking. In summary, this report provides a reference tool for analysis of mutational signatures, a comprehensive benchmarking of bioinformatics tools for extracting mutational signatures, and several novel mutational signatures including a signature putatively attributed to direct tobacco smoking mutagenesis in bladder cancer and in normal bladder epithelium.

bioinformatics↗

BiG-MAP: an automated pipeline to profile metabolic gene cluster abundance and expression in microbiomes

Microbial gene clusters encoding the biosynthesis of primary and secondary metabolites play key roles in shaping microbial ecosystems and driving microbiome-associated phenotypes. Although effective approaches exist to evaluate the metabolic potential of such bacteria through identification of metabolic gene clusters in their genomes, no automated pipelines exist to profile the abundance and expression levels of such gene clusters in microbiome samples to generate hypotheses about their functional roles and to find associations with phenotypes of interest. Here, we describe BiG-MAP, a bioinformatic tool to profile abundance and expression levels of gene clusters across metagenomic and metatranscriptomic data and evaluate their differential abundance and expression between different conditions. To illustrate its usefulness, we analyzed 47 metagenomic samples from healthy and caries-associated human oral microbiome samples and identified 58 gene clusters, including unreported ones, that were significantly more abundant in either phenotype. Among them, we found the muc operon, a gene cluster known to be associated to tooth decay. Additionally, we found a putative reuterin biosynthetic gene cluster from a Streptococcus strain to be enriched but not exclusively found in healthy samples; metabolomic data from the same samples showed masses with fragmentation patterns consistent with (poly)acrolein, which is known to spontaneously form from the products of the reuterin pathway and has been previously shown to inhibit pathogenic Streptococcus mutans strains. Thus, we show how BiG-MAP can be used to generate new hypotheses on potential drivers of microbiome-associated phenotypes and prioritize the experimental characterization of relevant gene clusters that may mediate them. ImportanceMicrobes play an increasingly recognized role in determining host-associated phenotypes by producing small molecules that interact with other microorganisms or host cells. The production of these molecules is often encoded in syntenic genomic regions, also known as gene clusters. With the increasing numbers of (multi-)omics datasets that can help understanding complex ecosystems at a much deeper level, there is a need to create tools that can automate the process of analyzing these gene clusters across omics datasets. The current study presents a new software tool called BiG-MAP, which allows assessing gene cluster abundance and expression in microbiome samples using metagenomic and metatranscriptomic data. In this manuscript, we describe the tool and its functionalities, and how it has been validated using a mock community. Finally, using an oral microbiome dataset, we show how it can be used to generate hypotheses regarding the functional roles of gene clusters in mediating host phenotypes.

bioinformatics↗

ViralRecall: A Flexible Command-Line Tool for the Detection of Giant Virus Signatures in Omic Data

Giant viruses are widespread in the biosphere and play important roles in biogeochemical cycling and host genome evolution. Also known as Nucleo-Cytoplasmic Large DNA Viruses (NCLDV), these eukaryotic viruses harbor the largest and most complex viral genomes known. Recent studies have shown that NCLDV are frequently abundant in metagenomic datasets, and that sequences derived from these viruses can also be found endogenized in diverse eukaryotic genomes. The accurate detection of sequences derived from NCLDV is therefore of great importance, but this task is challenging owing to both the high level of sequence divergence between NCLDV families and the extraordinarily high diversity of genes encoded in their genomes, including some encoding for metabolic or translation-related functions that are typically found only in cellular lineages. Here we present ViralRecall, a bioinformatic tool for the identification of NCLDV signatures in omic data. This tool leverages a library of Giant Virus Orthologous Groups (GVOGs) to identify sequences that bear signatures of NCLDV. We demonstrate that this tool can effectively identify NCLDV sequences with high sensitivity and specificity. Moreover, we show that it can be useful both for removing contaminating sequences in metagenome-assembled viral genomes as well as the identification of eukaryotic genomic loci that derived from NCLDV. ViralRecall is written in Python 3.5 and is freely available on GitHub: https://github.com/faylward/viralrecall.

bioinformatics↗