bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

KRSA: Network-based Prediction of Differential Kinase Activity from Kinome Array Data

MotivationPhosphorylation by serine-threonine and tyrosine kinases is critical for determining protein function. Array-based approaches for measuring multiple kinases allow for the testing of differential phosphorylation between conditions for distinct sub-kinomes. While bioinformatics tools exist for processing and analyzing such kinome array data, current open-source tools lack the automated approach of upstream kinase prediction and network modeling. The presented tool, alongside other tools and methods designed for gene expression and protein-protein interaction network analyses, help the user better understand the complex regulation of gene and protein activities that forms biological systems and cellular signaling networks. ResultsWe present the Kinome Random Sampling Analyzer (KRSA), a web-application for kinome array analysis. While the underlying algorithm has been experimentally validated in previous publications, we tested the full KRSA application on dorsolateral prefrontal cortex (DLPFC) in male (n=3) and female (n=3) subjects to identify differential phosphorylation and upstream kinase activity. Kinase activity differences between males and females were compared to a previously published kinome dataset (11 female and 7 male subjects) which showed similar patterns to the global phosphorylation signal. Additionally, kinase hits were compared to gene expression databases for in silico validation at the transcript level and showed differential gene expression of kinases. Availability and implementationKRSA as a web-based application can be found at http://bpg-n.utoledo.edu:3838/CDRL/KRSA/. The code and data are available at https://github.com/kalganem/KRSA. Supplementary informationSupplementary data are available online.

bioinformatics↗

Mapping ribonucleotides embedded in genomic DNA to single-nucleotide resolution using Ribose-Map

Ribose-Map is a user-friendly, standardized bioinformatics toolkit for the comprehensive analysis of ribonucleotide sequencing experiments. It allows researchers to map the locations of ribonucleotides in DNA to single-nucleotide resolution and identify biological signatures of ribonucleotide incorporation. In addition, it can be applied to data generated using any currently available high-throughput ribonucleotide sequencing technique, thus standardizing the analysis of ribonucleotide sequencing experiments and allowing direct comparisons of results. This protocol describes in detail how to use Ribose-Map to analyze raw ribonucleotide sequencing data, including preparing the reads for analysis, locating the genomic coordinates of ribonucleotides, exploring the genome-wide distribution of ribonucleotides, determining the nucleotide sequence context of ribonucleotides, and identifying hotspots of ribonucleotide incorporation. Ribose-Map does not require background knowledge of ribonucleotide sequencing analysis and assumes only basic command-line skills. The protocol requires less than 3 hr of computing time for most datasets and about 30 min of hands-on time.

bioinformatics↗

starmapVR: immersive visualisation of single cell spatial omic data

MotivationAdvances in high throughput single-cell and spatial omic technologies have enabled the profiling of molecular expression and phenotypic properties of hundreds of thousands of individual cells in the context of their two dimensional (2D) or three dimensional (3D) spatial endogenous arrangement. However, current visualisation techniques do not allow for effective display and exploration of the single cell data in their spatial context. With the widespread availability of low-cost virtual reality (VR) gadgets, such as Google Cardboard, we propose that an immersive visualisation strategy is useful. ResultsWe present starmapVR, a light-weight, cross-platform, web-based tool for visualising single-cell and spatial omic data. starmapVR supports a number of interaction methods, such as keyboard, mouse, wireless controller and voice control. The tool visualises single cells in a 3D space and each cell can be represented by a star plot (for molecular expression, phenotypic properties) or image (for single cell imaging). For spatial transcriptomic data, the 2D single cell expression data can be visualised alongside the histological image in a 2.5D format. The application of starmapVR is demonstrated through a series of case studies. Its scalability has been carefully evaluated across different platforms. Availability and implementationstarmapVR is freely accessible at https://holab-hku.github.io/starmapVR, with the corresponding source code available at https://github.com/holab-hku/starmapVR under the open source MIT license. Supplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Antibody Upstream Sequences Diversity and Its Biological Implications Revealed by Repertoire Sequencing

The sequence upstream of antibody variable region (Antibody Upstream Sequence, or AUS) consists of 5 untranslated region (5 UTR) and two leader regions, L-PART1 and L-PART2. The sequence variations in AUS affect the efficiency of PCR amplification, mRNA translation, and subsequent PCR-based antibody quantification as well as antibody engineering. Despite their importance, the diversity of AUSs has long been neglected. Utilizing the rapid amplification of cDNA ends (5RACE) and high-throughput antibody repertoire sequencing (Rep-Seq) technique, we acquired full-length AUSs for human, rhesus macaque (RM), cynomolgus macaque (CM), mouse, and rat. We designed a bioinformatics pipeline and discovered 2,957 unique AUSs, corresponding to 2,786 and 1,159 unique sequences for 5 UTR and leader, respectively. Comparing with the leader records in the international ImMunoGeneTics (IMGT), while 529 were identical, 313 were with single nucleotide polymorphisms (SNPs), 280 were totally new, and 37 updated the incomplete records. The diversity of AUSs impact on related antibody biology was also probed. Taken together, our findings would facilitate Rep-Seq primer design for capturing antibodies comprehensively and efficiently as well as provide a valuable resource for antibody engineering and the studies of antibody at the molecular level.

bioinformatics↗

Boosting the analysis of protein interfaces with Multiple Interface String Alignment: illustration on the spikes of coronaviruses

We introduce Multiple Interface String Alignment (MISA), a visualization tool to display coherently various sequence and structure based statistics at protein-protein interfaces (SSE elements, buried surface area, {Delta}ASA, B factor values, etc). The amino-acids supporting these annotations are obtained from Voronoi interface models. The benefit of MISA is to collate annotated sequences of (homologous) chains found in different biological contexts i.e. bound with different partners or unbound. The aggregated views MISA/SSE, MISA/BSA, MISA/{Delta} ASAetc make it trivial to identify commonalities and differences between chains, to infer key interface residues, and to understand where conformational changes occur upon binding. As such, they should prove of key relevance for knowledge based annotations of protein databases such as the Protein Data Bank. Illustrations are provided on the receptor binding domain (RBD) of coronaviruses, in complex with their cognate partner or (neutralizing) antibodies. MISA computed with a minimal number of structures complement and enrich findings previously reported. The corresponding package is available from the Structural Bioinformatics Library (http://sbl.inria.fr)

bioinformatics↗

MiREDiBase: a manually curated database of editing events in microRNAs

MicroRNAs (miRNAs) are regulatory small non-coding RNAs that function as translational repressors. MiRNAs are involved in most cellular processes, and their expression and function are presided by several factors. Amongst, miRNA editing is an epitranscriptional modification that alters the original nucleotide sequence of selected miRNAs, possibly influencing their biogenesis and target-binding ability. A-to-I and C-to-U RNA editing are recognized as the canonical types, with the A-to-I type being the predominant one. Albeit some bioinformatics resources have been implemented to collect RNA editing data, it still lacks a comprehensive resource explicitly dedicated to miRNA editing. Here, we present MiREDiBase, a manually curated catalog of editing events in miRNAs. The current version includes 3,059 unique validated and putative editing sites from 626 pre-miRNAs in humans and three primates. Editing events in mature human miRNAs are supplied with miRNA-target predictions and enrichment analysis, while minimum free energy structures are inferred for edited pre-miRNAs. MiREDiBase represents a valuable tool for cell biology and biomedical research and will be continuously updated and expanded at https://ncrnaome.osumc.edu/miredibase.

bioinformatics↗

Metagenomics Strain Resolution on Assembly Graphs

We introduce a novel bioinformatics pipeline, STrain Resolution ON assembly Graphs (STRONG), which identifies strains de novo, when multiple metagenome samples from the same community are available. STRONG performs coassembly, followed by binning into metagenome assembled genomes (MAGs), but uniquely it stores the coassembly graph prior to simplification of variants. This enables the subgraphs for individual single-copy core genes (SCGs) in each MAG to be extracted. It can then thread back reads from the samples to compute per sample coverages for the unitigs in these graphs. These graphs and their unitig coverages are then used in a Bayesian algorithm, BayesPaths, that determines the number of strains present, their sequences or haplotypes on the SCGs and their abundances in each of the samples. Our approach both avoids the ambiguities of read mapping and allows more of the information on co-occurrence of variants in reads to be utilised than if variants were treated independently, whilst at the same time exploiting the correlation of variants across samples that occurs when they are linked in the same strain. We compare STRONG to the current state of the art on synthetic communities and demonstrate that we can recover more strains, more accurately, and with a realistic estimate of uncertainty deriving from the variational Bayesian algorithm employed for the strain resolution. On a real anaerobic digestor time series we obtained strain-resolved SCGs for over 300 MAGs that for abundant community members match those observed from long Nanopore reads.

bioinformatics↗

rSWeeP: A R/Bioconductor package deal with SWeeP sequences representation

The rSWeeP package is an R implementation of the SWeeP model, designed to handle Big Data. rSweeP meets to the growing demand for efficient methods of heuristic representation in the field of Bioinformatics, on platforms accessible to the entire scientific community. We explored the implementation of rSWeeP using a dataset containing 31,386 viral proteomes, performing phylogenetic and principal component analysis. As a case study we analyze the viral strains closest to the SARS-CoV, responsible for the current pandemic of COVID-19, confirming that rSWeeP can accurately classify organisms taxonomically. rSWeeP package is freely available at https://bioconductor.org/packages/release/bioc/html/rSWeeP.html.

bioinformatics↗

CIAlign - A highly customisable command line tool to clean, interpret and visualise multiple sequence alignments.

BackgroundThroughout biology, multiple sequence alignments (MSAs) form the basis of much investigation into biological features and relationships. These alignments are at the heart of many bioinformatics analyses. However, sequences in MSAs are often incomplete or very divergent, which leads to poorly aligned regions or large gaps in alignments. This slows down computation and can impact conclusions without being biologically relevant. Therefore, cleaning the alignment by removing these regions can substantially improve analyses. Manual editing of MSAs is very widespread but is time-consuming and difficult to reproduce. ResultsWe present a comprehensive, user-friendly MSA trimming tool with multiple visualisation options. Our highly customisable command line tool aims to give intervention power to the user by offering various options, and outputs graphical representations of the alignment before and after processing to give the user a clear overview of what has been removed. The main functionalities of the tool include removing regions of low coverage due to insertions, removing gaps, cropping poorly aligned sequence ends and removing sequences that are too divergent or too short. The thresholds for each function can be specified by the user and parameters can be adjusted to each individual MSA. CIAlign is designed with an emphasis on solving specific and common alignment problems and on providing transparency to the user. ConclusionCIAlign effectively removes problematic regions and sequences from MSAs and provides novel visualisation options. This tool can be used to refine alignments for further analysis and processing. The tool is aimed at anyone who wishes to automatically clean up parts of an MSA and those requiring a new, accessible way of visualising large MSAs.

bioinformatics↗

SeqRepo: A system for managing local collections biological sequences

MotivationAccess to biological sequence data, such as genome, transcript, or protein sequence, is at the core of many bioinformatics analysis workflows. The National Center for Biotechnology Information (NCBI), Ensembl, and other sequence database maintainers provide methods to access sequences through network connections. For many users, the convenience and currency of remotely managed data are compelling, and the network latency is non-consequential. However, for high-throughput and clinical applications, local sequence collections are essential for performance, stability, privacy, and reproducibility. ResultsHere we describe SeqRepo, a novel system for building a local, high-performance, non-redundant collection of biological sequences. SeqRepo enables clients to use primary database identifiers and several digests to identify sequences and sequence alises. SeqRepo provides a native Python interface and a REST interface, which can run locally and enables access from other programming languages. SeqRepo also provides an alternative REST interface based on the GA4GH refget protocol. SeqRepo provides fast random access to sequence slices. We provide results that demonstrate that a local SeqRepo sequence collection yields significant performance benefits of up to 1300-fold over remote sequence collections. In our use case for a variant validation and normalization pipeline, SeqRepo improved throughput 50-fold relative to use with remote sequences. SeqRepo may be used with any species or sequence type. Regular snapshots of Human sequence collections are available. It is often convenient or necessary to use a computed digest as a sequence identifier. For example, a digest-based identifier may be used to refer to proprietary reference genomes or segments of a graph genome, for which conventional identifiers will not be available. Here we also introduce a convention for the application of the SHA-512 hashing algorithm with Base64 encoding to generate URL-safe identifiers. This convention, sha512t24u, combines a fast digest mechanism with a space-efficient representation that can be used for any object. Our report includes an analysis of timing and collision probabilities for sha512t24u. SeqRepo enables clients to use sha512t24u as identifiers, thereby seamlessly integrating public and private sequence sets. AvailabilitySeqRepo is released under the Apache License 2.0 and is available on github and PyPi. Docker images and database snapshots are also available. See https://github.com/biocommons/biocommons.seqrepo.

bioinformatics↗

MetaFusion: A high-confidence metacaller for filtering and prioritizing RNA-seq gene fusion candidates.

MotivationCurrent fusion detection tools use diverse calling approaches and provide varying results, making selection of the appropriate tool challenging. Ensemble fusion calling techniques appear promising; however, current options have limited accessibility and function. ResultsMetaFusion is a flexible meta-calling tool that amalgamates outputs from any number of fusion callers. Individual caller results are standardized by conversion into the new file type Common Fusion Format (CFF). Calls are annotated, merged using graph clustering, filtered, and ranked to provide a final output of high confidence candidates. MetaFusion consistently achieves higher precision and recall than individual callers on real and simulated datasets, and reaches up to 100% precision, indicating that ensemble calling is imperative for high confidence results. MetaFusion uses FusionAnnotator to annotate calls with information from cancer fusion databases, and is provided with a benchmarking toolkit to calibrate new callers. AvailabilityMetaFusion is freely available at https://github.com/ccmbioinfo/MetaFusion Contactarun.ramani@sickkids.ca Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

An Advanced Framework for Time-lapse Microscopy Image Analysis

Time-lapse microscopy is a powerful technique that generates large volumes of image-based information to quantify the behaviors of cell populations. This method has been applied to cancer studies to estimate the drug response for precision medicine and has great potential to address inter-patient (or intertumoral) heterogeneity. A couple of algorithms exist to analyze time-lapse microscopy images; however, most deal with very high-resolution images involving few cells (typically cell lines). There are currently no advanced and efficient computational frameworks available to process large-scale time-lapse microscopy imaging data to estimate patient-specific response to therapy based on a large population of primary cells. In this paper, we propose a robust and user-friendly pipeline to preprocess the images and track the behaviors of thousands of cancer cells simultaneously for a better drug response prediction of cancer patients. Availability and ImplementationSource code is available at: https://github.com/CompbioLabUCF/CellTrack ACM Reference FormatQibing Jiang, Praneeth Sudalagunta, Mark B. Meads, Khandakar Tanvir Ahmed, Tara Rutkowski, Ken Shain, Ariosto S. Silva, and Wei Zhang. 2020. An Advanced Framework for Time-lapse Microscopy Image Analysis. In Proceedings of BioKDD: 19th International Workshop on Data Mining In Bioinformatics (BioKDD). ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/nnnnnnn.nnnnnnn

bioinformatics↗

Investigation of COVID-19 comorbidities reveals genes and pathways coincident with the SARS-CoV-2 viral disease

The emergence of the SARS-CoV-2 virus and subsequent COVID-19 pandemic initiated intense research into the mechanisms of action for this virus. It was quickly noted that COVID-19 presents more seriously in conjunction with other human disease conditions such as hypertension, diabetes, and lung diseases. We conducted a bioinformatics analysis of COVID-19 comorbidity-associated gene sets, identifying genes and pathways shared among the comorbidities, and evaluated current knowledge about these genes and pathways as related to current information about SARS-CoV-2 infection. We performed our analysis using GeneWeaver (GW), Reactome, and several biomedical ontologies to represent and compare common COVID-19 comorbidities. Phenotypic analysis of shared genes revealed significant enrichment for immune system phenotypes and for cardiovascular-related phenotypes, which might point to alleles and phenotypes in mouse models that could be evaluated for clues to COVID-19 severity. Through pathway analysis, we identified enriched pathways shared by comorbidity datasets and datasets associated with SARS-CoV-2 infection.

bioinformatics↗

Dynamic Analysis of Alternative Polyadenylation from Single-Cell RNA-Seq(scDaPars) Reveals Cell Subpopulations Invisible to Gene Expression Analysis

Alternative polyadenylation (APA) is a major mechanism of post-transcriptional regulation in various cellular processes including cell proliferation and differentiation, but the APA heterogeneity among single cells remains largely unknown. Single-cell RNA sequencing (scRNA-seq) has been extensively used to define cell subpopulations at the transcription level. Yet, most scRNA-seq data have not been analyzed in an "APA-aware" manner. Here, we introduce scDaPars, a bioinformatics algorithm to accurately quantify APA events at both single-cell and single-gene resolution using standard scRNA-seq data. Validations in both real and simulated data indicate that scDaPars can robustly recover missing APA events caused by the low amounts of mRNA sequenced in single cells. When applied to cancer and human endoderm differentiation data, scDaPars not only revealed cell-type-specific APA regulation but also identified cell subpopulations that are otherwise invisible to conventional gene expression analysis. Thus, scDaPars will enable us to understand cellular heterogeneity at the post-transcriptional APA level.

bioinformatics↗

Insights into the mechanism of bovine spermiogenesis based on comparative transcriptomic studies

To reduce the reproductive loss caused by semen quality and provide theoretical guidance for the eradication of human male infertility, differential analysis of the bovine transcriptome among round spermatids, elongated spermatids, and epididymal sperm was carried out with the reference of the mouse transcriptome, and the homology trends of gene expression to the mouse were also analysed. First, to explore the physiological mechanism of spermiogenesis that profoundly affects semen quality, homological trends of differential genes were compared during spermiogenesis in dairy cattle and mice. Next, the Gene ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment, protein-protein interaction network (PPI network), and bioinformatics analysis uncovered the regulation network of acrosome formation during the transition from round to elongated spermatids. In addition, processes that regulate gene expression during spermiogenesis from elongated spermatid to epididymal sperm, such as ubiquitination, acetylation, deacetylation, glycosylation, and the functional gene ART3 may play an important role during spermiogenesis. Therefore, its localisation in the seminiferous tubule was investigated by immunofluorescent analysis, and its structure and function were also predicted. This study provides important data for revealing the mystery of life during spermiogenesis resulting from acrosome formation, histone replacement, and the fine regulation of gene expression.

bioinformatics↗

Inductive Inference of Gene Regulatory Network Using Supervised and Semi-supervised Graph Neural Networks

Discovering gene regulatory relationships and reconstructing gene regulatory networks (GRN) based on gene expression data is a classical, long-standing computational challenge in bioinformatics. Computationally inferring a possible regulatory relationship between two genes can be formulated as a link prediction problem between two nodes in a graph. Graph neural network (GNN) provides an opportunity to construct GRN by integrating topological neighbor propagation through the whole gene network. We propose an end-to-end gene regulatory graph neural network (GRGNN) approach to reconstruct GRNs from scratch utilizing the gene expression data, in both a supervised and a semi-supervised framework. To get better inductive generalization capability, GRN inference is formulated as a graph classification problem, to distinguish whether a subgraph centered at two nodes contains the link between the two nodes. A linked pair between a transcription factor (TF) and a target gene, and their neighbors are labeled as a positive subgraph, while an unlinked TF and target gene pair and their neighbors are labeled as a negative subgraph. A GNN model is constructed with node features from both explicit gene expression and graph embedding. We demonstrate a noisy starting graph structure built from partial information, such as Pearsons correlation coefficient and mutual information can help guide the GRN inference through an appropriate ensemble technique. Furthermore, a semi-supervised scheme is implemented to increase the quality of the classifier. When compared with established methods, GRGNN achieved state-of-the-art performance on the DREAM5 GRN inference benchmarks. GRGNN is publicly available at https://github.com/juexinwang/GRGNN. HighlightsO_LIWe present a novel formulation of graph classification in inferring gene regulatory relationships from gene expression and graph embedding. C_LIO_LIOur method leverages a powerful framework, gene regulatory graph neural network (GRGNN), which is flexible and powerful to ensemble statistical powers from a number of heuristic skeletons. C_LIO_LIOur results show GRGRNN outperforms previous supervised and unsupervised methods inductively on benchmarks. C_LIO_LIGRGNN can be interpreted and explained following the biological network motif hypothesis in gene regulatory networks. C_LI

bioinformatics↗

TALE: Transformer-based protein function Annotation with joint sequence-Label Embedding

MotivationFacing the increasing gap between high-throughput sequence data and limited functional insights, computational protein function annotation provides a high-throughput alternative to experimental approaches. However, current methods can have limited applicability while relying on data besides sequences, or lack generalizability to novel sequences, species and functions. ResultsTo overcome aforementioned barriers in applicability and generalizability, we propose a novel deep learning model, named Transformer-based protein function Annotation through joint sequence-Label Embedding (TALE). For generalizbility to novel sequences we use self attention-based transformers to capture global patterns in sequences. For generalizability to unseen or rarely seen functions, we also embed protein function labels (hierarchical GO terms on directed graphs) together with inputs/features (sequences) in a joint latent space. Combining TALE and a sequence similarity-based method, TALE+ outperformed competing methods when only sequence input is available. It even outperformed a state-of-the-art method using network information besides sequence, in two of the three gene ontologies. Furthermore, TALE and TALE+ showed superior generalizability to proteins of low homology and never/rarely annotated novel species or functions compared to training data, revealing deep insights into the protein sequence-function relationship. Ablation studies elucidated contributions of algorithmic components toward the accuracy and the generalizability. AvailabilityThe data, source codes and models are available at https://github.com/Shen-Lab/TALE Contactyshen@tamu.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Boolean Implication Analysis Improves Prediction Accuracy of In Silico Gene Reporting of Retinal Cell Types

The retina is a complex tissue containing multiple cell types that is essential for vision. Understanding the gene expression patterns of various retinal cell types has potential applications in regenerative medicine. Retinal organoids (optic vesicles) derived from pluripotent stem cells have begun to yield insights into the transcriptomics of developing retinal cell types in humans through single cell RNA-sequencing studies. Previous methods of gene reporting have relied upon techniques in vivo using microarray data, or correlational and dimension reduction methods for analyzing single cell RNA-sequencing data in silico. Here, we present a bioinformatic approach using Boolean implication to discover retinal cell type-specific genes. We apply this approach to previously published retina and retinal organoid datasets and improve upon previously published correlational methods. Our method improves the prediction accuracy and reproducibility of marker genes of retinal cell types and discovers several new high confidence cone and rod-specific genes. Furthermore, our method is general and can impact all areas of gene expression analyses in cancer and other human diseases. Significance StatementEfforts to derive retinal cell types from pluripotent stem cells to the end of curing retinal disease require robust characterization of these cell types gene expression patterns. The Boolean method described in this study improves prediction accuracy of earlier methods of gene reporting, and allows for the discovery and validation of retinal cell type-specific marker genes. The invariant nature of results from Boolean implication analysis can yield high-value molecular markers that can be used as biomarkers or drug targets. O_FIG_DISPLAY_L [Figure 1] M_FIG_DISPLAY C_FIG_DISPLAY

bioinformatics↗