bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

BeEM: fast and faithful conversion of mmCIF format structure files to PDB format

Although mmCIF is the current official format for deposition of protein and nucleic acid structures to the Protein Data Bank (PDB) database, the legacy PDB format is still the primary supported format for many structural bioinformatics tools. Therefore, reliable software to convert mmCIF structure files to PDB files is needed. Unfortunately, existing conversion programs fail to correctly convert many mmCIF files, especially those with many atoms and/or long chain identifies. This study proposed BeEM, which converts any mmCIF format structure files to PDB format. BeEM conversion faithfully retains all atomic and chain information, including chain IDs with more than 2 characters, which are not supported by any existing mmCIF to PDB converters. The conversion speed of BeEM is at least ten times faster than existing converters such as MAXIT and Phenix. BeEM is available under the BSD licence at https://github.com/kad-ecoli/BeEM/.

bioinformatics↗

Inferring a cell's capabilities from omics data with ImmCellFie

ImmCellFie is a user-friendly, web-based platform for comprehensive analysis of metabolic functions inferred from transcriptomic or proteomic data. It enables researchers to leverage the powerful mechanistic insight provided by complex genome-scale metabolic models with little to no bioinformatics training required. The platform has been integrated with a series of useful tools and richly annotated scientific visualizations for interactive exploration by the user. ImmCellFie pushes beyond simple statistical enrichment and incorporates complex biological mechanisms to quantify cell activity. Graphical abstract

bioinformatics↗

An Automatized Workflow to Study Mechanistic Indicators for Driver Gene Prediction with Moonlight

Prediction of tumor suppressors and oncogenes, also called driver genes, is an essential step in understanding cancer development and discovering potential novel treatments. We recently proposed Moonlight as a bioinformatics framework to predict driver genes and analyze them in a system-biology-oriented manner based on -omics integration. Moonlight uses gene expression as a primary data source and combines it with patterns related to cancer hallmarks and regulatory networks to identify oncogenic mediators. Once the oncogenic mediators are identified, it is important to include extra levels of evidence, called mechanistic indicators, to identify driver genes and to link the observed changes in gene expression to the underlying alteration that promotes them. Such a mechanistic indicator could be for example a mutation in the regulatory regions for the candidate gene or mutations in the regulator itself. In this work, we developed new functionalities and release Moonlight2, to provide the user with the mutation-based mechanistic indicator to streamline the analyses of this second layer of evidence. The function analyzes mutation information in a cancer cohort to classify them into driver and passenger mutations. Moreover, the function estimates the potential effect of a mutation on the transcriptional, translational, or protein structure/function level. Those oncogenic mediators with at least one driver mutation are retained as the final set of driver genes. We applied Moonlight2 and the newly developed function to a case study on Basal-like breast cancer subtype using data from The Cancer Genome Atlas. We found six oncogenes (SF3B4, EBNA1BP2, KRTCAP2, ZBTB8OS, RUNX2, and POLR2J) and ten tumor suppressor genes (KIF26B, NR5A2, ARHGAP25, EMCN, ARL15, PCOLCE, TPK1, TEK, KIR2DL4, and GMFG) containing a driver mutation in their promoter region, possibly explaining their deregulation. The Moonlight2R source code is available at https://github.com/ELELAB/Moonlight2R.

bioinformatics↗

Optimization-Based Decoding of Imaging Spatial Transcriptomics Data

MotivationImaging Spatial Transcriptomics (iST) techniques characterize gene expression in cells in their native context by imaging barcoded probes for mRNA with single molecule resolution. However, the need to acquire many rounds of high-magnification imaging data limits the throughput and impact of existing methods. ResultsWe describe the Joint Sparse method for Imaging Transcriptomics (JSIT), an algorithm for decoding lower magnification IT data than that used in standard experimental workflows. JSIT incorporates codebook knowledge and sparsity assumptions into an optimization problem which is less reliant on well separated optical signals than current pipelines. Using experimental data obtained by performing Multiplexed Error-Robust Fluorescence in situ Hybridization (MERFISH) on tissue from mouse motor cortex, we demonstrate that JSIT enables improved throughput and recovery performance over standard decoding methods. Contactyonina.eldar@weizmann.ac.il, sfarhi@broadinstitute.org, bcleary@bu.edu Availability and ImplementationSoftware implementation of JSIT, together with example files, are available at https://github.com/jpbryan13/JSIT. Supplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics↗

A Novel Approach to Topological Network Analysis for the Identification of Metrics and Signatures in Non-Small Cell Lung Cancer

Non-small cell lung cancer (NSCLC), the primary histological form of lung cancer, accounts for about 25% - the highest - of all cancer deaths. As NSCLC is often undetected until symptoms appear in the late stages, it is imperative to discover more effective tumor-associated biomarkers for early diagnosis. Topological data analysis is one of the most powerful methodologies applicable to biological networks. However, current studies fail to consider the biological significance of their quantitative methods and utilize popular scoring metrics without verification, leading to low performance. To extract meaningful insights from genomic data, it is essential to understand the relationship between geometric correlations and biological function mechanisms. Through bioinformatics and network analyses, we propose a novel composite selection index, the C-Index, that best captures significant pathways and interactions in gene networks to identify biomarkers with the highest efficiency and accuracy. Furthermore, we establish a 4-gene biomarker signature that serves as a promising therapeutic target for NSCLC and personalized medicine. We designed a Cascading machine learning model to validate both the C-Index and the biomarkers discovered. The methodology proposed for finding top metrics can be applied to effectively select biomarkers and early diagnose many diseases, revolutionizing the approach to topological network research for all cancers.

bioinformatics↗

SCALA: A web application for multimodal analysis of single cell next generation sequencing data

Analysis and interpretation of high-throughput transcriptional and chromatin accessibility data at single cell resolution are still open challenges in the biomedical field. In this article, we present SCALA, a bioinformatics tool for analysis and visualization of single cell RNA sequencing (scRNA-seq) and Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) datasets. SCALA combines standard types of analysis by integrating multiple software packages varying from quality control to identification of distinct cell population and cell states. Additional analysis options enable functional enrichment, cellular trajectory inference, ligand-receptor analysis and regulatory network reconstruction. SCALA is fully parameterizable at every step of the analysis, presenting data in tabular format and produces publication-ready 2D and 3D visualizations including heatmaps, barcharts, scatter, violin and volcano plots. We demonstrate the functionality of SCALA through two use-cases related to TNF-driven arthritic mice, handling data from both scRNA-seq and scATAC-seq experiments. SCALA is mainly developed in R, Shiny and JavaScript and is available as a web application at http://scala.pavlopouloslab.info or https://scala.fleming.gr.

bioinformatics↗

Developing best practices for genotyping-by-sequencing analysis using linkage maps as benchmarks

BackgroundGenotyping-by-Sequencing (GBS) provides affordable methods for genotyping hundreds of individuals using millions of markers. However, this challenges bioinformatic procedures that must overcome possible artifacts such as the bias generated by PCR duplicates and sequencing errors. Genotyping errors lead to data that deviate from what is expected from regular meiosis. This, in turn, leads to difficulties in grouping and ordering markers resulting in inflated and incorrect linkage maps. Therefore, genotyping errors can be easily detected by linkage map quality evaluations. ResultsWe developed and used the Reads2Map workflow to build linkage maps with simulated and empirical GBS data of diploid outcrossing populations. The workflows run GATK, Stacks, TASSEL, and Freebayes for SNP calling and updog, polyRAD, and SuperMASSA for genotype calling, and OneMap and GUSMap to build linkage maps. Using simulated data, we observed which genotype call software fails in identifying common errors in GBS sequencing data and proposed specific filters to better handle them. We tested whether it is possible to overcome errors in a linkage map using genotype probabilities from each software or global error rates to estimate genetic distances with an updated version of OneMap. We also evaluated the impact of segregation distortion, contaminant samples, and haplotype-based multiallelic markers in the final linkage maps. Through our evaluations, we observed that some of the approaches produce different results depending on the dataset (dataset-dependent) and others produce consistent advantageous results among them (dataset-independent). ConclusionsWe set as default in the Reads2Map workflows the approaches that showed to be dataset-independent for GBS datasets according to our results. This reduces the number required of tests to identify optimal pipelines and parameters for other empirical datasets. Using Reads2Map, users can select the pipeline and parameters that best fit their data context. The Reads2MapApp shiny app provides a graphical representation of the results to facilitate their interpretation.

bioinformatics↗

pLM-BLAST - distant homology detection based on direct comparison of sequence representations from protein language models

MotivationThe detection of homology through sequence comparison is a typical first step in the study of protein function and evolution. In this work, we explore the applicability of protein language models to this task. ResultsWe introduce pLM-BLAST, a tool inspired by BLAST, that detects distant homology by comparing single-sequence representations (embeddings) derived from a protein language model, ProtT5. Our benchmarks reveal that pLM-BLAST maintains a level of accuracy on par with HHsearch for both highly similar sequences (with over 50% identity) and markedly divergent sequences (with less than 30% identity), while being significantly faster. Additionally, pLM-BLAST stands out among other embedding-based tools due to its ability to compute local alignments. We show that these local alignments, produced by pLM-BLAST, often connect highly divergent proteins, thereby highlighting its potential to uncover previously undiscovered homologous relationships and improve protein annotation. Availability and ImplementationpLM-BLAST is accessible via the MPI Bioinformatics Toolkit as a web server for searching precomputed databases (https://toolkit.tuebingen.mpg.de/tools/plmblast). It is also available as a standalone tool for building custom databases and performing batch searches (https://github.com/labstructbioinf/pLM-BLAST).

bioinformatics↗

RNAlysis: analyze your RNA sequencing data without writing a single line of code

BackgroundAmongst the major challenges in next-generation sequencing experiments are exploratory data analysis, interpreting trends, identifying potential targets/candidates, and visualizing the results clearly and intuitively. These hurdles are further heightened for researchers who are not experienced in writing computer code, since the majority of available analysis tools require programming skills. Even for proficient computational biologists, an efficient and replicable system is warranted to generate standardized results. ResultsWe have developed RNAlysis, a modular Python-based analysis software for RNA sequencing data. RNAlysis allows users to build customized analysis pipelines suiting their specific research questions, going all the way from raw FASTQ files, through exploratory data analysis and data visualization, clustering analysis, and gene-set enrichment analysis. RNAlysis provides a friendly graphical user interface, allowing researchers to analyze data without writing code. We demonstrate the use of RNAlysis by analyzing RNA data from different studies using C. elegans nematodes. We note that the software is equally applicable to data obtained from any organism. ConclusionsRNAlysis is suitable for investigating a variety of biological questions, and allows researchers to more accurately and reproducibly run comprehensive bioinformatic analyses. It functions as a gateway into RNA sequencing analysis for less computer-savvy researchers, but can also help experienced bioinformaticians make their analyses more robust and efficient, as it offers diverse tools, scalability, automation, and standardization between analyses.

bioinformatics↗

Deep Learning for Predicting 16S rRNA Copy Number

BackgroundCulture-independent 16S rRNA gene metabarcoding is a commonly used method in microbiome profiling. However, this approach can only reflect the proportion of sequencing reads, rather than the actual cell fraction. To achieve more quantitative cell fraction estimates, we need to resolve the 16S gene copy numbers (GCN) for different community members. Currently, there are several bioinformatic tools available to estimate 16S GCN, either based on taxonomy assignment or phylogeny. MethodHere we develop a novel algorithm, Stacked Ensemble Model (SEM), that estimates 16S GCN directly from the 16S rRNA gene sequence strings, without resolving taxonomy or phylogeny. For accessibility, we developed a public, end-to-end, web-based tool based on the SEM model, named Artificial Neural Network Approximator for 16S rRNA Gene Copy Number (ANNA16). ResultsBased on 27,579 16S rRNA gene sequence data (rrnDB database), we show that ANNA16 outperforms the most commonly used 16S GCN prediction algorithms. The prediction error range in the 5-fold cross validation of SEM is completely lower than all other algorithms for the 16S full-length sequence and partially lower at 16S subregions. The final test and a mock community test indicate ANNA16 is more accurate than all currently available tools (i.e., rrnDB, CopyRighter, PICRUSt2, & PAPRICA). SHAP value analysis indicates ANNA16 mainly learns information from rare insertions. ConclusionANNA16 represents a deep learning based 16S GCN prediction tool. Compared to the traditional GCN prediction tools, ANNA16 has a simple structure, faster inference speed without precomputing, and higher accuracy. With increased 16S GCN data in the database, future studies could improve the prediction errors for rare, high-GCN taxa due to current under sampling.

bioinformatics↗

Inferring delays in partially observed gene regulatory networks

MotivationCell function is regulated by gene regulatory networks (GRNs) defined by protein-mediated interaction between constituent genes. Despite advances in experimental techniques, we can still measure only a fraction of the processes that govern GRN dynamics. To infer the properties of GRNs using partial observation, unobserved sequential processes can be replaced with distributed time delays, yielding non-Markovian models. Inference methods based on the resulting model suffer from the curse of dimensionality. ResultsWe develop a simulation-based Bayesian MCMC method for the efficient and accurate inference of GRN parameters when only some of their products are observed. We illustrate our approach using a two-step activation model: An activation signal leads to the accumulation of an unobserved regulatory protein, which triggers the expression of observed fluorescent proteins. With prior information about observed fluorescent protein synthesis, our method successfully infers the dynamics of the unobserved regulatory protein. We can estimate the delay and kinetic parameters characterizing target regulation including transcription, translation, and target searching of an unobserved protein from experimental measurements of the products of its target gene. Our method is scalable and can be used to analyze non-Markovian models with hidden components. AvailabilityAccompanying code in R is available at https://github.com/Mathbiomed/SimMCMC. Contactjaekkim@kaist.ac.kr or kresimir.josic@gmail.com or cbskust@korea.ac.kr Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Fast protein structure searching using structure graph embeddings

Comparing and searching protein structures independent of primary sequence has proved useful for remote homology detection, function annotation and protein classification. Fast and accurate methods to search with structures will be essential to make use of the vast databases that have recently become available, in the same way that fast protein sequence searching underpins much of bioinformatics. We train a simple graph neural network using supervised contrastive learning to learn a low-dimensional embedding of protein structure. The method, called Progres, is available as software at https://github.com/greener-group/progres and as a web server at https://progres.mrc-lmb.cam.ac.uk. It has accuracy comparable to the best current methods and can search the AlphaFold database TED domains in a tenth of a second per query on CPU.

bioinformatics↗

Gene Set Scoring on the Nearest Neighbor Graph (gssnng) for Single Cell RNA-seq (scRNA-seq)

Gene set scoring (or enrichment) is a common dimension reduction task in bioinformatics that can be focused on differences between groups or at the single sample level. Gene sets can represent biological functions, molecular pathways, cell identities, and more. Gene set scores are context dependent values that are useful for interpreting biological changes following experiments or perturbations. Single sample scoring produces a set of scores, one for each member of a group, which can be analyzed with statistical models that can include additional clinically important factors such as gender or age. However, the sparsity and technical noise of single cell expression measures create difficulties for these methods, which were originally designed for bulk expression profiling (microarrays, RNAseq). This can be greatly remedied by first applying a smoothing transformation that shares gene measure information within transcriptomic neighborhoods. In this work, we use the nearest neighbor graph of cells for matrix smoothing to produce high quality gene set scores on a per-cell, per-group, level which is useful for visualization and statistical analysis. Availability and implementationThe gssnng software is available using the python package index (PyPI) and works with Scanpy AnnData objects. It can be installed using pip install gssnng. More information and demo notebooks: See https://github.com/IlyaLab/gssnng

bioinformatics↗

MLcps: Machine Learning Cumulative Performance Score for classification problems

MotivationA performance metric is a tool to measure the correctness of a trained Machine Learning (ML) model. Numerous performance metrics have been developed for classification problems making it overwhelming to select the appropriate one since each of them represents a particular aspect of the model. Furthermore, selection of a performance metric becomes harder for problems with imbalanced and/or small datasets. Therefore, in clinical studies where datasets are frequently imbalanced and, in situations when the prevalence of a disease is low or the collection of patient samples is difficult, deciding on a suitable metric for performance evaluation of an ML model becomes quite challenging. The most common approach to address this problem is measuring multiple metrics and compare them to identify the best-performing ML model. However, comparison of multiple metrics is laborious and prone to user preference bias. Furthermore, evaluation metrics are also required by ML model optimization techniques such as hyperparameter tuning, where we train many models, each with different parameters, and compare their performances to identify the best-performing parameters. In such situations, it becomes almost impossible to assess different models by comparing multiple metrics. ResultsHere, we propose a new metric called Machine Learning Cumulative Performance Score (MLcps) as a Python package for classification problems. MLcps combines multiple pre-computed performance metrics into one metric that conserves the essence of all pre-computed metrics for a particular model. We tested MLcps on 4 different publicly available biological datasets and the results reveal that it provides a comprehensive picture of overall model robustness. AvailabilityMLcps is available at https://pypi.org/project/MLcps/ and cases of use are available at https://mybinder.org/v2/gh/FunctionalUrology/MLcps.git/main. Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Fast and accurate protein function prediction from sequence through pretrained language model and homology-based label diffusion

Protein function prediction is an essential task in bioinformatics which benefits disease mechanism elucidation and drug target discovery. Due to the explosive growth of proteins in sequence databases and the diversity of their functions, it remains challenging to fast and accurately predict protein functions from sequences alone. Although many methods have integrated protein structures, biological networks or literature information to improve performance, these extra features are often unavailable for most proteins. Here, we propose SPROF-GO, a Sequence-based alignment-free PROtein Function predictor which leverages a pretrained language model to efficiently extract informative sequence embeddings and employs self-attention pooling to focus on important residues. The prediction is further advanced by exploiting the homology information and accounting for the overlapping communities of proteins with related functions through the label diffusion algorithm. SPROF-GO was shown to surpass state-of-the-art sequence-based and even network-based approaches by more than 14.5%, 27.3% and 10.1% in AUPR on the three sub-ontology test sets, respectively. Our method was also demonstrated to generalize well on non-homologous proteins and unseen species. Finally, visualization based on the attention mechanism indicated that SPROF-GO is able to capture sequence domains useful for function prediction. Key pointsO_LISPROF-GO is a sequence-based protein function predictor which leverages a pretrained language model to efficiently extract informative sequence embeddings, thus bypassing expensive database searches. C_LIO_LISPROF-GO employs self-attention pooling to capture sequence domains useful for function prediction and provide interpretability. C_LIO_LISPROF-GO applies hierarchical learning strategy to produce consistent predictions and label diffusion to exploit the homology information. C_LIO_LISPROF-GO is accurate and robust, with better performance than state-of-the-art sequence-based and even network-based approaches, and great generalization ability on non-homologous proteins and unseen species C_LI

bioinformatics↗

Discovering Haematoma-Stimulated Circuits for Secondary Brain Injury after Intraventricular Haemorrhage by Spatial Transcriptome Analysis

Central nervous system (CNS) diseases, such as neurodegenerative disorders and brain diseases caused by acute injuries, are important yet challenging to study due to disease leisions locations and other complexities. Utilizing the powerful spatial transcriptome analysis together with novel algorithms we developed for the study, we report here for the first time a 3D trajectory map of gene expression changes in the brain following acute neural injury using a mouse model of intraventricular haemorrhage (IVH). IVH is a common and representative complication after various acute brain injuries with severe mortality and mobility implications. Our data identified three main 3D global pseudospace-time trajectory bundles which represent the main neural circuits from the lateral ventricle to the hippocampus and primary cortex induced by experimental intraventricular haematoma stimulation. Further analysis indicated a rapid response in the primary cortex, as well as a direct and integrated effect on the hippocampus after IVH. These results are informative in understanding the pathophysiological changes, including the spatial and temporal patterns of gene expression changes, in IVH patients after acute brain injuries and strategizing more effective clinical management regimens. The bioinformatics strategies will also be useful for the study of other CNS diseases.

bioinformatics↗

Unsupervised reference-free inference reveals unrecognized regulated transcriptomic complexity in human single cells

Myriad mechanisms diversify the sequence content of eukaryotic transcripts at both the DNA and RNA levels, leading to profound functional consequences. Examples of this diversity include RNA splicing and V(D)J recombination. Currently, these mechanisms are detected using fragmented bioinformatic tools that require predefining a form of transcript diversification and rely on alignment to an incomplete reference genome, filtering out unaligned sequences, potentially crucial for novel discoveries. Here, we present SPLASH+, significantly advancing biological discovery possible with SPLASH, our recently introduced efficient, reference-free statistical approach. Integrating a micro-assembly and biological interpretation framework, SPLASH+ enables new discoveries including broad and novel examples of transcript diversification in single cells de novo, without the need for cell type metadata, which is impossible with current algorithms. Applied to 10,326 primary human single cells across 19 tissues profiled with SmartSeq2, SPLASH+ discovers a set of splicing and histone regulators with highly conserved intronic regions that are themselves subject to complex splicing regulation. Additionally, it reveals unreported transcript diversity in the heat shock protein HSP90AA1, as well as diversification in centromeric RNA expression, V(D)J recombination, RNA editing, and repeat expansion, all missed by existing methods. SPLASH+ is highly efficient, enabling the discovery of an unprecedented breadth of RNA regulation and diversification in single cells through a new automated paradigm of unbiased transcriptomic analysis.

bioinformatics↗

Bayesian clustering with uncertain data

Clustering is widely used in bioinformatics and many other fields, with applications from exploratory analysis to prediction. Many types of data have associated uncertainty or measurement error, but this is rarely used to inform the clustering. We present Dirichlet Process Mixtures with Uncertainty (DPMUnc), an extension of a Bayesian nonparametric clustering algorithm which makes use of the uncertainty associated with data points. We show that DPMUnc out-performs existing methods on simulated data. We cluster immune-mediated diseases (IMD) using GWAS summary statistics, which have uncertainty linked with the sample size of the study. DPMUnc separates autoimmune from autoinflammatory diseases and isolates other subgroups such as adult-onset arthritis. We additionally consider how DPMUnc can be used to cluster gene expression datasets that have been summarised using gene signatures. We first introduce a novel procedure for generating a summary of a gene signature on a dataset different to the one where it was discovered. Since the genes in the gene signature are unlikely to be as strongly correlated as in the original dataset, it is important to quantify the variance of the gene signature for each individual. We summarise three public gene expression datasets containing patients with a range of IMD, using three relevant gene signatures. We find association between disease and the clusters returned by DPMUnc, with clustering structure replicated across the datasets. The significance of this work is two-fold. Firstly, we demonstrate that when data has associated uncertainty, this uncertainty should be used to inform clustering and we present a method which does this, DPMUnc. Secondly, we present a procedure for using gene signatures in datasets other than where they were originally defined. We show the value of this procedure by summarising gene expression data from patients with immune-mediated diseases using relevant gene signatures, and clustering these patients using DPMUnc. Author SummaryIdentifying groups of items that are similar to each other, a process called clustering, has a range of applications. For example, if patients split into two distinct groups this suggests that a disease may have subtypes which should be treated differently. Real data often has measurement error associated with it, but this error is frequently discarded by clustering methods. We propose a clustering method which makes use of the measurement error and use it to cluster diseases linked to the immune system. Gene expression datasets measure the activity level of all ~20,000 genes in the human genome. We propose a procedure for summarising gene expression data using gene signatures, lists of genes produced by highly focused studies. For example, a study might list the genes which increase activity after exposure to a particular virus. The genes in the gene signature may not be as tightly correlated in a new dataset, and so our procedure measures the strength of the gene signature in the new dataset, effectively defining measurement error for the summary. We summarise gene expression datasets related to the immune system using relevant gene signatures and find that our method groups patients with the same disease.

bioinformatics↗