bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Metheor: Ultrafast DNA methylation heterogeneity calculation from bisulfite read alignments

MotivationPhased DNA methylation states within bisulfite sequencing reads are valuable source of information that can be used to estimate epigenetic diversity across cells as well as epigenomic instability in individual cells. Various measures capturing the heterogeneity of DNA methylation states have been proposed for a decade. However, in routine analyses on DNA methylation, this heterogeneity is often overlooked by computing average methylation levels at CpG sites. In this study, to facilitate the application of the DNA methylation heterogeneity measures in downstream epigenomic analyses, we present a Rust-based, extremely fast and lightweight bioinformatics toolkit called Metheor. ResultsWe benchmark the performance of Metheor against existing code implementation for DNA methylation heterogeneity measures in three different scenarios of simulated bisulfite sequencing datasets. Metheor was shown to dramatically reduce the execution time up to 300-fold and memory footprint up to 60-fold, while producing the same results with the original implementation. AvailabilitySource code for Metheor is at https://github.com/dohlee/metheor and is freely available for non-commercial users. Contactsunkim.bioinfo@snu.ac.kr

bioinformatics↗

Deep self-supervised learning for biosynthetic gene cluster detection and product classification

Natural products are chemical compounds that form the basis of many therapeutics used in the pharmaceutical industry. In microbes, natural products are synthesized by groups of colocalized genes called biosynthetic gene clusters (BGCs). With advances in high-throughput sequencing, there has been an increase of complete microbial isolate genomes and metagenomes, from which a vast number of BGCs are undiscovered. Here, we introduce a self-supervised learning approach designed to identify and characterize BGCs from such data. To do this, we represent BGCs as chains of functional protein domains and train a masked language model on these domains. We assess the ability of our approach to detect BGCs and characterize BGC properties in bacterial genomes. We also demonstrate that our model can learn meaningful representations of BGCs and their constituent domains, detect BGCs in microbial genomes, and predict BGC product classes. These results highlight self-supervised neural networks as a promising framework for improving BGC prediction and classification. Author summaryBiosynthetic gene clusters (BGCs) encode for natural products of diverse chemical structures and function, but they are often difficult to discover and characterize. Many bioinformatic and deep learning approaches have leveraged the abundance of genomic data to recognize BGCs in bacterial genomes. However, the characterization of BGC properties remains the main bottleneck in identifying novel BGCs and their natural products. In this paper, we present a self-supervised masked language model that learns meaningful representations of BGCs with improved downstream detection and classification.

bioinformatics↗

Paralog Explorer: a resource for mining information about paralogs in common research organisms

Paralogs are genes which arose via gene duplication, and when such paralogs retain overlapping or redundant function, this poses a challenge to functional genetics research. Recent technological advancements have made it possible to systematically probe gene function for redundant genes using dual or multiplex gene perturbation, and there is a need for a simple bioinformatic tool to identify putative paralogs of a gene(s) of interest. We have developed Paralog Explorer (https://www.flyrnai.org/tools/paralogs/), an online resource that allows researchers to quickly and accurately identify candidate paralogous genes in the genomes of the model organisms D. melanogaster, C. elegans, D. rerio, M. musculus, and H. sapiens. Paralog Explorer deploys an effective between-species ortholog prediction software, DIOPT, to analyze within-species paralogs. Paralog Explorer allows users to identify candidate paralogs, and to navigate relevant databases regarding gene co-expression, protein-protein and genetic interaction, as well as gene ontology and phenotype annotations. Altogether, this tool extends the value of current ortholog prediction resources by providing sophisticated features useful for identification and study of paralogous genes.

bioinformatics↗

PEMT: A patent enrichment tool for drug discovery

MotivationDrug discovery practitioners in industry and academia use semantic tools to extract information from online scientific literature to generate new insights into targets, therapeutics and diseases. However, due to complexities in access and analysis, patent-based literature is often overlooked as a source of information. As drug discovery is a highly competitive field, naturally, tools that tap into patent literature can provide any actor in the field an advantage in terms of better informed decision making. Hence, we aim to facilitate access to patent literature through the creation of an automatic tool for extracting information from patents described in existing public resources. ResultsHere, we present PEMT, a novel patent enrichment tool, that takes advantage of public databases like ChEMBL and SureChEMBL to extract relevant patent information linked to chemical structures and/or gene names described through FAIR principles and metadata annotations. PEMT aims at supporting drug discovery and research by establishing a patent landscape around genes of interest. The pharmaceutical focus of the tool is mainly due to the subselection of International Patent Classification (IPC) codes, but in principle, it can be used for other patent fields, provided that a link between a concept and chemical structure is investigated. Finally, we demonstrate a use-case in rare diseases by generating a gene-patent list based on the epidemiological prevalence of these diseases and exploring their underlying patent landscapes. Availability and implementationPEMT is an open-source Python tool and its source code and PyPi package are available at https://github.com/Fraunhofer-ITMP/PEMT and https://pvpi.org/project/PEMT/ respectively. Contactyojana.gadiya@itmp.fraunhofer.de Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Reference BioImaging to assess the phenotypic trait diversity of bryophytes within the family Scapaniaceae

Macro- and microscopic images of organisms are pivotal in biodiversity research. Despite that bioimages have manifold applications such as for assessing the diversity of form and function, FAIR bioimaging data in the context of biodiversity are still very scarce, especially for difficult taxonomic groups such as bryophytes. Here, we present a high-quality reference dataset containing macroscopic and bright-field microscopic images documenting various phenotypic attributes of the species belonging to the family of Scapaniaceae occurring in Europe. To encourage data reuse in biodiversity and adjacent research areas, we annotated the imaging data with machine-actionable meta-data using community-accepted semantics. Furthermore, raw imaging data are retained and any contextual image processing like multi-focus image fusion and stitching were documented to foster good scientific practices through source tracking and provenance. The information contained in the raw images are also of particular interest for machine learning and image segmentation used in bioinformatics and computational ecology. We expect that this richly annotated reference dataset will encourage future studies to follow our principles.

bioinformatics↗

A Graph Coarsening Algorithm for Compressing Representations of Single-Cell Data with Clinical or Experimental Attributes

Graph-based algorithms have become essential in the analysis of single-cell data for numerous tasks, such as automated cell-phenotyping and identifying cellular correlates of experimental perturbations or disease states. In large multi-patient, multi-sample single-cell datasets, the analysis of cell-cell similarity graphs representations of these data becomes computationally prohibitive. Here, we introduce cytocoarsening, a novel graph-coarsening algorithm that significantly reduces the size of single-cell graph representations, which can then used as input to downstream bioinformatics algorithms for improved computational efficiency. Uniquely, cytocoarsening considers both phenotypical similarity of cells and similarity of cells associated clinical or experimental attributes in order to more readily identify condition-specific cell populations. The resulting coarse graph representations were evaluated based on both their structural correctness and the capacity of downstream algorithms to uncover the same biological conclusions as if the full graph had been used. Cytocoarsening is provided as open source code at https://github.com/ChenCookie/cytocoarsening.

bioinformatics↗

HSDatabase - a database of highly similar duplicate genes from plants, animals, and algae.

Gene duplication is an important evolutionary mechanism capable of providing new genetic material, which can help organisms adapt to various environmental conditions. Recent studies, for example, have indicated that highly similar duplicated genes (HSDs) are involved in adaptation to extreme conditions via gene dosage. However, HSDs in most genomes remain uncharacterized. Here, we collected and curated HSDs in nuclear genomes from a diversity of species and indexed them in an online, open-access sequence repository called HSDatabase. Currently, this database contains 117,864 curated HSDs from 40 eukaryotic genomes, and it includes information on the total HSD number, gene copy number/length, and alignments of gene copies. HSDatabase also allows users to download sequences of gene copies, access genome browsers, and link out to other databases, such as Pfam and KEGG. Whats more, a built-in Basic Local Alignment Search Tool (BLAST) option is available to conveniently explore potential homologous sequences of interest within and across species. HSDatabase is presented with a user-friendly interface and provides easy access to the source data. It can be used on its own for comparative analyses of gene duplicates or in conjunction with HSDFinder, a newly developed bioinformatics tool for identifying, annotating, categorizing, and visualizing HSDs. Database URLhttp://hsdfinder.com/database/

bioinformatics↗

Hierarchical Interleaved Bloom Filter: Enabling ultrafast, approximate sequence queries

Searching sequences in large, distributed databases is the most widely used bioinformatics analysis done. This basic task is in dire need for solutions that deal with the exponential growth of sequence repositories and perform approximate queries very fast. In this paper, we present a novel data structure: the Hierarchical Interleaved Bloom Filter (HIBF). It is extremely fast and space efficient, yet so general that it has the potential to serve as the underlying engine for many applications. We show that the HIBF is superior in build time, index size and search time while achieving a comparable or better accuracy compared to other state-of-the art tools (Mantis and Bifrost). The HIBF builds an index up to 211 times faster, using up to 14 times less space and can answer approximate membership queries faster by a factor of up to 129. This can be considered a quantum leap that opens the door to indexing complete sequence archives like the European Nucleotide Archive or even larger metagenomics data sets.

bioinformatics↗

Molecular formula discovery via bottom-up MS/MS interrogation

A substantial fraction of metabolic features remains undetermined in mass spectrometry (MS)-based metabolomics. Here we present bottom-up tandem MS (MS/MS) interrogation to illuminate the unidentified features via accurate molecular formula annotation. Our approach prioritizes MS/MS-explainable formula candidates, implements machine-learned ranking, and offers false discovery rate estimation. Compared to the existing MS1-initiated formula annotation, our approach shrinks the formula candidate space by 42.8% on average. The superior annotation accuracy of our bottom-up interrogation was demonstrated on reference MS/MS libraries and real metabolomics datasets. Applied on 155,321 annotated recurrent unidentified spectra (ARUS), our approach confidently annotated >5,000 novel molecular formulae unarchived in chemical databases. Beyond the level of individual metabolic features, we combined bottom-up MS/MS interrogation with global peak annotation. This approach reveals peak interrelationships, allowing the systematic annotation of 37 fatty acid amide molecules in human fecal data, among other applications. All bioinformatics pipelines are available in a standalone software, BUDDY (https://github.com/HuanLab/BUDDY/).

bioinformatics↗

LambdaPP: Fast and accessible protein-specific phenotype predictions

The availability of accurate and fast Artificial Intelligence (AI) solutions predicting aspects of proteins are revolutionizing experimental and computational molecular biology. The webserver LambdaPP aspires to supersede PredictProtein, the first internet server making AI protein predictions available in 1992. Given a protein sequence as input, LambdaPP provides easily accessible visualizations of protein 3D structure, along with predictions at the protein level (GeneOntology, subcellular location), and the residue level (binding to metal ions, small molecules, and nucleotides; conservation; intrinsic disorder; secondary structure; alpha-helical and beta-barrel transmembrane segments; signal-peptides; variant effect) in seconds. The structure prediction provided by LambdaPP - leveraging ColabFold and computed in minutes - is based on MMseqs2 multiple sequence alignments. All other feature prediction methods are based on the pLM ProtT5. Queried by a protein sequence, LambdaPP computes protein and residue predictions almost instantly for various phenotypes, including 3D structure and aspects of protein function. Accessibility StatementLambdaPP is freely available for everyone to use under embed.predictprotein.org, the interactive results for the case study can be found under https://embed.predictprotein.org/o/Q9NZC2. The frontend of LambdaPP can be found on GitHub (github.com/sacdallago/embed.predictprotein.org), and can be freely used and distributed under the academic free use license (AFL-2). For high-throughput applications, all methods can be executed locally via the bio-embeddings (bioembeddings.com) python package, or docker image at ghcr.io/bioembeddings/bio_embeddings, which also includes the backend of LambdaPP. Impact StatementWe introduce LambdaPP, a webserver integrating fast and accurate sequence-only protein feature predictions based on embeddings from protein Language Models (pLMs) available in seconds along with high-quality protein structure predictions. The intuitive interface invites experts and novices to benefit from the latest machine learning tools. LambdaPPs unique combination of predicted features may help in formulating hypotheses for experiments and as input to bioinformatics pipelines.

bioinformatics↗

Concatenated 16S rRNA Sequence Analysis Improve Bacterial Taxonomy

Microscopic, biochemical, molecular, and computer-based approaches are extensively used to identify and classify bacterial populations. Further, advances in DNA sequencing and bioinformatics workflows facilitated sophisticated genome-based methods for microbial taxonomy. Although sequencing of 16S rRNA gene is widely employed to identify and classify the bacterial community as a cost-effective and single-gene approach. However, the accuracy of the 16S rRNA sequence-based species identification is limited by multiple copies of the gene and their higher sequence identity between closely related species. Availability of a large volume of bacterial whole-genome data provided an opportunity to develop comprehensive species-specific 16S rRNA reference libraries. With defined rules, we have concatenated four 16S rRNA gene copy variants to develop a species-specific reference library. Using this approach, species-specific 16S rRNA gene libraries were developed for four closely related Streptococcus species (S. gordonii, S. mitis, S. oralis, and S. pneumoniae). Sequence similarity and phylogenetic analysis of concatenated 16S rRNA copies yielded better resolution than single gene copy approaches. The approach is very effective to classify genetically related species, and it may reduce misclassification of bacterial species and genome assemblies.

bioinformatics↗

metGWAS 1.0: An R workflow for network-driven over-representation analysis between independent metabolomic and meta-genome wide association studies

BackgroundMany diseases may result from disrupted metabolic regulation. Metabolite-GWAS studies assess the association of polymorphic variants with metabolite levels in body fluids. While these studies are successful, they have a high cost and technical expertise burden due to combining the analytical biochemistry of metabolomics with the computational genetics of GWAS. Currently, there are 100s of standalone metabolomics and GWAS studies related to similar diseases or phenotypes. A method that could statically evaluate these independent studies to find novel metabolites-genes association is of high interest. Although such an analysis is limited to genes with known metabolite interactions due to the unpaired nature of the data sets, any discovered associations may represent biomarkers and druggable targets for treatment and prevention. MethodsWe developed a bioinformatics tool, metGWAS 1.0, that generates and statistically compares metabolic and genomic gene sets using a hypergeometric test. Metabolic gene sets are generated by mapping disease-associated metabolites to interacting proteins (genes) via online databases. Genomic gene sets are identified from a network representation of the GWAS Catalog comprising 100s of studies. ResultsThe metGWAS 1.0 tool was evaluated using standalone metabolomics datasets extracted from two metabolomics-GWAS case studies. In case-study 1, a cardiovascular disease association study, we identified nine genes (APOA5, PLA2G5, PLA2G2D, PLA2G2E, PLA2G2F, LRAT, PLA2G2A, PLB1, and PLA2G7) that interact with metabolites in the KEGG glycerophospholipid metabolism pathway and contain polymorphic variants associated with cardiovascular disease (P < 0.005). The gene APOA5 was matched from the original metabolomics-GWAS study. In case study 2, a urine metabolome study of kidney metabolism in healthy subjects, we found marginal significance (P = 0.10 and P = 0.13) for glycine, serine, and threonine metabolism and alanine, aspartate, and glutamate metabolism pathways to GWAS data relating to kidney disease. ConclusionThe metGWAS 1.0 platform provides insight into developing methods that bridge standalone metabolomics and disease and phenotype GWAS data. We show the potential to reproduce findings of paired metabolomics-GWAS data and provide novel associations of gene variation and metabolite expression.

bioinformatics↗

MONI-k: An index for efficient pangenome-to-pangenome comparison

Maximal exact matches (MEMs) are widely used in bioinformatics, originally for genome-to-genome comparison but especially for DNA alignment ever since Li (2013) presented BWA-MEM. Building on work by Bannai, Gagie and I (2018) and again targeting alignment, Rossi et al. (2022) recently built an index called MONI that is based on the run-length compressed Burrows-Wheeler Transform and can find MEMs efficiently with respect to pangenomes. In this paper we define k-MEMs to be maximal substrings of a pattern that each occur exactly at least k times in a text (so a MEM is a 1-MEM) and briefly explain why computing k-MEMs could be useful for pangenome-to-pangenome comparison. We then show that, when k is given at construction time, MONI can easily be extended to find k-MEMs efficiently as well.

bioinformatics↗

Cytocipher detects significantly different populations of cells in single cell RNA-seq data

Identification of cell types using single cell RNA-seq (scRNA-seq) is revolutionising the study of multicellular organisms. However, typical scRNA-seq analysis often involves post hoc manual curation to ensure clusters are transcriptionally distinct, which is time-consuming, error-prone, and irreproducible. To overcome these obstacles, we developed Cytocipher, a bioinformatics method and scverse compatible software package that statistically determines significant clusters. Application of Cytocipher to normal tissue, development, disease, and large-scale atlas data reveals the broad applicability and power of Cytocipher to generate biological insights in numerous contexts. This included the identification of cell types not previously described in the datasets analyzed, such as CD8+ T cell subtypes in human peripheral blood mononuclear cells; cell lineage intermediate states during mouse pancreas development; and subpopulations of luminal epithelial cells over-represented in prostate cancer. Cytocipher also scales to large datasets with high test performance, as shown by application to the Tabula Sapiens Atlas representing >480,000 cells. Cytocipher is a novel and generalisable method that statistically determines transcriptionally distinct and programmatically reproducible clusters from single cell data. Cytocipher is available at https://github.com/BradBalderson/Cytocipher.

bioinformatics↗

Accurately modeling biased random walks on weighted networks using node2vec+

MotivationAccurately representing biological networks in a low-dimensional space, also known as network embedding, is a critical step in network-based machine learning and is carried out widely using node2vec, an unsupervised method based on biased random walks. However, while many networks, including functional gene interaction networks, are dense, weighted graphs, node2vec is fundamentally limited in its ability to use edge weights during the biased random walk generation process, thus under-using all the information in the network. ResultsHere, we present node2vec+, a natural extension of node2vec that accounts for edge weights when calculating walk biases and reduces to node2vec in the cases of unweighted graphs or unbiased walks. Using two synthetic datasets, we empirically show that node2vec+ is more robust to additive noise than node2vec in weighted graphs. Then, using genome-scale functional gene networks to solve a wide range of gene function and disease prediction tasks, we demonstrate the superior performance of node2vec+ over node2vec in the case of weighted graphs. Notably, due to the limited amount of training data in the gene classification tasks, graph neural networks such as GCN and GraphSAGE are outperformed by both node2vec and node2vec+ Contactarjun.krishnan@cuanschutz.edu Code Availabilityhttps://github.com/krishnanlab/node2vecplus_benchmarks Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Identification of divergent botulinum neurotoxin homologs in Paeniclostridium ghonii

Botulinum neurotoxins (BoNTs) are the most potent family of toxins known to science. Bioinformatic studies in recent years have revealed that they are members of a broader toxin family, with an increasing number of divergent homologs identified in genomes of organisms outside of the Clostridium genus. Here, we report the identification of two putative divergent BoNT-like homologs in the genomes of two strains of Paeniclostridium ghonii. We designated them PG-toxin 1 (PGT1) and PG-toxin 2 (PGT2), which share ~54% protein sequence identity. Unlike any other known BoNT homologs, PGT1 and PGT2 are composed of two separate subunits encoded on two neighboring genes: one encoding the protease domain (light chain, LC) with a conserved HExxH motif, and the second encoding the heavy-chain (HC) containing the putative translocation domain and receptor-binding domain. Phylogenetic analysis of both the LC and HC reveal that it is a divergent member of the lineage of BoNT that also includes BoNT/X, BoNT/En and the insecticidal PMP1. The gene clusters harboring PGT1 and PGT2 also include a putative insecticidal delta-endotoxin, Cry8Ea1, as well as putative endolysin and bacteriocin genes that may facilitate lytic toxin secretion, suggesting a possibility that this gene cluster might serve an insecticidal purpose.

bioinformatics↗

Comparison of de novo and reference genome-based transcriptome assembly pipelines for differential expression analysis of RNA sequencing data

ObjectiveAs sequencing technologies become more accessible and bioinformatic tools improve, genomic resources are increasingly available for non-model species. Using a draft genome to guide transcriptome assembly from RNA sequencing data, rather than performing assembly de novo, affects downstream analyses. Yet, direct comparisons of these approaches are rare. Here, we compare the results of the standard de novo assembly pipeline ( Trinity) and two reference genome-based pipelines ( Tuxedo and the new Tuxedo) for differential expression and gene ontology enrichment analysis of a companion study on Atlantic cod (Gadus morhua). ResultsThe new Tuxedo pipeline produced a higher quality assembly than the Tuxedo suite. However, greater enrichment of Trinity-identified differentially expressed genes suggests that a higher proportion of them represent biologically meaningful differences in transcription, as opposed to transcriptional noise or false positives. Coupled with the ability to annotate novel loci, the increased sensitivity of the Trinity pipeline might make it preferable over the reference genome-based approaches for studies aimed at broadly characterizing variation in the magnitude of expression differences and biological processes. However, the new Tuxedo pipeline might be appropriate when a more conservative approach is warranted, such as for the identification of candidate genes.

bioinformatics↗

SlowMoMan: A web app for discovery of important features along user-drawn trajectories in 2D embeddings

Nonlinear low-dimensional embeddings allow humans to visualize high-dimensional data, as is often seen in bioinformatics, where data sets may have tens of thousands of dimensions. However, relating the axes of a nonlinear embedding to the original dimensions is a nontrivial problem. In particular, humans may identify patterns or interesting subsections in the embedding, but cannot easily identify what those patterns correspond to in the original data. Thus, we present SlowMoMan (SLOW Motions on MANifolds), a web application which allows the user to draw a 1-dimensional path onto a 2-dimensional embedding. Then, by back-projecting the manifold to the original, high-dimensional space, we sort the original features such that those most discriminative along the manifold are ranked highly. We show a number of pertinent use cases for our tool, including trajectory inference, spatial transcriptomics, and automatic cell classification. Software availabilityhttps://yunwilliamyu.github.io/SlowMoMan/ Code availabilityhttps://github.com/yunwilliamyu/SlowMoMan

bioinformatics↗