bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96Linked to original sources

State-of-the-art structural variant calling: What went conceptually wrong and how to fix it?

Structural variant (SV) calling belongs to the standard tools of modern bioinformatics for identifying and describing alterations in genomes. Initially, this work presents several complex genomic rearrangements that reveal conceptual ambiguities inherent to the SV representations of state-of-the-art SV callers. We contextualize these ambiguities theoretically as well as practically and propose a graph-based approach for resolving them. Our graph model unifies both genomic strands by using the concept of skew-symmetry; it supports graph genomes in general and pan genomes in specific. Instances of our model are inferred directly from seeds instead of the commonly used alignments that conflict with various types of SV as reported here. For yeast genomes, we practically compute adjacency matrices of our graph model and demonstrate that they provide highly accurate descriptions of one genome in terms of another. An open-source prototype implementation of our approach is available under the MIT license at https://github.com/ITBE-Lab/MA.

bioinformatics↗

An optimized FM-index library for nucleotide and amino acid search

Pattern matching is a key step in a variety of biological sequence analysis pipelines. The FM-index is a compressed data structure for pattern matching, with search run time that is independent of the length of the database text. We present AvxWindowedFMindex (AWFM-index), an open-source, thread-parallel FM-index library written in C that is optimized for indexing nucleotide and amino acid sequences. AWFM-index is easy to incorporate into bioinformatics software and is able to perform exact match count and locate queries ~2-4x faster than SeqAn3s FM-index implementation for nucleotide search, and ~2-6x faster for amino acid search in a single-threaded context. This performance is due to (i) a new approach to storing FM-index data in a strided bit-vector format that enables extremely efficient computation of the FM-index occurrence function via AVX2 bitwise instructions, and (ii) inclusion of a cache-efficient lookup table for partial k-mer searches. AWFM-index also trivially parallelizes to multiple threads with good scaling, and enables efficient on-disk storage of the memory-intensive suffix array. The open-source library is available for download at https://github.com/TravisWheelerLab/AvxWindowFmIndex.

bioinformatics↗

The statistics of k-mers from a sequence undergoing a simple mutation process without spurious matches

K-mer-based methods are widely used in bioinformatics, but there are many gaps in our understanding of their statistical properties. Here, we consider the simple model where a sequence S (e.g. a genome or a read) undergoes a simple mutation process whereby each nucleotide is mutated independently with some probability r, under the assumption that there are no spurious k-mer matches. How does this process affect the k-mers of S? We derive the expectation and variance of the number of mutated k-mers and of the number of islands (a maximal interval of mutated k-mers) and oceans (a maximal interval of non-mutated k-mers). We then derive hypothesis tests and confidence intervals for r given an observed number of mutated k-mers, or, alternatively, given the Jaccard similarity (with or without minhash). We demonstrate the usefulness of our results using a few select applications: obtaining a confidence interval to supplement the Mash distance point estimate, filtering out reads during alignment by Minimap2, and rating long read alignments to a de Bruijn graph by Jabba.

bioinformatics↗

SAINT: automatic taxonomy embedding and categorization by Siamese triplet network

MotivationUnderstanding the phylogenetic relationship among organisms is the key in contemporary evolutionary study and sequence analysis is the workhorse towards this goal. Conventional approaches to sequence analysis are based on sequence alignment, which is neither scalable to large-scale datasets due to computational inefficiency nor adaptive to next-generation sequencing (NGS) data. Alignment-free approaches are typically used as computationally effective alternatives yet still suffering the high demand of memory consumption. One desirable sequence comparison method at large-scale requires succinctly-organized sequence data management, as well as prompt sequence retrieval given a never-before-seen sequence as query. ResultsIn this paper, we proposed a novel approach, referred to as SAINT, for efficient and accurate alignment-free sequence comparison. Compared to existing alignment-free sequence comparison methods, SAINT offers advantages in two aspects: (1) SAINT is a weakly-supervised learning method where the embedding function is learned automatically from the easily-acquired data; (2) SAINT utilizes the non-linear deep learning-based model which potentially better captures the complicated relationship among genome sequences. We have applied SAINT to real-world datasets to demonstrate its empirical utility, both qualitatively and quantitatively. Considering the extensive applicability of alignment-free sequence comparison methods, we expect SAINT to motivate a more extensive set of applications in sequence comparison at large scale. AvailabilityThe open source, Apache licensed, python-implemented code will be available upon acceptance. Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

MS2AI: Automated repurposing of public peptide LC-MS data for ML applications

MotivationLiquid-chromatography mass-spectrometry (LC-MS) is the established standard for analyzing the proteome in biological samples by identification and quantification of thousands of proteins. Machine learning (ML) promises to considerably improve the analysis of the resulting data, however, there is yet to be any tool that mediates the path from raw data to modern ML applications. More specifically, ML applications are currently hampered by three major limitations: (1) absence of balanced training data with large sample size; (2) unclear definition of sufficiently information-rich data representations for e.g. peptide identification; (3) lack of benchmarking of ML methods on specific LC-MS problems. ResultsWe created the MS2AI pipeline that automates the process of gathering vast quantities of mass spectrometry (MS) data for large scale ML applications. The software retrieves raw data from either in-house sources or from the proteomics identifications database, PRIDE. Subsequently, the raw data is stored in a standardized format amenable for ML encompassing MS1/MS2 spectra and peptide identifications. This tool bridges the gap between MS and AI, and to this effect we also present an ML application in the form of a convolutional neural network for the identification of oxidized peptides. AvailabilityAn open source implementation of the software can be found freely available for non-commercial use at https://gitlab.com/roettgerlab/ms2ai. Contactveits@bmb.sdu.dk Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Machine Translation between paired Single Cell Multi Omics Data

BackgroundSingle-cell multi-omics technologies allow the profiling of different data modalities from the same cell. However, while isolated modalities only capture one view of the total information of a biological cell, an integrative analysis capturing the different modalities is challenging. In response, bioinformatics and machine learning methodologies have been developed for multi-omics single-cell analysis. Nevertheless, it is unclear if current tools can address the dual aspect of modality integration and prediction across modalities without requiring extensive parameter finetuning. ResultsWe designed LIBRA, a Neural Network based framework, to learn a translation between paired multi-omics profiles such that a shared latent space is constructed. LIBRA is a state-of-the-art tool when evaluating the ability to increase cell-type (clustering) resolution in the latent space. When assessing the predictive power across data modalities, LIBRA outperforms existing tools. Finally, considering the importance of hyperparameters, we implemented an adaptative-tuning strategy, labelled aLIBRA, in the LIBRA package. As expected, adaptive parameter optimization significantly boosts the performance of learning predictive models from paired datasets. Additionally, aLIBRA provides parameter combinations balancing the integrative and predictive tasks. ConclusionsLIBRA is a versatile tool, uniquely targeting both integration and prediction tasks of Single-cell multi-omics data. LIBRA is a data-driven robust platform that includes an adaptive learning scheme. Furthermore, LIBRA is freely available as R and Python libraries (https://github.com/TranslationalBioinformaticsUnit/LIBRA).

bioinformatics↗

Strobemers: an alternative to k-mers for sequence comparison

K-mer-based methods are widely used in bioinformatics for various types of sequence comparison. However, a single mutation will mutate k consecutive k-mers and makes most k-mer based applications for sequence comparison sensitive to variable mutation rates. Many techniques have been studied to overcome this sensitivity, e.g., spaced k-mers and k-mer permutation techniques, but these techniques do not handle indels well. For indels, pairs or groups of small k-mers are commonly used, but these methods first produce k-mer matches, and only in a second step, a pairing or grouping of k-mers is performed. Such techniques produce many redundant k-mer matches due to the size of k. Here, we propose strobemers as an alternative to k-mers for sequence comparison. Intuitively, strobemers consist of linked minimizers. We use simulated data to show that strobemers provide more evenly distributed sequence matches and are less sensitive to different mutation rates than k-mers and spaced k-mers. Strobemers also produce a higher match coverage across sequences. We further implement a proof-of-concept sequence matching tool StrobeMap, and use synthetic and biological Oxford Nanopore sequencing data to show the utility of using strobemers for sequence comparison in different contexts such as sequence clustering and alignment scenarios. A reference implementation of our tool StrobeMap together with code for analyses is available at https://github.com/ksahlin/strobemers.

bioinformatics↗

ZoomQA: Residue-Level Single-Model QA Support Vector Machine Utilizing Sequential and 3D Structural Features

MotivationThe Estimation of Model Accuracy problem is a cornerstone problem in the field of Bioinformatics. When predictions are made for proteins of which we do not know the native structure, we run into an issue to tell how good a tertiary structure prediction is, especially the protein binding regions, which are useful for drug discovery. Currently, most methods only evaluate the overall quality of a protein decoy, and few can work on residue level and protein complex. Here we introduce ZoomQA, a novel, single-model method for assessing the accuracy of a tertiary protein structure / complex prediction at residue level. ZoomQA differs from others by considering the change in chemical and physical features of a fragment structure (a portion of a protein within a radius r of the target amino acid) as the radius of contact increases. Fourteen physical and chemical properties of amino acids are used to build a comprehensive representation of every residue within a protein and grades their placement within the protein as a whole. Moreover, ZoomQA can evaluate the quality of protein complex, which is unique. ResultsWe benchmark ZoomQA on CASP14, it outperforms other state of the art local QA methods and rivals state of the art QA methods in global prediction metrics. Our experiment shows the efficacy of these new features, and shows our method is able to match the performance of other state-of-the-art methods without the use of homology searching against database or PSSM matrix. Availabilityhttp://zoomQA.renzhitech.com Contactcaora@plu.edu

bioinformatics↗

HashSeq: A Simple, Scalable, and Conservative De Novo Variant Caller for 16S rRNA Gene Datasets

16S rRNA gene sequencing is a common and cost-effective technique for characterization of microbial communities. Recent bioinformatics methods enable high-resolution detection of sequence variants of only one nucleotide difference. In this manuscript, we utilize a very fast HashMap-based approach to detect sequence variants in six publicly available 16S rRNA gene datasets. We then use the normal distribution combined with LOESS regression to estimate background error rates as a function of sequencing depth for individual clusters of sequences. This method is computationally efficient and produces inference that yields sets of variants that are conservative and well supported by reference databases. We argue that this approach to inference is fast, simple, scalable to large datasets, and provides a high-resolution set of sequence variants which are less likely to be the result of sequencing error.

bioinformatics↗

Telomere maintenance pathway activity analysis enables tissue- and gene-level inferences

Telomere maintenance is one of the mechanisms ensuring indefinite divisions of cancer and stem cells. Good understanding of telomere maintenance mechanisms (TMM) is important for studying cancers and designing therapies. However, molecular factors triggering selective activation of either the telomerase dependent (TEL) or the alternative lengthening of telomeres (ALT) pathway are poorly understood. In addition, more accurate and easy-to-use methodologies are required for TMM phenotyping. In this study, we have performed literature based reconstruction of signaling pathways for the ALT and TEL TMMs. Gene expression data were used for computational assessment of TMM pathway activities and compared with experimental assays for TEL and ALT. Explicit consideration of pathway topology makes bioinformatics analysis more informative compared to computational methods based on simple summary measures of gene expression. Application to healthy human tissues showed high ALT and TEL pathway activities in testis, and identified genes and pathways that may trigger TMM activation. Our approach offers a novel option for systematic investigation of TMM activation patterns across cancers and healthy tissues for dissecting pathway-based molecular markers with diagnostic impact.

bioinformatics↗

Discovery of 17 conserved structural RNAs in fungi

Many non-coding RNAs with known functions are structurally conserved: their intramolecular secondary and tertiary interactions are maintained across evolutionary time. Consequently, the presence of conserved structure in multiple sequence alignments can be used to identify candidate functional non-coding RNAs. Here, we present a bioinformatics method that couples iterative homology search with covariation analysis to assess whether a genomic region has evidence of conserved RNA structure. We used this method to examine all unannotated regions of five well-studied fungal genomes (Saccharomyces cerevisiae, Candida albicans, Neurospora crassa, Aspergillus fumigatus, and Schizosaccharomyces pombe). We identified 17 novel structurally conserved non-coding RNA candidates, which include 4 H/ACA box small nucleolar RNAs, 4 intergenic RNAs, and 9 RNA structures located within the introns and untranslated regions (UTRs) of mRNAs. For the two structures in the 3' UTRs of the metabolic genes GLY1 and MET13, we performed experiments that provide evidence against them being eukaryotic riboswitches.

bioinformatics↗

Genotyping Copy Number Alterations from single-cell RNA sequencing

MotivationCancers are composed by several heterogeneous subpopulations, each one harbouring different genetic and epigenetic somatic alterations that contribute to disease onset and therapy response. In recent years, copy number alterations leading to tumour aneuploidy have been identified as potential key drivers of such populations, but the definition of the precise makeup of cancer subclones from sequencing assays remains challenging. In the end, little is known about the mapping between complex copy number alterations and their effect on cancer phenotypes. ResultsWe introduce CONGAS, a Bayesian probabilistic method to phase bulk DNA and single-cell RNA measurements from independent assays. CONGAS jointly identifies clusters of single cells with subclonal copy number alterations, and differences in RNA expression. The model builds statistical priors leveraging bulk DNA sequencing data, does not require a normal reference and scales fast thanks to a GPU backend and variational inference. We test CONGAS on both simulated and real data, and find that it can determine the tumour subclonal composition at the single-cell level together with clone-specific RNA phenotypes in tumour data generated from both 10x and Smart-Seq assays. AvailabilityCONGAS is available as 2 packages: CONGAS (https://github.com/caravagnalab/congas), which implements the model in Python, and RCONGAS (https://caravagnalab.github.io/rcongas/), which provides R functions to process inputs, outputs, and run CONGAS fits. The analysis of real data and scripts to generate figures of this paper are available via RCONGAS; code associated to simulations is available at https://github.com/caravagnalab/rcongas_test. Contactgcaravagna@units.it Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

ChIPuana: from raw data to epigenomic dynamics

We present ChIPflow, a Snakemake-based pipeline for epigenomic data from the raw fastq files to the differential analysis. It can be applied to any chromatin factor, e.g. histone modification or transcription factor, which can be profiled with ChIP-seq. ChIPflow streamlines critical steps like the quality assessment of the immunoprecipitation using cross-correlation and the replicate comparison for both narrow and broad peaks. For the differential analysis ChIPflow provides linear and nonlinear methods for normalisation between samples as well as conservative and stringent models for estimating the variance and testing the significance of the observed binding/marking differences. ChIPflow can process in parallel multiple chromatin factors with different experimental designs, number of biological replicates and/or conditions. It also facilitates the specific parametrisation of each dataset allowing both narrow or broad peak calling, as well as comparisons between the conditions using multiple statistical settings. Finally, complete reports are produced at the end of the bioinformatic and the statistical part of the analysis, which facilitate the data quality control and the interpretation of the results. We explored the discriminative power of the statistical settings for the differential analysis, using a published dataset of three histone marks (H3K4me3, H3K27ac and H3K4me1) and two transcription factors (Oct4 and Klf4) profiled with ChIP-seq in two biological conditions (shControl and shUbc9). We show that distinct results are obtained depending on the sources of ChIP-seq variability and the dynamics of the chromatin factor under study. We propose that ChIPflow can be used to measure the richness of the epigenomic landscape underlying a biological process by identifying diverse regulatory regimes and the associated genes sets.

bioinformatics↗

SplicingFactory - Splicing diversity analysis for transcriptome data

MotivationAlternative splicing contributes to the diversity of RNA found in biological samples. Current tools investigating patterns of alternative splicing check for coordinated changes in the expression or relative ratio of RNA isoforms where specific isoforms are up- or downregulated in a condition. However, the molecular process of splicing is stochastic and changes in RNA isoform diversity for a gene might arise between samples or conditions. A specific condition can be dominated by a single isoform, while multiple isoforms with similar expression levels can be present in a different condition. These changes might be the result of mutations, drug treatments or differences in the cellular or tissue environment. Here, we present a tool for the characterization and analysis of RNA isoform diversity using isoform level expression measurements. ResultsWe developed an R package called SplicingFactory, to calculate various RNA isoform diversity metrics, and compare them across conditions. Using the package, we tested the effect of RNA-seq quantification tools, quantification uncertainty, gene expression levels, and isoform numbers on the isoform diversity calculation. We analyzed a set of CD34+ hematopoietic stem cells and myelodysplastic syndrome samples and found a set of genes whose isoform diversity change is associated with SF3B1 mutations. Availability and implementationThe SplicingFactory package is freely available under the GPL-3.0 license from Bioconductor for the Windows, MacOS and Linux operating systems (https://www.bioconductor.org/packages/release/bioc/html/SplicingFactory.html). Contactsebestyen.endre@med.semmelweis-univ.hu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Practical selection of representative sets of RNA-seq samples using a hierarchical approach

Despite numerous RNA-seq samples available at large databases, most RNA-seq analysis tools are evaluated on a limited number of RNA-seq samples. This drives a need for methods to select a representative subset from all available RNA-seq samples to facilitate comprehensive, unbiased evaluation of bioinformatics tools. In sequence-based approaches for representative set selection (e.g. a k-mer counting approach that selects a subset based on k-mer similarities between RNA-seq samples), because of the huge number of available RNA-seq samples and the large number of k-mers/sequences in each sample, computing the full similarity matrix between all samples using k-mers/sequences for the entire set of RNA-seq samples in a large database (e.g. the SRA) has memory and runtime challenges, making direct representative set selection infeasible with limited computing resources. Therefore, we developed a novel computational method called "hierarchical representative set selection" to handle this challenge. Hierarchical representative set selection is a divide-and-conquer-like algorithm that breaks the representative set selection into sub-selections and hierarchically selects representative samples through multiple levels. We demonstrate that hierarchical representative set selection can achieve performance close to that of direct representative set selection, while largely reducing the runtime and memory requirements of computing the full similarity matrix (up to 8.4X runtime reduction and 4.7X memory reduction for 10000 samples that could be practically run with direct subset selection). We show that hierarchical representative set selection substantially outperforms random sampling on the entire SRA set of RNA-seq samples, making it a practical solution to representative set selection on large databases such as the SRA.

bioinformatics↗

Epstein-Barr virus long non-coding RNA RPMS1 full-length spliceome in transformed epithelial tissue

Epstein-Barr virus is associated with two types of epithelial neoplasms, nasopharyngeal carcinoma and gastric adenocarcinoma. The viral long non-coding RNA RPMS1 is the most abundantly expressed poly-adenylated viral RNA in these malignant tissues. The RPMS1 gene is known to contain two cassette exons, exon Ia and Ib, and several alternative splicing variants have been described in low-throughput studies. To characterize the entire RPMS1 spliceome we combined long-read sequencing data from the nasopharyngeal cell line C666-1 and a primary gastric adenocarcinoma, with complementary short-read sequencing datasets. We developed FLAME, a Python-based bioinformatics package that can generate complete high resolution characterization of RNA splicing at full-length. Using FLAME, we identified 32 novel exons in the RPMS1 gene, primarily within the large constitutive exons III, V and VII. Two of the novel exons contained retention of the intron between exon III and exon IV, and a novel cassette exon was identified between VI and exon VII. All previously described transcript variants of RPMS1 containing putative ORFs were identified at various levels. Similarly, native transcripts with the potential to form previously reported circular RNA elements were detected. Our work illuminates the multifaceted nature of viral transcriptional repertoires. FLAME provides a comprehensive overview of the relative abundance of alternative splice variants and allows for a wealth of previously unknown splicing events to be unveiled.

bioinformatics↗

Customized de novo mutation detection for any variant calling pipeline: SynthDNM

MotivationAs sequencing technologies and analysis pipelines evolve, DNM calling tools must be adapted. Therefore, a flexible approach is needed that can accurately identify de novo mutations from genome or exome sequences from a variety of datasets and variant calling pipelines. ResultsHere, we describe SynthDNM, a random-forest based classifier that can be readily adapted to new sequencing or variant-calling pipelines by applying a flexible approach to constructing simulated training examples from real data. The optimized SynthDNM classifiers predict de novo SNPs and indels with robust accuracy across multiple methods of variant calling. AvailabilitySynthDNM is freely available on Github (https://github.com/james-guevara/synthdnm) Contactjsebat@ucsd.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Learning Sparse Log-Ratios for High-Throughput Sequencing Data

The automatic discovery of sparse biomarkers that are associated with an outcome of interest is a central goal of bioinformatics. In the context of high-throughput sequencing (HTS) data, and compositional data (CoDa) more generally, an important class of biomarkers are the log-ratios between the input variables. However, identifying predictive log-ratio biomarkers from HTS data is a combinatorial optimization problem, which is computationally challenging. Existing methods are slow to run and scale poorly with the dimension of the input, which has limited their application to low- and moderate-dimensional metagenomic datasets. Building on recent advances from the field of deep learning, we present CoDaCoRe, a novel learning algorithm that identifies sparse, interpretable, and predictive log-ratio biomarkers. Our algorithm exploits a continuous relaxation to approximate the underlying combinatorial optimization problem. This relaxation can then be optimized efficiently using the modern ML toolbox, in particular, gradient descent. As a result, CoDaCoRe runs several orders of magnitude faster than competing methods, all while achieving state-of-the-art performance in terms of predictive accuracy and sparsity. We verify the outperformance of CoDaCoRe across a wide range of microbiome, metabolite, and microRNA benchmark datasets, as well as a particularly high-dimensional dataset that is outright computationally intractable for existing sparse log-ratio selection methods.1

bioinformatics↗