bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Genome ARTIST_v2 software - a support for annotation of class II natural transposons in new sequenced genomes

Transposon annotation is a very dynamic field of genomics and various tools assigned to support this bioinformatics endeavor were reported. Genome ARTIST (GA) software was initially developed for mapping artificial transposons mobilized during insertional mutagenesis projects. Now, the new functions of GA_v2 qualify it as an effective companion for mapping and annotation of class II natural transposons in assembled genomes, contigs or sequencing reads. Tabular export of mapping and annotation data for subsequent high-throughput data analysis, the export of a list of flanking sequences around either the coordinates of insertion or around the target site duplications (TSDs) and generation of a consensus sequence for the respective flanking sequences are all key assets of GA_v2. Additionally, we developed two accompanying short scripts that enable the user to annotate transposons existent in assembled genomes and to use various annotation offered by FlyBase for Drosophila melanogaster genome. Herein, we present the applicability of GA_v2 for a preliminary annotation of the class II transposon P-element in the genome of D. melanogaster strain Horezu, Romania, which was sequenced with Nanopore technology in our laboratory. Our results point that GA_v2 is a reliable tool to be integrated in pipelines designed to perform transposon annotation in new sequenced genomes. GA_v2 is open source software compatible with Ubuntu, Mac OS and Windows and is available at https://github.com/genomeartist/genomeartist and at www.genomeartist.ro.

bioinformatics↗

MetaPop: A pipeline for macro- and micro-diversity analyses and visualization of microbial and viral metagenome-derived populations

BackgroundMicrobes and their viruses are hidden engines driving Earths ecosystems from the oceans and soils to humans and bioreactors. Though gene marker approaches can now be complemented by genome-resolved studies of inter- (macrodiversity) and intra- (microdiversity) population variation, analytical tools to do so remain scattered or under-developed. ResultsHere we introduce MetaPop, an open-source bioinformatic pipeline that provides a single interface to analyze and visualize microbial and viral community metagenomes at both the macro- and micro-diversity levels. Macrodiversity estimates include population abundances and - and {beta}-diversity. Microdiversity calculations include identification of single nucleotide polymorphisms, novel codon-constrained linkage of SNPs, nucleotide diversity ({pi} and {theta}) and selective pressures (pN/pS and Tajimas D) within and fixation indices (FST) between populations. MetaPop will also identify genes with distinct codon usage. Following rigorous validation, we applied MetaPop to the gut viromes of autistic children that underwent fecal microbiota transfers and their neurotypical peers. The macrodiversity results confirmed our prior findings for viral populations (microbial shotgun metagenomes were not available), that diversity did not significantly differ between autistic and neurotypical children. However, by also quantifying microdiversity, MetaPop revealed lower average viral nucleotide diversity ({pi}) in autistic children. Analysis of the percentage of genomes detected under positive selection was also lower among autistic children, suggesting that higher viral {pi} in neurotypical children may be beneficial because it allows populations to better bet hedge in changing environments. Further, comparisons of microdiversity pre- and post-FMT in the autistic children revealed that the delivery FMT method (oral versus rectal) may influence viral activity and engraftment of microdiverse viral populations, with children who received their FMT rectally having higher microdiversity post-FMT. Overall, these results show that analyses at the macro-level alone can miss important biological differences. ConclusionsThese findings suggest that standardized population and genetic variation analyses will be invaluable for maximizing biological inference, and MetaPop provides a convenient tools package to explore the dual impact of macro- and micro-diversity across microbial communities.

bioinformatics↗

Immunopeptidomics for Dummies: Detailed Experimental Protocols and Rapid, User-Friendly Visualization of MHC I and II Ligand Datasets with MhcVizPipe

Immunopeptidomics refers to the science of investigating the composition and dynamics of peptides presented by major histocompatibility complex (MHC) class I and class II molecules using mass spectrometry (MS). Here, we aim to provide a technical report to any non-expert in the field wishing to establish and/or optimize an immunopeptidomic workflow with relatively limited computational knowledge and resources. To this end, we thoroughly describe step-by-step instructions to isolate MHC class I and II-associated peptides from various biological sources, including mouse and human biospecimens. Most notably, we created MhcVizPipe (MVP) (https://github.com/CaronLab/MhcVizPipe), a new and easy-to-use open-source software tool to rapidly assess the quality and the specific enrichment of immunopeptidomic datasets upon the establishment of new workflows. In fact, MVP enables intuitive visualization of multiple immunopeptidomic datasets upon testing sample preparation protocols and new antibodies for the isolation of MHC class I and II peptides. In addition, MVP enables the identification of unexpected binding motifs and facilitates the analysis of non-canonical MHC peptides. We anticipate that the experimental and bioinformatic resources provided herein will represent a great starting point for any non-expert and will therefore foster the accessibility and expansion of the field to ultimately boost its maturity and impact.

bioinformatics↗

BASE: a novel workflow to integrate non-ubiquitous genes in genomics analyses for selection.

Inferring the selective forces that different ortholog genes underwent across different lineages can make us understand the evolutionary processes which shaped their extant diversity. The more widespread metric to estimate coding sequences selection regimes across across their sites and species phylogeny is the ratio of nonsynonymous to synonymous substitutions (dN/dS, also known as{omega} ). Nowadays, modern sequencing technologies and the large amount of already available sequence data allow the retrieval of thousands of genes orthology groups across large numbers of species. Nonetheless, the tools available to explore selection regimes are not designed to automatically process all orthogroups and practical usage is often restricted to those consisting of single-copy genes which are ubiquitous across the species considered (i.e. the subset of genes which is shared by all the species considered). This approach limits the scale of the analysis to a fraction of single-copy genes, which can be as lower as an order of magnitude in respect to non-ubiquitous ones (i.e. those which are not present across all the species considered). Here we present a workflow named BASE that - leveraging the CodeML framework - ease the inference and interpretation of selection regimes in the context of comparative genomics. Although a number of bioinformatics tools have already been developed to facilitate this kind of analyses, BASE is the first to be specifically designed to ease the integration of non-ubiquitous genes orthogroups. The workflow - along with all the relevant documentation - is available at github.com/for-giobbe/BASE.

bioinformatics↗

IRIS-FGM: an integrative single-cell RNA-Seq interpretation system for functional gene module analysis

SummarySingle-cell RNA-Seq (scRNA-Seq) data is useful in discovering cell heterogeneity and signature genes in specific cell populations in cancer and other complex diseases. Specifically, the investigation of functional gene modules (FGM) can help to understand gene interactive networks and complex biological processes. QUBIC2 is recognized as one of the most efficient and effective tools for FGM identification from scRNA-Seq data. However, its limited availability to a C implementation restricted its application to only a few downstream analyses functionalities. We developed an R package named IRIS-FGM (Integrative scRNA-Seq Interpretation System for Functional Gene Module analysis) to support the investigation of FGMs and cell clustering using scRNA-Seq data. Empowered by QUBIC2, IRIS-FGM can effectively identify co-expressed and co-regulated FGMs, predict cell types/clusters, uncover differentially expressed genes, and perform functional enrichment analysis. It is noteworthy that IRIS-FGM can also takes Seurat objects as input, which facilitate easy integration with existing analysis pipeline. Availability and ImplementationIRIS-FGM is implemented in R environment (as of version 3.6) with the source code freely available at https://github.com/OSU-BMBL/IRIS-FGM Contactqin.ma@osumc.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

CRISPRIdentify: Identification of CRISPR arrays using machine learning approach

CRISPR-Cas are adaptive immune systems that degrade foreign genetic elements in archaea and bacteria. In carrying out their immune functions, CRISPR- Cas systems heavily rely on RNA components. These CRISPR (cr) RNAs are repeat-spacer units that are produced by processing of pre-crRNA, the transcript of CRISPR arrays, and guide Cas protein(s) to the cognate invading nucleic acids, enabling their destruction. Several bioinformatics tools have been developed to detect CRISPR arrays based solely on DNA sequences, but all these tools employ the same strategy of looking for repetitive patterns, which might correspond to CRISPR array repeats. The identified patterns are evaluated using a fixed, built-in scoring function, and arrays exceeding a cut-off value are reported. Here, we instead introduce a data-driven approach that uses machine learning to detect and differentiate true CRISPR arrays from false ones based on several features. Our CRISPR detection tool, CRISPRidentify, performs three steps: detection, feature extraction and classification based on manually curated sets of positive and negative examples of CRISPR arrays. The identified CRISPR arrays are then reported to the user accompanied by detailed annotation. We demonstrate that our approach identifies not only previously detected CRISPR arrays, but also CRISPR array candidates not detected by other tools. Compared to other methods, our tool has a drastically reduced false positive rate. In contrast to the existing tools, our approach not only provides the user with the basic statistics on the identified CRISPR arrays but also produces a certainty score as a practical measure of the likelihood that a given genomic region is a CRISPR array.

bioinformatics↗

SwarmTCR: a computational approach to predict the specificity of T Cell Receptors

MotivationComputationally predicting the specificity of T cell receptors can be a powerful tool to shed light on the immune response against infectious diseases and cancers, autoimmunity, cancer immunotherapy, and immunopathology. With more T cell receptor sequence data becoming available, the need for bioinformatics approaches to tackle this problem is even more pressing. Here we present SwarmTCR, a method that uses labeled sequence data to predict the specificity of T cell receptors using a nearest-neighbor approach. SwarmTCR works by optimizing the weights of the individual CDR regions to maximize classification performance. ResultsWe compared the performance of SwarmTCR against a state-of-the-art method (TCRdist) and showed that SwarmTCR performed significantly better on epitopes EBV-BRLF1300, EBV-BRLF1109, NS4B214-222 with single cell data and epitopes EBV-BRLF1300, EBV-BRLF1109, IAV-M158 with bulk sequencing data ( and {beta} chains). In addition, we show that the weights returned by SwarmTCR are biologically interpretable. AvailabilitySwarmTCR is distributed freely under the terms of the GPL-3 license. The source code and all sequencing data are available at GitHub (https://github.com/thecodingdoc/SwarmTCR) Contactdghersi@unomaha.edu

bioinformatics↗

cblaster: a remote search tool for rapid identification and visualisation of homologous gene clusters

Genes involved in coordinated biological pathways, including metabolism, drug resistance and virulence, are often collocalised as gene clusters. Identifying homologous gene clusters aids in the study of their function and evolution, however existing tools are limited to searching local sequence databases. Tools for remotely searching public databases are necessary to keep pace with the rapid growth of online genomic data. Here, we present cblaster, a Python based tool to rapidly detect collocated genes in local and remote databases. cblaster is easy to use, offering both a command line and a user-friendly graphical user interface (GUI). It generates outputs that enable intuitive visualisations of large datasets, and can be readily incorporated into larger bioinformatic pipelines. cblaster is a significant update to the comparative genomics toolbox. cblaster source code and documentation is freely available from GitHub under the MIT license (github.com/gamcil/cblaster).

bioinformatics↗

Optimizing Consensus Generation Algorithms for Highly Variable Amino Acid Sequence Clusters

Producing a functional consensus sequence is a preliminary bioinformatics task, which is a necessity for many research purposes. However, the existence of hypervariable regions in the input multiple sequence alignment files causes complications in generating a useful consensus sequence. The current methods for consensus generation, Threshold, and majority algorithms, have several problems, which exclude them as applicable algorithms for such highly variable sequence clusters. Hence, we designed a novel alternative algorithm for the same purpose. The algorithm was explained both using a mathematical formula and a practical implementation in Python programming language. A sequence set from HCV genotype 1b E2 protein has been utilized as a practical example for evaluating the algorithms performance. A few in silico tests have been performed on the output sequence and the results have been compared to results from other algorithms. Epitope-mapping analysis indicates the functionality of this algorithm, by preserving the hotspot residues in the consensus sequence, and the antigenicity index shows significant antigenicity of the consensus sequence. Moreover, phylogenetic analysis shows no significant change in the placement of the new consensus sequence on the phylogenetic tree compared to other algorithms. This approach will have several implications in designing a new vaccine for highly variable viruses such as HIV-1, Influenza, and Hepatitis C Viruses (HCV).

bioinformatics↗

SPDE: A Multi-functional Software for Sequence Processing and Data Extraction

Efficiently extracting information from biological big data can be a huge challenge for people (especially those who lack programming skills). We developed Sequence Processing and Data Extraction (SPDE) as an integrated tool for sequence processing and data extraction for gene family and omics analyses. Currently, SPDE has seven modules comprising 100 basic functions that range from single gene processing (e.g., translation, reverse complement, and primer design) to genome information extraction. All SPDE functions can be used without the need for programming or command lines. The SPDE interface has enough prompt information to help users run SPDE without barriers. In addition to its own functions, SPDE also incorporates the publicly available analyses tools (such as, NCBI-blast, HMMER, Primer3 and SAMtools), thereby making SPDE a comprehensive bioinformatics platform for big biological data analysis. AvailabilitySPDE was built using Python and can be run on 32-bit, 64-bit Windows and macOS systems. It is an open-source software that can be downloaded from https://github.com/simon19891216/SPDEv1.2.git. Contactxudongzhuanyong@163.com

bioinformatics↗

KusakiDB v1.0: a novel approach for validation and completeness of protein orthologous groups

Plants have quite a low coverage in the major protein databases despite their roughly 350,000 species. Moreover, the agricultural sector is one of the main categories in bioeconomy. In order to manipulate and/or engineer plant-based products, it is important to understand the essential fabric of an organism, its proteins. Therefore, we created KusakiDB, which is a database of orthologous proteins, in plants, that correlates three major databases, OrthoDB, UniProt and RefSeq. KusakiDB has an orthologs assessment and management tools in order to compare orthologous groups, which can provide insights not only under an evolutionary point of view but also evaluate structural gene prediction quality and completeness among plant species. KusakiDB could be a new approach to reduce error propagation of functional annotation in plant species. Additionally, this method could, potentially, bring to light some orthologs unique to a few species or families that could have evolved at a high evolutionary rate or could have been a result of a horizontal gene transfer. Availability and ImplementationThe software is implemented in R. It is available at http://pgdbjsnp.kazusa.or.jp/app/kusakidb and at https://hub.docker.com/r/ghelfi/kusakidb under the MIT license. Contactandreaghelfi@kazusa.or.jp Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

A Bayesian Nonparametric Model for Inferring Subclonal Populations from Structured DNA Sequencing Data

There are distinguishing features or "hallmarks" of cancer that are found across tumors, individuals, and types of cancer, and these hallmarks can be driven by specific genetic mutations. Yet, within a single tumor there is often extensive genetic heterogeneity as evidenced by single-cell and bulk DNA sequencing data. The goal of this work is to jointly infer the underlying genotypes of tumor subpopulations and the distribution of those subpopulations in individual tumors by integrating single-cell and bulk sequencing data. Understanding the genetic composition of the tumor at the time of treatment is important in the personalized design of targeted therapeutic combinations and monitoring for possible recurrence after treatment. We propose a hierarchical Dirichlet process mixture model that incorporates the correlation structure induced by a structured sampling arrangement and we show that this model improves the quality of inference. We develop a representation of the hierarchical Dirichlet process prior as a Gamma-Poisson hierarchy and we use this representation to derive a fast Gibbs sampling inference algorithm using the augment-and-marginalize method. Experiments with simulation data show that our model outperforms standard numerical and statistical methods for decomposing admixed count data. Analyses of real acute lymphoblastic leukemia cancer sequencing dataset shows that our model improves upon state-of-the-art bioinformatic methods. An interpretation of the results of our model on this real dataset reveals co-mutated loci across samples.

bioinformatics↗

FASTAFS: file system virtualisation of random access compressed FASTA files

BackgroundThe FASTA file format used to store polymeric sequence data has become a bioinformatics file standard used for decades. The relatively large files require additional files beyond the scope of the original format, to identify sequences and provide random access. Currently, multiple compressors have been developed to archive FASTA files back and forth, but these lack direct access to targeted content or metadata of the archive. Moreover, these solutions are not directly backwards compatible to FASTA files, resulting in limited software integration. ResultsWe designed linux based a toolkit using Filesystem in Userspace (FUSE) that virtualises the content of DNA, RNA and protein FASTA archives into the filesystem. This guarantees in-sync virtualised metadata files and offers fast random-access decompression using Zstandard (zstd). The toolkit, FASTAFS, can track all system wide running instances, allows file integrity verification and can provide, instantly, scriptable access to sequence files and is easy to use and deploy. ConclusionsFASTAFS is a user-friendly and easy to deploy backwards compatible generic purpose solution to store and access compressed FASTA files, since it offers file system access to FASTA files as well as in-sync metadata files through file virtualisation. Using virtual filesystems as in-between layer offers the possibility to design format conversion without the need to rewrite code into different languages while preserving compatibility. Code Availabilityhttps://github.com/yhoogstrate/fastafs

bioinformatics↗

snpXplorer: a web application to explore SNP-associations and annotate SNP-sets

Genetic association studies are frequently used to study the genetic basis of numerous human phenotypes. However, the rapid interrogation of how well a certain genomic region associates across traits as well as the interpretation of genetic associations is often complex and requires the integration of multiple sources of annotation, which involves advanced bioinformatic skills. We developed snpXplorer, an easy-to-use web-server application for exploring Single Nucleotide Polymorphisms (SNP) association statistics and to functionally annotate sets of SNPs. snpXplorer can superimpose association statistics from multiple studies, and displays regional information including SNP associations, structural variations, recombination rates, eQTL, linkage disequilibrium patterns, genes and gene-expressions per tissue. By overlaying multiple GWAS studies, snpXplorer can be used to compare levels of association across different traits, which may help the interpretation of variant consequences. Given a list of SNPs, snpXplorer can also be used to perform variant-to-gene mapping and gene-set enrichment analysis to identify molecular pathways that are overrepresented in the list of input SNPs. snpXplorer is freely available at https://snpxplorer.net. Source code, documentation, example files and tutorial videos are available within the Help section of snpXplorer and at https://github.com/TesiNicco/snpXplorer. Key points: O_LIsnpXplorer shows GWAS summary statistics, regional information and helps deciphering GWAS outcomes C_LIO_LIsnpXplorer interactively compares association levels of a genomic region across phenotypes C_LIO_LIsnpXplorer performs variant-to-gene mapping and gene-set enrichment analysis C_LI O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=151 SRC="FIGDIR/small/377879v2_ufig1.gif" ALT="Figure 1"> View larger version (47K): org.highwire.dtl.DTLVardef@13a9ba5org.highwire.dtl.DTLVardef@c07b9aorg.highwire.dtl.DTLVardef@f2ce4forg.highwire.dtl.DTLVardef@c6d228_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

A simplified workflow for the analysis of whole-genome sequencing data from Pristionchus pacificus mutant lines

Nematodes are attractive model systems to understand the genetic basis of various biological processes ranging from development to complex behaviors. In particular, mutagenesis experiments combined with whole-genome sequencing has been proven as one of the most effective methods to identify core players of multiple biological pathways. To enable experimentalists to apply such integrative genetic and bioinformatic analysis in the case of the satellite model organism Pristionchus pacificus, I present a simplified workflow for the analysis of whole-genome data from mutant lines and corresponding mapping panels. Individual components are based on well-maintained and widely used software packages and are extended by 50 lines of code for the analysis and visualization of allele frequencies. The effectiveness of this workflow is demonstrated by an application to recently generated data of a P. pacificus mutant line, where it reduced the number of candidate mutations from an initial set of 3,500 single nucleotide variants to ten.

bioinformatics↗

Trecode: a FAIR eco-system for the analysis and archiving of omics data in a combined diagnostic and research setting

MotivationThe increase in speed, reliability and cost-effectiveness of high-throughput sequencing has led to the widespread clinical application of genome (WGS), exome (WXS) and transcriptome analysis. WXS and RNA sequencing is now being implemented as standard of care for patients and for patients included in clinical studies. To keep track of sample relationships and analyses, a platform is needed that can unify metadata for diverse sequencing strategies with sample metadata whilst supporting automated and reproducible analyses. In essence ensuring that analysis is conducted consistently, and data is Findable, Accessible, Interoperable and Reusable (FAIR). ResultsWe present "Trecode", a framework that records both clinical and research sample (meta) data and manages computational genome analysis workflows executed for both settings. Thereby achieving tight integration between analyses results and sample metadata. With complete, consistent and FAIR (meta) data management in a single platform, stacked bioinformatic analyses are performed automatically and tracked by the database ensuring data provenance, reproducibility and reusability which is key in worldwide collaborative translational research. Availability and implementationThe Trecode data model, codebooks, NGS workflows and client programs are currently being cleared from local compute infrastructure dependencies and will become publicly available in spring 2021. Contactp.kemmeren@prinsesmaximacentrum.nl

bioinformatics↗

Fast Alignment-Free Similarity Estimation By Tensor Sketching

The sharp increase in next-generation sequencing technologies capacity has created a demand for algorithms capable of quickly searching a large corpus of biological sequences. The complexity of biological variability and the magnitude of existing data sets have impeded finding algorithms with guaranteed accuracy that efficiently run in practice. Our main contribution is the Tensor Sketch method that efficiently and accurately estimates edit distances. In our experiments, Tensor Sketch had 0.956 Spearmans rank correlation with the exact edit distance, improving its best competitor Ordered MinHash by 23%, while running almost 5 times faster. Finally, all sketches can be updated dynamically if the input is a sequence stream, making it appealing for large-scale applications where data cannot fit into memory. Conceptually, our approach has three steps: 1) represent sequences as tensors over their sub-sequences, 2) apply tensor sketching that preserves tensor inner products, 3) implicitly compute the sketch. The sub-sequences, which are not necessarily contiguous pieces of the sequence, allow us to outperform k-mer-based methods, such as min-hash sketching over a set of k-mers. Typically, the number of sub-sequences grows exponentially with the sub-sequence length, introducing both memory and time overheads. We directly address this problem in steps 2 and 3 of our method. While the sketching of rank-1 or super-symmetric tensors is known to admit efficient sketching, the sub-sequence tensor does not satisfy either of these properties. Hence, we propose a new sketching scheme that completely avoids the need for constructing the ambient space. Our tensor-sketching techniques main advantages are three-fold: 1) Tensor Sketch has higher accuracy than any of the other assessed sketching methods used in practice. 2) All sketches can be computed in a streaming fashion, leading to significant time and memory savings when there is overlap between input sequences. 3) It is straightforward to extend tensor sketching to different settings leading to efficient methods for related sequence analysis tasks. We view tensor sketching as a framework to tackle a wide range of relevant bioinformatics problems, and we are confident that it can bring significant improvements for applications based on edit distance estimation.

bioinformatics↗

CRISPAltRations: a validated cloud-based approach for interrogation of double-strand break repair mediated by CRISPR genome editing

CRISPR systems enable targeted genome editing in a wide variety of organisms by introducing single- or double-strand DNA breaks, which are repaired using endogenous molecular pathways. Characterization of on- and off-target editing events from CRISPR proteins can be evaluated using targeted genome resequencing. We characterized DNA repair footprints that result from non-homologous end joining (NHEJ) after double stranded breaks (DSBs) were introduced by Cas9 or Cas12a for >500 paired treatment/control experiments. We found that building our understanding into a novel analysis tool (CRISPAltRations) improved results quality. We validated our software using simulated rhAmpSeq amplicon sequencing data (11 gRNAs and 603 on- and off-target locations) and demonstrate that CRISPAltRations outperforms other publicly available software tools in accurately annotating CRISPR-associated indels and homology directed repair (HDR) events. We enable non-bioinformaticians to use CRISPAltRations by developing a web-accessible, cloud-hosted deployment, which allows rapid batch processing of samples in a graphical user-interface (GUI) and complies with HIPAA security standards. By ensuring that our software is thoroughly tested, version controlled, and supported with a UI we enable resequencing analysis of CRISPR genome editing experiments to researchers no matter their skill in bioinformatics.

bioinformatics↗