bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

A cloud-based platform for the analysis of single cell RNA sequencing data.

MotivationSingle-cell RNA sequencing (scRNA-seq) is a recent technology that has provided many valuable biological insights. Notable uses include identifying novel cell-types, measuring the cellular response to treatment, and tracking trajectories of distinct cell lineages in time. The raw data generated in this process typically amounts to hundreds of millions of sequencing reads and requires substantial computational infrastructure for downstream analysis, a major hurdle for a biological research lab. Fortunately, the preprocessing step that converts this huge sequence data into manageable cell-specific expression profiles is standardized and can be performed in the cloud. We demonstrate how a cloud-based computational framework can be used to transform the raw data into biologically interpretable cell-type-specific information, using either 3 or 5 transcriptome libraries from 10x Genomics. The processed data which is an order of magnitude smaller in size can be easily downloaded to a laptop for customized analysis to gain deeper biological insights. ResultsWe produced an automated and easily extensible pipeline in the cloud for the analysis of single-cell RNA-seq data which provides a convenient method to handle post-processing of scRNA sequencing using next generation sequencing platforms. The basic step provides the transformation of the scRNA-seq data to cell-type-specific expression profiles and computes the quality control metrics for the dataset. The extensibility of the platform is demonstrated by adding a doublet-removal algorithm and recomputing the clustering of the cells. Any additional computational steps that take a cell-type expression counts matrix as input can be easily added to this framework with minimal effort. AvailabilityThe framework and its documentation for installation is available at the Github repository http://github.com/nj3252/CB-Source/ Contactkyun@houstonmethodist.org Supplementary informationSupplementary data available at Bioinformatics online.

bioinformatics↗

A novel algorithm to accurately classify metagenomic sequences

Widespread availability of next-generation sequencing (NGS) technologies has prompted a recent surge in interest in the microbiome. As a consequence, metagenomics is a fast growing field in bioinformatics and computational biology. An important problem in analyzing metagenomic sequenced data is to identify the microbes present in the sample and figure out their relative abundances. In this article we propose a highly efficient algorithm dubbed as "Hybrid Metagenomic Sequence Classifier" (HMSC) to accurately detect microbes and their relative abundances in a metagenomic sample. The algorithmic approach is fundamentally different from other state-of-the-art algorithms currently existing in this domain. HMSC judiciously exploits both alignment-free and alignment-based approaches to accurately characterize metagenomic sequenced data. To demonstrate the effectiveness of HMSC we used 8 metagenomic sequencing datasets (2 mock and 6 in silico bacterial communities) produced by 3 different sequencing technologies (e.g., HiSeq, MiSeq, and NovaSeq) with realistic error models and abundance distribution. Rigorous experimental evaluations show that HMSC is indeed an effective, scalable, and efficient algorithm compared to the other state-of-the-art methods in terms of accuracy, memory, and runtime. Availability of data and materialsThe implementations and the datasets we used are freely available for non-commercial purposes. They can be downloaded from: https://drive.google.com/drive/folders/132k5E5xqpkw7olFjzYwjWNjyHFrqJITe?usp=sharing

bioinformatics↗

Parallelized calculation of permutation tests

MotivationPermutation tests offer a straight forward framework to assess the significance of differences in sample statistics. A significant advantage of permutation tests are the relatively few assumptions about the distribution of the test statistic are needed, as they rely on the assumption of exchangeability of the group labels. They have great value, as they allow a sensitivity analysis to determine the extent to which the assumed broad sample distribution of the test statistic applies. However, in this situation, permutation tests are rarely applied because the running time of naive implementations is too slow and grows exponentially with the sample size. Nevertheless, continued development in the 1980s introduced dynamic programming algorithms that compute exact permutation tests in polynomial time. Albeit this significant running time reduction, the exact test has not yet become one of the predominant statistical tests for medium sample size. Here, we propose a computational parallelization of one such dynamic programming-based permutation test, the Green algorithm, which makes the permutation test more attractive. ResultsParallelization of the Green algorithm was found possible by nontrivial rearrangement of the structure of the algorithm. A speed-up - by orders of magnitude - is achievable by executing the parallelized algorithm on a GPU. We demonstrate that the execution time essentially becomes a non-issue for sample sizes, even as high as hundreds of samples. This improvement makes our method an attractive alternative to, e.g., the widely used asymptotic Mann-Whitney U-test. AvailabilityIn Python 3 code from the GitHub repository https://github.com/statisticalbiotechnology/parallelPermutationTest under an Apache 2.0 license. Contactlukask@kth.se Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

CancerMIRNome: a web server for interactive analysis and visualization of cancer miRNome data

MicroRNAs (miRNAs), which play critical roles in gene regulatory networks, have emerged as promising diagnostic and prognostic biomarkers for human cancer. In particular, circulating miRNAs that are secreted into circulation exist in remarkably stable forms, and have enormous potential to be leveraged as non-invasive biomarkers for early cancer detection. Novel and user-friendly tools are desperately needed to facilitate data mining of the vast amount of miRNA expression data from The Cancer Genome Atlas (TCGA) and large-scale circulating miRNA profiling studies. To fill this void, we developed CancerMIRNome, a comprehensive database for the interactive analysis and visualization of miRNA expression profiles based on 10,998 samples from 33 TCGA projects and 21,993 samples from 40 public circulating miRNome datasets. A series of cutting-edge bioinformatics tools and machine learning algorithms have been packaged in CancerMIRNome, allowing for the pan-cancer analysis of a miRNA of interest across multiple cancer types and the comprehensive analysis of miRNome profiles to identify dysregulated miRNAs and develop diagnostic or prognostic signatures. The data analysis and visualization modules will greatly facilitate the exploit of the valuable resources and promote translational application of miRNA biomarkers in cancer. The CancerMIRNome database is publicly available at http://bioinfo.jialab-ucr.org/CancerMIRNome.

bioinformatics↗

Expanding the Galaxy's reference data

SummaryProperly and effectively managing reference datasets is an important task for many bioinformatics analyses. Refgenie is a reference asset management system that allows to easily organize, retrieve, and share such datasets. Here, we describe the integration of refgenie into the Galaxy platform. Server administrators are able to configure Galaxy to make use of reference datasets made available on a refgenie instance. Additionally, a Galaxy Data Manager tool has been developed to provide a graphical interface to refgenies remote reference retrieval functionality. A large collection of reference datasets has also been made available using the CVMFS repository from GalaxyProject.org, with mirrors across the United States, Canada, Europe, and Australia, enabling easy use outside of Galaxy. Availability and implementationThe ability of Galaxy to use refgenie assets was added to the core Galaxy framework in version 20.05, which is available from https://github.com/galaxyproject/galaxy under the Academic Free License version 3.0. The refgenie Data Manager tool can be installed via the Galaxy ToolShed, with source code managed at https://github.com/BlankenbergLab/galaxy-tools-blankenberg/tree/main/data_managers/data_manager_refgenie_pull and released using an MIT license.

bioinformatics↗

HextractoR: an R package for automatic extraction of hairpins from genome-wide data

Extracting stem-loop sequences (hairpins) from genome-wide data is very important nowadays for some data mining tasks in bioinformatics. The genome preprocessing is very important because it has a strong influence on the later steps and the final results. For example, for novel miRNA prediction, all well-known hairpins must be properly located. Although there are some scripts that can be adapted and put together to achieve this task, they are outdated, none of them guarantees finding correspondence to well-known structures in the genome under analysis, and they do not take advantage of the latest advances in secondary structure prediction. We present here an R package for automatic extraction of hairpins from genome-wide data (HextractorR). HextractoR makes an exhaustive and smart analysis of the genome in order to obtain a very good set of short sequences for further processing. Moreover, genomes can be processed in parallel and with low memory requirements. Results obtained showed that HextractoR has effectively outperformed other methods. HextractoR it is freely available at CRAN and Sourceforge.

bioinformatics↗

Gene-level metagenomics identifies genome islands associated with immunotherapy response

Researchers must be able to generate experimentally testable hypotheses from sequencing-based observational microbiome experiments to discover the mechanisms underlying the influence of gut microbes on human health. We describe a novel bioinformatics tool for identifying testable hypotheses based on gene-level metagenomic analysis of WGS microbiome data (geneshot). By applying geneshot to two independent previously published cohorts, we identified microbial genomic islands consistently associated with response to immune checkpoint inhibitor (ICI)-based cancer treatment in culturable type strains. The identified genomic islands are within operons involved in type II secretion, TonB-dependent transport, and bacteriophage growth. These results, as well as the underlying methodology, inform further mechanistic studies and facilitate the development of microbiome-enhanced therapeutics.

bioinformatics↗

Compact and evenly distributed k-mer binning for genomic sequences

The processing of k-mers (subsequences of length k) is at the foundation of many sequence processing algorithms in bioinformatics, including k-mer counting for genome size estimation, genome assembly, and taxonomic classification for metagenomics. Minimizers - ordered m-mers where m < k - are often used to group k-mers into bins as a first step in such processing. However, minimizers are known to generate bins of very different sizes, which can pose challenges for distributed and parallel processing, as well as generally increase memory requirements. Furthermore, although various minimizer orderings have been proposed, their practical value for improving tool efficiency has not yet been fully explored. Here we present Discount, a distributed k-mer counting tool based on Apache Spark, which we use to investigate the behaviour of various minimizer orderings in practice when applied to metagenomics data. Using this tool, we then introduce the universal frequency ordering, a new combination of frequency counted minimizers and universal k-mer hitting sets, which yields both evenly distributed binning and small bin sizes. We show that this ordering allows Discount to perform distributed k-mer counting on a large dataset in as little as 1/8 of the memory of comparable approaches, making it the most efficient out-of-core distributed k-mer counting method available.

bioinformatics↗

Multivariate Meta-Analysis of Differential Principal Components underlying Human Primed and Naive-like Pluripotent States

The ground or naive pluripotent state of human pluripotent stem cells (hPSCs), which was initially established in mouse embryonic stem cells (mESCs), is an emerging and tentative concept. To verify this important concept in hPSCs, we performed a multivariate meta-analysis of major hPSC datasets via the combined analytic powers of percentile normalization, principal component analysis (PCA), t-distributed stochastic neighbor embedding (t-SNE), and SC3 consensus clustering. This vigorous bioinformatics approach has significantly improved the predictive values of the current meta-analysis. Accordingly, we were able to reveal various similarities between some naive-like hPSCs (NLPs) and their human and mouse in vitro counterparts. Moreover, we also showed numerous fundamental inconsistencies between diverse naive-like states, which are likely attributed to interlaboratory protocol differences. Collectively, our meta-analysis failed to provide global transcriptomic markers that support a bona fide human naive pluripotent state, rather suggesting the existence of altered pluripotent states under current naive-like growth protocols.

bioinformatics↗

Cuttlefish: Fast, parallel, and low-memory compaction of de Bruijn graphs from large-scale genome collections

MotivationThe construction of the compacted de Bruijn graph from collections of reference genomes is a task of increasing interest in genomic analyses. These graphs are increasingly used as sequence indices for short and long read alignment. Also, as we sequence and assemble a greater diversity of genomes, the colored compacted de Bruijn graph is being used as the basis for efficient methods to perform comparative genomic analyses on these genomes. Therefore, designing time and memory efficient algorithms for the construction of this graph from reference sequences is an important problem. ResultsWe introduce a new algorithm, implemented in the toolCuttlefish, to construct the (colored) compacted de Bruijn graph from a collection of one or more genome references. Cuttlefish introduces a novel approach of modeling de Bruijn graph vertices as finite-state automata; it constrains these automatas state-space to enable tracking their transitioning states with very low memory usage. Cuttlefish is fast and highly parallelizable. Experimental results demonstrate that it scales much better than existing approaches, especially as the number and the scale of the input references grow. On our test hardware, Cuttlefish constructed the graph for 100 human genomes in under 9 hours, using ~29 GB of memory while no other tested tool completed this task. On 11 diverse conifer genomes, the compacted graph was constructed by Cuttlefish in under 9 hours, using ~84 GB of memory, while the only other tested tool that completed this construction on our hardware took over 16 hours and ~289 GB of memory. AvailabilityCuttlefish is written in C++14, and is available under an open source license at https://github.com/COMBINE-lab/cuttlefish. Contactrob@cs.umd.edu Supplementary informationSupplementary text are available at Bioinformatics online.

bioinformatics↗

Databiology Lab CORONAHACK: Collection of Public COVID-19 Data

COVID-19 has had an unprecedented global impact in health and economy affecting millions of persons world-wide. To support and enable a collaborative response from the global research communities, we created a data collection for different public sources for anonymized patient clinical data, imaging datasets, molecular data as nucleotide and protein sequences for the SARS-CoV-2 virus, reports of count of cases and deaths per city/country, and other economic indicators in Databiology Lab (https://www.lab.databiology.net/) where researchers could access these data assets and use the hundreds of available open source bioinformatic applications to analyze them. These data assets are regularly updated and was used in a successful virtual 3-day hackathon organized by Databiology Ltd and Mindstream-AI where hundreds of attendees to work collaboratively to analyze these data collections.

bioinformatics↗

In silico secretome characterization of clinical Mycobacterium abscessus isolates provides insights into antigenic differences

Mycobacterium abscessus (MAB) is a widely disseminated pathogenic non-tuberculous mycobacterium (NTM). Like with M. tuberculosis complex (MTBC), excreted / secreted (ES) proteins play an essential role for its virulence and survival inside the host. ES proteins contain highly immunogenic proteins, which are of interest for novel diagnostic assays and vaccines. Here, we used a robust bioinformatics pipeline to predict the secretome of the M. abscessus ATCC 19977 reference strain and fifteen clinical isolates belonging to all three MAB subspecies, M. abscessus subsp. abscessus, M. abscessus subsp. bolletii, and M. abscessus subsp. massiliense. We found that ~18% of the proteins encoded in the MAB genomes were predicted as secreted and that the three MAB subspecies shared > 85 % of the predicted secretomes. MAB isolates with a rough (R) colony morphotype showed larger predicted secretomes than isolates with a smooth (S) morphotype. Additionally, proteins exclusive to the secretomes of MAB R variants had higher antigenic densities than those exclusive to S variants, independently of the subspecies. For all investigated isolates, ES proteins had a significantly higher antigenic density than non-ES proteins. We identified 337 MAB ES proteins with homologues in previously investigated M. tuberculosis secretomes. Among these, 222 have previous experimental support of secretion, and some proteins showed homology with protein drug targets reported in the DrugBank database. The predicted MAB secretomes showed a higher abundance of proteins related to quorum-sensing and Mce domains as compared to MTBC indicating the importance of these pathways for MAB pathogenicity and virulence. Comparison of the predicted secretome of M. abscessus ATCC 19977 with the list of essential genes revealed that 99 secreted proteins corresponded to essential proteins required for in vitro growth. All predicted secretomes were deposited in the Secret-AAR web-server (http://microbiomics.ibt.unam.mx/tools/aar/index.php).

bioinformatics↗

Functional Microswitches of Mammalian G Protein-Coupled Bitter-Taste Receptors

Bitter taste receptors (TAS2Rs) are a poorly understood subgroup of G protein-coupled receptors (GPCRs). The experimental structure of these receptors has yet to be determined, and key-residues controlling their function remain mostly unknown. We designed an integrative approach to improve comparative modeling of TAS2Rs. Using current knowledge on class A GPCRs and existing experimental data in the literature as constraints, we pinpointed conserved motifs to entirely re-align the amino-acid sequences of TAS2Rs. We constructed accurate homology models of human TAS2Rs. As a test case, we examined the accuracy of the TAS2R16 model with site-directed mutagenesis and in vitro functional assays. This combination of in silico and in vitro results clarify sequence-function relationships and identify the functional molecular switches that encode agonist sensing and downstream signaling mechanisms within mammalian TAS2Rs sequences. ClassificationBiological sciences, Computational biology, and bioinformatics

bioinformatics↗

ShinyCell: Simple and sharable visualisation of single-cell gene expression data

MotivationAs the generation of complex single-cell RNA sequencing datasets becomes more commonplace it is the responsibility of researchers to provide access to these data in a way that can be easily explored and shared. Whilst it is often the case that data is deposited for future bioinformatic analysis many studies do not release their data in a way that is easy to explore by non-computational researchers. ResultsIn order to help address this we have developed ShinyCell, an R package that converts single-cell RNA sequencing datasets into explorable and shareable interactive interfaces. These interfaces can be easily customised in order to maximise their usability and can be easily uploaded to online platforms to facilitate wider access to published data. AvailabilityShinyCell is available at https://github.com/SGDDNB/ShinyCell. Contactowen.rackham@duke-nus.edu.sg

bioinformatics↗

Tracking cytosine depletion in SARS-CoV-2

MotivationDanchin et al. have pointed out that cytosine drives the evolution of SARS-CoV-2. A depletion of cytosine might lead to the attenuation of SARS-CoV-2. ResultsWe built a website to track the composition change of mono-, di-, and tri-nucleotide of SARS-CoV-2 over time. The website downloads new strains available from GISAID and updates its results daily. Our analysis suggests that the composition of cytosine in coronaviruses is related to their reported mortality. Using 137,315 SARS-CoV-2 strains collected in ten months, we observed cytosine depletion at a rate of about one cytosine loss per month from the whole genome. AvailabilityThe website is available at http://www.bio8.cs.hku.hk/sarscov2/. Contactrbluo@cs.hku.hk Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

simplifyEnrichment: an R/Bioconductor package for Clustering and Visualizing Functional Enrichment Results

Functional enrichment analysis or gene set enrichment analysis is a basic bioinformatics method that evaluates biological importance of a list of genes of interest. However, it may produce a long list of significant terms with highly redundant information that is difficult to summarize. Current tools to simplify enrichment results by clustering them into groups either still produce redundancy between clusters or do not retain consistent term similarities within clusters. We propose a new method named binary cut for clustering similarity matrices of functional terms. Through comprehensive benchmarks on both simulated and real-world datasets, we demonstrated that binary cut can efficiently cluster functional terms into groups where terms showed more consistent similarities within groups and were more mutually exclusive between groups. We compared binary cut clustering on the similarity matrices obtained from different similarity measures and found that the semantic similarity worked well with binary cut while similarity matrices based on gene overlap showed less consistent patterns. We implemented the binary cut algorithm in the R package simplifyEnrichment which additionally provides functionalities for visualizing, summarizing and comparing the clusterings.

bioinformatics↗

CharPlant: A De Novo Open Chromatin Region (OCR) Prediction Tool for Plant Genomes

Chromatin accessibility is a highly informative structural feature for understanding gene transcription regulation because it indicates the degree to which nuclear macromolecules such as proteins and RNA can access chromosomal DNA. Studies show that chromatin accessibility is highly dynamic during stress response, stimulus response, and developmental transition. Moreover, physical access to chromosomal DNA in eukaryotes is highly cell-specific. Therefore, current technologies such as DNase-seq, ATAC-seq, and FAIRE-seq reveal only a portion of the open chromatin regions (OCRs) present in a given species. Thus, the genome-wide distribution of OCRs remains unknown. In this study, we developed a bioinformatics tool called CharPlant for the de novo prediction of chromatin accessible regions in plant genomes. To develop this tool, we constructed a three-layer convolutional neural network (CNN) and subsequently trained the CNN using DNase-seq and ATAC-seq datasets of four plant species. The model simultaneously learns the sequence motifs and regulatory logics, which are jointly used to determine DNA accessibility. All of these steps are integrated into CharPlant, which can be run using a simple command line. The results of data analysis using CharPlant in this study demonstrate its prediction power and computational efficiency. To our knowledge, CharPlant is the first de novo prediction tool that can identify potential OCRs in the whole genome. The source code of CharPlant and supporting files are freely downloadable from https://github.com/Yin-Shen/CharPlant.

bioinformatics↗

Genome ARTIST_v2 software - a support for annotation of class II natural transposons in new sequenced genomes

Transposon annotation is a very dynamic field of genomics and various tools assigned to support this bioinformatics endeavor were reported. Genome ARTIST (GA) software was initially developed for mapping artificial transposons mobilized during insertional mutagenesis projects. Now, the new functions of GA_v2 qualify it as an effective companion for mapping and annotation of class II natural transposons in assembled genomes, contigs or sequencing reads. Tabular export of mapping and annotation data for subsequent high-throughput data analysis, the export of a list of flanking sequences around either the coordinates of insertion or around the target site duplications (TSDs) and generation of a consensus sequence for the respective flanking sequences are all key assets of GA_v2. Additionally, we developed two accompanying short scripts that enable the user to annotate transposons existent in assembled genomes and to use various annotation offered by FlyBase for Drosophila melanogaster genome. Herein, we present the applicability of GA_v2 for a preliminary annotation of the class II transposon P-element in the genome of D. melanogaster strain Horezu, Romania, which was sequenced with Nanopore technology in our laboratory. Our results point that GA_v2 is a reliable tool to be integrated in pipelines designed to perform transposon annotation in new sequenced genomes. GA_v2 is open source software compatible with Ubuntu, Mac OS and Windows and is available at https://github.com/genomeartist/genomeartist and at www.genomeartist.ro.

bioinformatics↗