bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

BrowseVCF: a web-based application and workflow to quickly prioritise disease-causative variants in VCF files.

As sequencing costs associated with fast advancing Next Generation Sequencing (NGS) technologies continue to decrease, variant discovery is becoming a more affordable and popular analysis method among research laboratories. Following variant calling and annotation, accurate variant filtering is a crucial step to extract meaningful biological information from sequencing data and to investigate disease etiology. However, the standard variant call file format (VCF) used to store this valuable information is not easy to handle without bioinformatics skills, thus preventing many investigators from directly analysing their data. Here, we present BrowseVCF, an easy-to-use stand-alone software that enables researchers to browse, query and filter millions of variants in a few seconds. Key features include the possibility to store intermediate search results, to query user-defined gene lists, to group samples for family or tumour/normal studies, to download a report of the filters applied, and to export the filtered variants in spreadsheet format. Additionally, BrowseVCF is suitable for any DNA variant analysis (exome, whole-genome and targeted sequencing), can be used also for non-diploid genomes, and is able to discriminate between Single Nucleotide Polymorphisms (SNPs), Insertions/Deletions (InDels), and Multiple Nucleotide Polymorphisms (MNPs). Owing to its portable implementation, BrowseVCF can be used either on personal computers or as part of automated analysis pipelines. The software can be initialised with a few clicks on any operating system without any special administrative or installation permissions. It is actively developed and maintained, and freely available for download from https://github.com/BSGOxford/BrowseVCF/releases/latest.

Bioinformatics

High-throughput pipeline for de-novo assembly and drug resistance mutations identification from Next-Generation Sequencing viral data of residual diagnostic samples

MotivationThe underlying genomic variation of a large number of pathogenic viruses can give rise to drug resistant mutations resulting in treatment failure. Next generation sequencing (NGS) enables the identification of viral quasi-species and the quantification of minority variants in clinical samples; therefore, it can be of direct benefit by detecting drug resistant mutations and devising optimal treatment strategies for individual patients.\n\nResultsThe ICONIC (InfeCtion respONse through vIrus genomiCs) project has developed an automated, portable and customisable high-throughput computational pipeline to assemble de novo whole viral genomes, either segmented or non-segmented, and quantify minority variants using residual diagnostic samples. The pipeline has been benchmarked on a dedicated High-Performance Computing cluster using paired-end reads from RSV and Influenza clinical samples. The median length of generated genomes was 96% for the RSV dataset and 100% for each Influenza segment. The analysis of each set lasted less than 12 hours; each sample took around 3 hours and required a maximum memory of 10 GB. The pipeline can be easily ported to a dedicated server or cluster through either an installation script or a docker image. As it enables the subtyping of viral samples and the detection of relevant drug resistance mutations within three days of sample collection, our pipeline could operate within existing clinical reporting time frames and potentially be used as a decision support tool towards more effective personalised patient treatments.\n\nAvailabilityThe software and its documentation are available from https://github.com/ICONIC-UCL/pipeline\n\nContactt.cassarino@ucl.ac.uk, pk5@sanger.ac.uk\n\nSupplementary informationSupplementary data are available at Briefings in Bioinformatics online.

Bioinformatics

Canvas: versatile and scalable detection of copy number variants

Motivation: Increased throughput and diverse experimental designs of large-scale sequencing studies necessi-tate versatile, scalable and robust variant calling tools. In particular, identification of copy number changes re-mains a challenging task due to their complexity, susceptibility to sequencing biases, variation in coverage data and dependence on genome-wide sample properties, such as tumor polyploidy or polyclonality in cancer samples.\n\nResults: We have developed a new tool, Canvas, for identification of copy number changes from diverse se-quencing experiments including whole-genome matched tumor-normal and single-sample normal re-sequencing, as well as whole-exome matched and unmatched tumor-normal studies. In addition to variant calling, Canvas infers genome-wide parameters such as cancer ploidy, purity and heterogeneity. It provides fast and simple to execute workflows that can scale to thousands of samples and can be easily incorporated into existing variant calling pipelines.\n\nAvailability: Canvas is distributed under an open source license and can be downloaded from https://github.com/Illumina/canvas.\n\nContact: eroller@illumina.com\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

Bioinformatics

Robust Lineage Reconstruction from High-Dimensional Single-Cell Data

Single-cell gene expression data provide invaluable resources for systematic characterization of cellular hierarchy in multi-cellular organisms. However, cell lineage reconstruction is still often associated with significant uncertainty due to technological constraints. Such uncertainties have not been taken into account in current methods. We present ECLAIR, a novel computational method for the statistical inference of cell lineage relationships from single-cell gene expression data. ECLAIR uses an ensemble approach to improve the robustness of lineage predictions, and provides a quantitative estimate of the uncertainty of lineage branchings. We show that the application of ECLAIR to published datasets successfully reconstructs known lineage relationships and significantly improves the robustness of predictions. In conclusion, ECLAIR is a powerful bioinformatics tool for single-cell data analysis. It can be used for robust lineage reconstruction with quantitative estimate of prediction accuracy.

Bioinformatics

Seqping: Gene Prediction Pipeline for Plant Genomes using Self-Trained Gene Models and Transcriptomic Data

SummaryAlthough various software are available for gene prediction, none of the currently available gene-finders have a universal Hidden Markov Models (HMM) that can perform gene prediction for all organisms equally well in an automatic fashion. Here, we report an automated pipeline that performs gene prediction using selftrained HMM models and transcriptomic data. The program processes the genome and transcriptome sequences of a target species through GlimmerHMM, SNAP, and AUGUSTUS training pipeline that ends with the program MAKER2 combining the predictions from the three models in association with the transcriptomic evidence. The pipeline generates species-specific HMMs and is able to predict genes that are not biased to other model organisms. Our evaluation of the program revealed that it performed better than the use of the closest related HMM from a standalone program.\n\nAvailability and ImplementationDistributed under the GNU license with free download at http://sourceforge.net/projects/seqping and http://genomsawit.mpob.gov.my.\n\nContactchankl@mpob.gov.my\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

PoPoolationTE2: comparative population genomics of transposable elements using Pool-Seq

The evolutionary dynamics of transposable elements (TEs) are still poorly understood. One reason is that TE abundance needs to be studied at the population level, and despite recent advances in sequencing technologies, characterizing TE abundance in multiple populations by sequencing individuals separately is still too expensive. While sequencing pools of individuals (Pool-Seq) dramatically reduces sequencing costs, a comparison of TE abundance between pooled samples has been difficult, if not impossible, due to various biases. Here, we introduce a novel bioinformatic tool, PoPoolationTE2, which is specifically tailored for the comparison of TE abundance among pooled population samples or different tissues. Using computer simulations we demonstrate that PoPoolationTE2 not only faithfully recovers TE insertion frequencies and positions but, by homogenizing the power to identify TEs acrosss samples, it provides an unbiased comparison of TE abundance between pooled population samples. We anticipate that PoPoolationTE2 will greatly facilitate the analysis of TE insertion patterns in a broad range of applications.

Bioinformatics

annotatr: Associating genomic regions with genomic annotations

MotivationAnalysis of next-generation sequencing data often results in a list of genomic regions. These may include differentially methylated CpGs/regions, transcription factor binding sites, interacting chromatin regions, or GWAS-associated SNPs, among others. A common analysis step is to annotate such genomic regions to genomic annotations (promoters, exons, enhancers, etc.). Existing tools are limited by a lack of annotation sources and flexible options, the time it takes to annotate regions, an artificial one-to-one region-to-annotation mapping, a lack of visualization options to easily summarize data, or some combination thereof.\n\nResultsWe developed the annotatr Bioconductor package to flexibly and quickly summarize and plot annotations of genomic regions. The annotatr package reports all intersections of regions and annotations, giving a better understanding of the genomic context of the regions. A variety of graphics functions are implemented to easily plot numerical or categorical data associated with the regions across the annotations, and across annotation intersections, providing insight into how characteristics of the regions differ across the annotations. We demonstrate that annotatr is up to 27x faster than comparable R packages. Overall, annotatr enables a richer biological interpretation of experiments.\n\nAvailabilityhttp://bioconductor.org/packages/annotatr/\n\nContactrcavalca@umich.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

MetaCycle: an integrated R package to evaluate periodicity in large scale data

SummaryDetecting periodicity in large scale data remains a challenge. Different algorithms offer strengths and weaknesses in statistical power, sensitivity to outliers, ease of use, and sampling requirements. While efforts have been made to identify best of breed algorithms, relatively little research has gone into integrating these methods in a generalizable method. Here we present MetaCycle, an R package that incorporates ARSER, JTK_CYCLE, and Lomb-Scargle to conveniently evaluate periodicity in time-series data.\n\nAvailability and implementationMetaCycle package is available on the CRAN repository (https://cran.r-project.org/web/packages/MetaCycle/index.html) and GitHub (https://github.com/gangwug/MetaCycle).\n\nContacthogenesch@gmail.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Kronos: a workflow assembler for genome analytics and informatics

BackgroundThe field of next generation sequencing informatics has matured to a point where algorithmic advances in sequence alignment and individual feature detection methods have stabilized. Practical and robust implementation of complex analytical workflows (where such tools are structured into best practices for automated analysis of NGS datasets) still requires significant programming investment and expertise.\n\nResultsWe present Kronos, a software platform for automating the development and execution of reproducible, auditable and distributable bioinformatics workflows. Kronos obviates the need for explicit coding of workflows by compiling a text configuration file into executable Python applications. The framework of each workflow includes a run manager to execute the encoded workflows locally (or on a cluster or cloud), parallelize tasks, and log all runtime events. Resulting workflows are highly modular and configurable by construction, facilitating flexible and extensible meta-applications which can be modified easily through configuration file editing. The workflows are fully encoded for ease of distribution and can be instantiated on external systems, promoting and facilitating reproducible research and comparative analyses. We introduce a framework for building Kronos components which function as shareable, modular nodes in Kronos workflows.\n\nConclusionThe Kronos platform provides a standard framework for developers to implement custom tools, reuse existing tools, and contribute to the community at large. Kronos is shipped with both Docker and Amazon AWS machine images. It is free, open source and available through PyPI (Python Package Index) and https://github.com/jtaghiyar/kronos.

Bioinformatics

A little walk from physical to biological complexity: protein folding and stability

As an example of topic where biology and physics meet, we present the issue of protein folding and stability, and the development of thermodynamics-based bioinformatics tools that predict the stability and thermal resistance of proteins and the change of these quantities upon amino acid substitutions. These methods are based on knowledge-driven statistical potentials, derived from experimental protein structures using the inverse Boltzmann law. We also describe an application of these predictors, which contributed to the understanding of the mechanisms of aggregation of a particular protein known to cause a neuronal disease.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=150 SRC=\"FIGDIR/small/043737_ufig1.gif\" ALT=\"Figure 1000\">\nView larger version (98K):\norg.highwire.dtl.DTLVardef@1fb262borg.highwire.dtl.DTLVardef@186eaaforg.highwire.dtl.DTLVardef@99b828org.highwire.dtl.DTLVardef@953a2_HPS_FORMAT_FIGEXP M_FIG C_FIG

Bioinformatics

SnoVault and encodeD: A novel object-based storage system and applications to ENCODE metadata

The Encyclopedia of DNA elements (ENCODE) project is an ongoing collaborative effort[1-6] to create a comprehensive catalog of functional elements initiated shortly after the completion of the Human Genome Project[7][1]. The current database exceeds 6500 experiments across more than 450 cell lines and tissues using a wide array of experimental techniques to study the chromatin structure, regulatory and transcriptional landscape of the H. sapiens and M. musculus genomes. All ENCODE experimental data, metadata, and associated computational analyses are submitted to the ENCODE Data Coordination Center (DCC) for validation, tracking, storage, unified processing, and distribution to community resources and the scientific community. As the volume of data increases, the identification and organization of experimental details becomes increasingly intricate and demands careful curation. The ENCODE DCC[8-10] has created a general purpose software system, known as SnoVault, that supports metadata and file submission, a database used for metadata storage, web pages for displaying the metadata and a robust API for querying the metadata. The software is fully open-source, code and installation instructions can be found at: http://github.com/ENCODE-DCC/snovault/ (for the generic database) and http://github.com/ENCODE-DCC/encoded/ to store genomic data in the manner of ENCODE. The core database engine, SnoVault (which is completely independent of ENCODE, genomic data, or bioinformatic data) has been released as a separate Python package.\n\nDatabase URL: https://www.encodeproject.org/

Bioinformatics

FOCUS2: agile and sensitive classification of metagenomics data using a reduced database

SummaryMetagenomics approaches rely on identifying the presence of organisms in the microbial community from a set of unknown DNA sequences. Sequence classification has valuable applications in multiple important areas of medical and environmental research. Here we introduce FOCUS2, an update of the previously published computational method FOCUS. FOCUS2 was tested with 10 simulated and 543 real metagenomes demonstrating that the program is more sensitive, faster, and more computationally efficient than existing methods.\n\nAvailabilityThe Python implementation is freely available at https://edwards.sdsu.edu/FOCUS2.\n\nSupplementary informationavailable at Bioinformatics online.

Bioinformatics

Genome-wide generalized additive models

MotivationChromatin immunoprecipitation followed by deep sequencing (ChIP-Seq) is a widely used approach to study protein-DNA interactions. Often, the quantities of interest are the differential occupancies relative to controls, between genetic backgrounds, treatments, or combinations thereof. Current methods for differential occupancy of ChIP-seq data rely however on binning or sliding window techniques, for which the choice of the window and bin sizes are subjective.\n\nResultsHere, we present GenoGAM (Genome-wide Generalized Additive Model), which brings the well-established and flexible generalized additive models framework to genomic applications using a data parallelism strategy. We model ChIP-Seq read count frequencies as products of smooth functions along chromosomes. Smoothing parameters are objectively estimated from the data by cross-validation, eliminating ad-hoc binning and windowing needed by current approaches. GenoGAM provides base-level and region-level significance testing for full factorial designs. Application to a ChIP-Seq dataset in yeast showed increased sensitivity over existing differential occupancy methods while controlling for type I error rate. By analyzing a set of DNA methylation data and illustrating an extension to a peak caller, we further demonstrate the potential of GenoGAM as a generic statistical modeling tool for genome-wide assays.\n\nAvailabilitySoftware is available from Bioconductor: https://www.bioconductor.org/packages/release/bioc/html/GenoGAM.html\n\nContactgagneur@in.tum.de\n\nSupplementary informationSupplementary information is available at Bioinformatics online.

Bioinformatics

Metavisitor, a suite of Galaxy tools for simple and rapid detection and discovery of viruses in deep sequence data

We present user-friendly and adaptable software to provide biologists, clinical researchers and possibly diagnostic clinicians with the ability to robustly detect and reconstruct viral genomes from complex deep sequence datasets. A set of modular bioinformatic tools and workflows was implemented as the Metavisitor package in the Galaxy framework. Using the graphical Galaxy workflow editor, users with minimal computational skills can use existing Metavisitor workflows or adapt them to suit specific needs by adding or modifying analysis modules. Metavisitor can be used on our Mississippi server, or can be installed on any Galaxy server instance and a pre-configured Metavisitor server image is provided. Metavisitor works with DNA, RNA or small RNA sequencing data over a range of read lengths and can use a combination of de novo and guided approaches to assemble genomes from sequencing reads. We show that the software has the potential for quick diagnosis as well as discovery of viruses from a vast array of organisms. Importantly, we provide here executable Metavisitor use cases, which increase the accessibility and transparency of the software, ultimately enabling biologists or clinicians to focus on biological or medical questions.

Bioinformatics

Heat*seq: an interactive web tool for high-throughput sequencing experiment comparison with public data

Better protocols and decreasing costs have made high-throughput sequencing experiments now accessible even to small experimental laboratories. However, comparing one or few experiments generated by an individual lab to the vast amount of relevant data freely available in the public domain might be limited due to lack of bioinformatics expertise. Though several tools, including genome browsers, allow such comparison at a single gene level, they do not provide a genome-wide view. We developed Heat*seq, a web-tool that allows genome scale comparison of high throughput experiments (ChIP-seq, RNA-seq and CAGE) provided by a user, to the data in the public domain. Heat*seq currently contains over 12,000 experiments across diverse tissue and cell types in human, mouse and drosophila. Heat*seq displays interactive correlation heatmaps, with an ability to dynamically subset datasets to contextualise user experiments. High quality figures and tables are produced and can be downloaded in multiple formats.\n\nAvailabilityWeb application: www.heatstarseq.roslin.ed.ac.uk/. Source code: https://github.com/gdevailly.\n\nContactGuillaume.Devailly@roslin.ed.ac.uk; Anagha.Joshi@roslin.ed.ac.uk

Bioinformatics

Eliminating redundancy among protein sequences using submodular optimization

AbstactO_ST_ABSMotivationC_ST_ABSSubmodular optimization, a discrete analogue to continuous convex optimization, has been used with great success in many fields but is not yet widely used in biology. We apply submodular optimization to the problem of removing redundancy in protein sequence data sets. This is a common step in many bioinformatics and structural biology workflows, including creation of non-redundant training sets for sequence and structural models as well as selection of \"operational taxonomic units\" from metagenomics data.\n\nResultsWe demonstrate that the submodular optimization approach results in representative protein sequence subsets with greater structural diversity than sets chosen by existing methods. In particular, we compare to a widely used, heuristic algorithm implemented in software tools such as CD-HIT, as well to as a variety of standard clustering methods, using as a gold standard the SCOPe library of protein domain structures. In this setting, submodular optimization consistently yields protein sequence subsets that include more SCOPe domain families than sets of the same size selected by competing approaches. We also show how the optimization framework allows us to design a mixture objective function that performs well for both large and small representative sets. The framework we describe is theoretically optimal under some assumptions, and it is flexible and intuitive because it applies generic methods to optimize one of a variety of objective functions. This application serves as a model for how submodular optimization can be applied to other discrete problems in biology.\n\nAvailabilitySource code is available at https://github.com/mlibbrecht/submodular_sequence_repset.\n\nContactwilliam-noble@uw.edu

Bioinformatics

UMI-tools: Modelling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy

Unique Molecular Identifiers (UMIs) are random oligonucleotide barcodes that are increasingly used in high-throughout sequencing experiments. Through a UMI, identical copies arising from distinct molecules can be distinguished from those arising through PCR amplification of the same molecule. However, bioinformatic methods to leverage the information from UMIs have yet to be formalised. In particular, sequencing errors in the UMI sequence are often ignored, or else resolved in an ad-hoc manner. We show that errors in the UMI sequence are common and introduce network-based methods to account for these errors when identifying PCR duplicates. Using these methods, we demonstrate improved quantification accuracy both under simulated conditions and real iCLIP and single cell RNA-Seq datasets. Reproducibility between iCLIP replicates and single cell RNA-Seq clustering are both improved using our proposed network-based method, demonstrating the value of properly accounting for errors in UMIs. These methods are implemented in the open source UMI-tools software package (https://github.com/CGATOxford/UMI-tools).

Bioinformatics

GeMSTONE: Orchestrated Prioritization of Human Germline Mutations in the Cloud

Integrative analysis of whole-genome/exome-sequencing data has been challenging, especially for the non-programming research community, as it requires leveraging an inordinate number of computational tools. Even computational biologists find it unexpectedly difficult to reproduce results from others or optimize their own strategies in an end-to-end workflow. We introduce Germline Mutation Scoring Tool fOr Next-generation sEquencing data (GeMSTONE), a cloud-based variant prioritization tool with high-level customization and a comprehensive collection of bioinformatics tools and data libraries (http://gemstone.yulab.org/). GeMSTONE generates and readily accepts a sharable \"recipe\" file for each run to either replicate existing results or analyze new data with identical parameters.

Bioinformatics