bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

InterARTIC: an interactive web application for whole-genome nanopore sequencing analysis of SARS-CoV-2 and other viruses

MotivationInterARTIC is an interactive web application for the analysis of viral whole-genome sequencing (WGS) data generated on Oxford Nanopore Technologies (ONT) devices. A graphical interface enables users with no bioinformatics expertise to analyse WGS experiments and reconstruct consensus genome sequences from individual isolates of viruses, such as SARS-CoV-2. InterARTIC is intended to facilitate widespread adoption and standardisation of ONT sequencing for viral surveillance and molecular epidemiology. Worked exampleWe demonstrate the use of InterARTIC for the analysis of ONT viral WGS data from SARS-CoV-2 and Ebola virus, using a laptop computer or the internal computer on an ONT GridION sequencing device. We showcase the intuitive graphical interface, workflow customisation capabilities and job-scheduling system that facilitate execution of small- and large-scale WGS projects on any common virus. ImplementationInterARTIC is a free, open-source web application implemented in Python. The application can be downloaded as a set of pre-compiled binaries that are compatible with all common Ubuntu distributions, or built from source. For further details please visit: https://github.com/Psy-Fer/interARTIC/.

bioinformatics↗

Assembly-free rapid differential gene expression analysis in non-model organisms using DNA-protein alignment

BackgroundRNA-seq is being increasingly adopted for gene expression studies in a panoply of non-model organisms, with applications spanning the fields of agriculture, aquaculture, ecology, and environment. For organisms that lack a well-annotated reference genome or transcriptome, a conventional RNA-seq data analysis workflow requires constructing a de-novo transcriptome assembly and annotating it against a high-confidence protein database. The assembly serves as a reference for read mapping, and the annotation is necessary for functional analysis of genes found to be differentially expressed. However, assembly is computationally expensive. It is also prone to errors that impact expression analysis, especially since sequencing depth is typically much lower for expression studies than for transcript discovery. ResultsWe propose a shortcut, in which we obtain counts for differential expression analysis by directly aligning RNA-seq reads to the high-confidence proteome that would have been otherwise used for annotation. By avoiding assembly, we drastically cut down computational costs - the running time on a typical dataset improves from the order of tens of hours to under half an hour, and the memory requirement is reduced from the order of tens of Gbytes to tens of Mbytes. We show through experiments on simulated and real data that our pipeline not only reduces computational costs, but has higher sensitivity and precision than a typical assembly-based pipeline. A Snakemake implementation of our workflow is available at: https://bitbucket.org/project_samar/samar ConclusionsThe flip side of RNA-seq becoming accessible to even modestly resourced labs has been that the time, labor, and infrastructure cost of bioinformatics analysis has become a bottleneck. Assembly is one such resource-hungry process, and we show here that it can be avoided for quick and easy, yet more sensitive and precise, differential gene expression analysis in non-model organisms.

bioinformatics↗

Predicted structural mimicry of spike receptor-binding motifs from highly pathogenic human coronaviruses

Viruses often encode proteins that mimic host proteins in order to facilitate infection. Little work has been done to understand the potential mimicry of the SARS-CoV-2, SARS-CoV, and MERS-CoV spike proteins, particularly the receptor-binding motifs, which could be important in determining tropism of the virus. Here, we use structural bioinformatics software to characterize potential mimicry of the three coronavirus spike protein receptor-binding motifs. We utilize sequence-independent alignment tools to compare structurally known or predicted three-dimensional protein models with the receptor-binding motifs and verify potential mimicry with protein docking simulations. Both human and non-human proteins were found to be similar to all three receptor-binding motifs. Similarity to human proteins may reveal which pathways the spike protein is co-opting, while analogous non-human proteins may indicate shared host interaction partners and overlapping antibody cross-reactivity. These findings can help guide experimental efforts to further understand potential interactions between human and coronavirus proteins. HighlightsO_LIPotential coronavirus spike protein mimicry revealed by structural comparison C_LIO_LIHuman and non-human protein potential interactions with virus identified C_LIO_LIPredicted structural mimicry corroborated by protein-protein docking C_LIO_LIEpitope-based alignments may help guide vaccine efforts C_LI Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/441187v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@1f09454org.highwire.dtl.DTLVardef@19a5557org.highwire.dtl.DTLVardef@158d3fdorg.highwire.dtl.DTLVardef@c59511_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Towards community-driven metadata standards for light microscopy: tiered specifications extending the OME model

1 -Digital light microscopy provides powerful tools for quantitatively probing the real-time dynamics of subcellular structures. While the power of modern microscopy techniques is undeniable, rigorous record-keeping and quality control are required to ensure that imaging data may be properly interpreted (quality), reproduced (reproducibility), and used to extract reliable information and scientific knowledge which can be shared for further analysis (value). Keeping notes on microscopy experiments and quality control procedures ought to be straightforward, as the microscope is a machine whose components are defined and the performance measurable. Nevertheless, to this date, no universally adopted community-driven specifications exist that delineate the required information about the microscope hardware and acquisition settings (i.e., microscopy "data provenance" metadata) and the minimally accepted calibration metrics (i.e., microscopy quality control metadata) that should be automatically recorded by both commercial microscope manufacturers and customized microscope developers. In the absence of agreed guidelines, it is inherently difficult for scientists to create comprehensive records of imaging experiments and ensure the quality of resulting image data or for manufacturers to incorporate standardized reporting and performance metrics. To add to the confusion, microscopy experiments vary greatly in aim and complexity, ranging from purely descriptive work to complex, quantitative and even sub-resolution studies that require more detailed reporting and quality control measures. To solve this problem, the 4D Nucleome Initiative (4DN) (1, 2) Imaging Standards Working Group (IWG), working in conjunction with the BioImaging North America (BINA) Quality Control and Data Management Working Group (QC-DM-WG) (3), here propose light Microscopy Metadata specifications that scale with experimental intent and with the complexity of the instrumentation and analytical requirements. They consist of a revision of the Core of the Open Microscopy Environment (OME) Data Model, which forms the basis for the widely adopted Bio-Formats library (4-6), accompanied by a suite of three extensions, each with three tiers, allowing the classification of imaging experiments into levels of increasing imaging and analytical complexity (7, 8). Hence these specifications not only provide an OME-based comprehensive set of metadata elements that should be recorded, but they also specify which subset of the full list should be recorded for a given experimental tier. In order to evaluate the extent of community interest, an extensive outreach effort was conducted to present the proposed metadata specifications to members of several core-facilities and international bioimaging initiatives including the European Light Microscopy Initiative (ELMI), Global BioImaging (GBI), and European Molecular Biology Laboratory (EMBL) - European Bioinformatics Institute (EBI). Consequently, close ties were established between our endeavour and the undertakings of the recently established QUAlity Assessment and REProducibility for Instruments and Images in Light Microscopy global community initiative (9). As a result this flexible 4DN-BINA-OME (NBO namespace) framework (7, 8) represents a turning point towards achieving community-driven Microscopy Metadata standards that will increase data fidelity, improve repeatability and reproducibility, ease future analysis and facilitate the verifiable comparison of different datasets, experimental setups, and assays, and it demonstrates the method for future extensions. Such universally accepted microscopy standards would serve a similar purpose as the Encode guidelines successfully adopted by the genomic community (10, 11). The intention of this proposal is therefore to encourage participation, critiques and contributions from the entire imaging community and all stakeholders, including research and imaging scientists, facility personnel, instrument manufacturers, software developers, standards organizations, scientific publishers, and funders.

bioinformatics↗

Toward comprehensive functional analysis of gene lists weighted by gene essentiality scores

Gene functional enrichment analysis represents one of the most popular bioinformatics methods for annotating the pathways and function categories of a given gene list. Current algorithms for enrichment computation such as Fishers exact test and hypergeometric test totally depend on the category count numbers of the gene list and one gene set. In this case, whatever the genes are, they were treated equally. However, actually genes show different scores in their essentiality in a gene list and in a gene set. It is thus hypothesized that the essentiality scores could be important and should be considered in gene functional analysis. For this purpose, here we proposed WEAT (https://www.cuilab.cn/weat/), a weighted gene set enrichment algorithm and online tool by weighting genes using essentiality scores. We confirmed the usefulness of WEAT using two case studies, the functional analysis of one aging-related gene list and one gene list involved in Lung Squamous Cell Carcinoma (LUSC). Finally, we believe that the WEAT method and tool could provide more possibilities for further exploring the functions of given gene lists.

bioinformatics↗

ZGA: a flexible pipeline for read processing, de novo assembly and annotation of prokaryotic genomes

MotivationWhole genome sequencing (WGS) became a routine method in modern days and may be applied to study a wide spectrum of scientific problems. Despite increasing availability of genome sequencing by itself, genome assembly and annotation could be a challenge for an inexperienced researcher. ResultsZGA is a computational pipeline to assemble and annotate prokaryotic genomes. The pipeline supports several modern sequencing platforms and may be used for hybrid genome assembling. Resulting genome assembly is ready for deposition to an INSDC database or for further analysis. AvailabilityZGA was written in Python, the source code is freely available at https://github.com/laxeye/zga/. ZGA can be installed via Anaconda Cloud and Python Package Index. Contactoscypek@ya.ru Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

DeepRegFinder: Deep Learning-Based Regulatory Elements Finder

MotivationEnhancers and promoters are important classes of DNA regulatory elements that control gene expression. Identifying them at the genomic scale is a critical and challenging task in bioinformatics. The most successful method so far is to train machine learning models on known enhancer and promoter sites and predict them at other genomic regions using ChIP-seq and related data. ResultsWe have developed a highly customizable program called DeepRegFinder which automates data processing, model training and genome-wide prediction of enhancers and promoters using convolutional and recurrent neural networks. Our program further classifies the enhancers and promoters into active and poised states to facilitate downstream analysis. Based on mean average precision scores of different classes across multiple cell types, our method significantly outperforms the existing algorithms. Availabilityhttps://github.com/shenlab-sinai/DeepRegFinder

bioinformatics↗

Regulation of Lysosome-Associated Membrane Protein 3 (LAMP3) in Lung Epithelial Cells by Coronaviruses (SARS-CoV-1/2) and Type I Interferon Signaling

Severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) infection is a major risk factor for mortality and morbidity in critical care hospitals around the world. Lung epithelial type II cells play a major role in several physiological processes, including recognition and clearance of respiratory viruses as well as repair of lung injury in response to environmental toxicants. Gene expression profiling of lung epithelial type II-specific genes led to the identification of lysosomal-associated membrane protein 3 (LAMP3). Intracellular locations of LAMP3 include plasma membrane, endosomes, and lysosomes. These intracellular organelles are involved in vesicular transport and facilitate viral entry and release of the viral RNA into the host cell cytoplasm. In this study, regulation of LAMP3 expression in human lung epithelial cells by several respiratory viruses and type I interferon signaling was investigated. Coronaviruses including SARS-CoV-1 and SARS-CoV-2 significantly induced LAMP3 expression in lung epithelial cells within 24 hours after infection that required the presence of ACE2 viral entry receptor. Time-course experiments revealed that the induced expression of LAMP3 by SARS-CoV-2 was correlated with the induced expression of interferon-beta1 (IFNB1) and signal transducers and activator of transcription 1 (STAT1) mRNA levels. LAMP3 was also induced by direct IFN-beta treatment or by infection with influenza virus lacking the non-structural protein1(NS1) in NHBE bronchial epithelial cells. LAMP3 expression was induced in human lung epithelial cells by several respiratory viruses, including respiratory syncytial virus (RSV) and the human parainfluenza virus 3 (HPIV3). Location in lysosomes and endosomes as well as induction by respiratory viruses and type I Interferon suggests that LAMP3 may have an important role in inter-organellar regulation of innate immunity and a potential target for therapeutic modulation in health and disease. Furthermore, bioinformatics revealed that a subset of lung type II cell genes were differentially regulated in the lungs of COVID-19 patients.

bioinformatics↗

spatialLIBD: an R/Bioconductor package to visualize spatially-resolved transcriptomics data

MotivationSpatially-resolved transcriptomics has now enabled the quantification of high-throughput and transcriptome-wide gene expression in intact tissue while also retaining the spatial coordinates. Incorporating the precise spatial mapping of gene activity advances our understanding of intact tissuespecific biological processes. In order to interpret these novel spatial data types, interactive visualization tools are necessary. ResultsWe describe spatialLIBD, an R/Bioconductor package to interactively explore spatially-resolved transcriptomics data generated with the 10x Genomics Visium platform. The package contains functions to interactively access, visualize, and inspect the observed spatial gene expression data and data-driven clusters identified with supervised or unsupervised analyses, either on the users computer or through a web application. AvailabilityspatialLIBD is available at bioconductor.org/packages/spatialLIBD. Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

HyMM: Hybrid method for disease-gene prediction by integrating multiscale module structures

MotivationIdentifying disease-related genes is important for the study of human complex diseases. Module structures or community structures are ubiquitous in biological networks. Although the modular nature of human diseases can provide useful insights, the mining of information hidden in multiscale module structures has received less attention in disease-gene prediction. ResultsWe propose a hybrid method, HyMM, to predict disease-related genes more effectively by integrating the information from multiscale module structures. HyMM consists of three key steps: extraction of multiscale modules, gene rankings based on multiscale modules and integration of multiple gene rankings. The statistical analysis of multiscale modules extracted by three multiscale-module-decomposition algorithms (MO, AS and HC) shows that the functional consistency of the modules gradually improves as the resolution increases. This suggests the existence of different levels of functional relationships in the multiscale modules, which may help reveal disease-gene associations. We display the effectiveness of multiscale module information in the disease-gene prediction and confirm the excellent performance of HyMM by 5-fold cross-validation and independent test. Specifically, HyMM with MO can more effectively enhance the ability of disease-gene prediction; HyMM (MO, RWR) and HyMM (MO, RWRH) are especially preferred due to their excellent comprehensive performance, and HyMM (AS, RWRH) is also good choice due to its local performance. We anticipate that this work could provide useful insights for disease-module analysis and disease-gene prediction based on multi-scale module structures. Availabilityhttps://github.com/xiangiu0208/HvMM Contactlimin@mail.csu.edu.cn Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Simplifying the development of portable, scalable, and reproducible workflows

Command-line software plays a critical role in biology research. However, processes for installing and executing software differ widely. The Common Workflow Language (CWL) is a community standard that addresses this problem. Using CWL, tool developers can formally describe a tools inputs, outputs, and other execution details in a manner that fosters use of shared computational methods and reproducibility of complex analyses. CWL documents can include instructions for executing tools inside software containers--isolated, operating-system environments. Accordingly, CWL tools are portable--they can be executed on diverse computers--including personal workstations, high-performance clusters, or the cloud. This portability enables easier adoption of bioinformatics pipelines. CWL supports workflows, which describe dependencies among tools and using outputs from one tool as inputs to others. To date, CWL has been used primarily for batch processing of large datasets, especially in genomics. But it can also be used for analytical steps of a study. This article explains key concepts about CWL and software containers and provides examples for using CWL in biology research. CWL documents are text-based, so they can be created manually, without computer programming. However, ensuring that these documents confirm to the CWL specification may prevent some users from adopting it. To address this gap, we created ToolJig, a Web application that enables researchers to create CWL documents interactively. ToolJig validates information provided by the user to ensure it is complete and valid. After creating a CWL tool or workflow, the user can create "input-object" files, which store values for a particular invocation of a tool or workflow. In addition, ToolJig provides examples of how to execute the tool or workflow via a workflow engine.

bioinformatics↗

Retention Time Standardization and Registration (RTStaR): An algorithm that matches corresponding and identifies unique species in nanoliquid chromatography-nanoelectrospray ionization-mass spectrometry lipidomic datasets

Bioinformatic tools capable of registering, rapidly and reproducibly, large numbers of nanoliquid chromatography-nanoelectrospray ionization-tandem mass spectrometry (nLC-nESI-MS/MS) lipidomic datasets are lacking. We provide here a freely available Retention Time Standardization and Registration (RTStaR) algorithm that aligns nLC-nESI-MS/MS spectra within a single dataset and compares these aligned retention times across multiple datasets. This two-step calibration matches corresponding and identifies unique lipid species in different lipidomes from different matrices and organisms. RTStaR was developed using a population-based study of 1001 human serum samples composed of 71 distinct glycerophosphocholine metabolites comprising a total of 68,572 analytes. Platform and matrix independence were validated using different MS instruments, nLC methodologies, and mammalian lipidomes. The complete algorithm is packaged in two modular ExcelTM workbook templates for easy implementation. RTStaR is freely available from the India Taylor Lipidomics Research Platform http://www.neurolipidomics.ca/rtstar/rtstar.html. Technical support is provided through ldomic@uottawa.ca

bioinformatics↗

Unbiased lexicometry analyses illuminate plague dynamics during the second pandemic

Knowledge of the second plague pandemic that swept over Europe during the 14th-19th centuries, mainly relies on the exegesis of contemporary texts, prone to interpretive bias. Leveraging bioinformatic tools routinely used in biology, we here developed a quantitative lexicography of 32 texts describing two major plague outbreaks using contemporary plague-unrelated texts as negative controls. Nested, network and category analyses of a 207-word pan-lexicome over-represented in plague-related texts, indicated that "buboes" and "carbuncles" words were significantly associated with plague, signaling ectoparasite- borne plague. Moreover, plague-related words were associated with the words "merchandise", "movable", "tatters", "bed" and "clothes", while no association was found with rats and fleas. These results support the hypothesis that during the second plague pandemic, human ectoparasites were the major drivers of plague. Analyzing ancient texts using here reported method would certify plague-related historical records and indicate particularities of each plague outbreak, including sources of the causative Yersinia pestis.

bioinformatics↗

poreCov - an easy to use, fast, and robust workflow for SARS-CoV-2 genome reconstruction via nanopore sequencing

In response to the SARS-CoV-2 pandemic, a highly increased sequencing effort has been established worldwide to track and trace ongoing viral evolution. Technologies such as nanopore sequencing via the ARTIC protocol are used to reliably generate genomes from raw sequencing data as a crucial base for molecular surveillance. However, for many labs that perform SARS-CoV-2 sequencing, bioinformatics is still a major bottleneck, especially if hundreds of samples need to be processed in a recurring fashion. Pipelines developed for short-read data cannot be applied to nanopore data. Therefore, specific long-read tools and parameter settings need to be orchestrated to enable accurate genotyping and robust reference-based genome reconstruction of SARS-CoV-2 genomes from nanopore data. Here we present poreCov, a highly parallel workflow written in Nextflow, using containers to wrap all the tools necessary for a routine SARS-CoV-2 sequencing lab into one program. The ease of installation, combined with concise summary reports that clearly highlight all relevant information, enables rapid and reliable analysis of hundreds of SARS-CoV-2 raw sequence data sets or genomes. poreCov is freely available on GitHub under the GNUv3 license: github.com/replikation/poreCov.

bioinformatics↗

OPUS-X: An Open-Source Toolkit for Protein Torsion Angles, Secondary Structure, Solvent Accessibility, Contact Map Predictions, and 3D Folding

In this paper, we report an open-source toolkit for protein 3D structure modeling, named OPUS-X. It contains three modules: OPUS-TASS2, which predicts protein torsion angles, secondary structure and solvent accessibility; OPUS-Contact, which measures the distance and orientations information between different residue pairs; and OPUS-Fold2, which uses the constraints derived from the first two modules to guide folding. OPUS-TASS2 is an upgraded version of our previous method OPUSS-TASS (Bioinformatics 2020, 36 (20), 5021-5026). OPUS-TASS2 integrates protein global structure information and significantly outperforms OPUS-TASS. OPUS-Contact combines multiple raw co-evolutionary features with protein 1D features predicted by OPUS-TASS2, and delivers better results than the open-source state-of-the-art method trRosetta. OPUS-Fold2 is a complementary version of our previous method OPUS-Fold (J. Chem. Theory Comput. 2020, 16 (6), 3970-3976). OPUS-Fold2 is a gradient-based protein folding framework based on the differentiable energy terms in opposed to OPUS-Fold that is a sampling-based method used to deal with the non-differentiable terms. OPUS-Fold2 exhibits comparable performance to the Rosetta folding protocol in trRosetta when using identical inputs. OPUS-Fold2 is written in Python and TensorFlow2.4, which is user-friendly to any source-code level modification. The code and pre-trained models of OPUS-X can be downloaded from https://github.com/OPUS-MaLab/opus_x.

bioinformatics↗

Prowler: A novel trimming algorithm for Oxford Nanopore sequence data

MotivationQuality control (QC) tools are critical in DNA sequencing analysis because they increase the accuracy of sequence alignments and thus the reliability of results. Oxford Nanopore Technologies (ONT) QC is currently rudimentary, generally based on whole read average quality. This results in discarding reads that contain regions of high quality sequence. Here we propose Prowler, a multi-window approach inspired by algorithms used to QC short read data. Importantly, we retain the phase and read length information by optionally replacing trimmed sections with Ns. ResultsProwler was applied to mammalian and bacterial datasets, to assess effects on alignment and assembly respectively. Compared to Nanofilt, alignments of data QCed with Prowler had lower error rates and more mapped reads. Assemblies of Prowler QCed data had a lower error rate than Nanofilt QCed data however this came at some cost to assembly contiguity. Availability and implementationProwler is implemented in Python and is available at: https://github.com/ProwlerForNanopore/ProwlerTrimmer Contacte.ross@uq.edu.au Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

riboviz 2: A flexible and robust ribosome profiling data analysis and visualization workflow

MotivationRibosome profiling, or Ribo-seq, is the state of the art method for quantifying protein synthesis in living cells. Computational analysis of Ribo-seq data remains challenging due to the complexity of the procedure, as well as variations introduced for specific organisms or specialized analyses. Many bioinformatic pipelines have been developed, but these pipelines have key limitations in terms of functionality or usability. ResultsWe present riboviz 2, an updated riboviz package, for the comprehensive transcript-centric analysis and visualization of Ribo-seq data. riboviz 2 includes an analysis workflow built on the Nextflow workflow management system, combining freely available software with custom code. The package is extensively documented and provides example configuration files for organisms spanning the domains of life. riboviz 2 is distinguished by clear separation of concerns between annotation and analysis: prior to a run, the user chooses a transcriptome in FASTA format, paired with annotation for the CDS locations in GFF3 format. The user is empowered to choose the relevant transcriptome for their biological question, or to run alternative analyses that address distinct questions. riboviz 2 has been extensively tested on various library preparation strategies, including multiplexed samples. riboviz 2 is flexible and uses open, documented file formats, allowing users to integrate new analyses with the pipeline. Availabilityriboviz 2 is freely available at github.com/riboviz/riboviz. Supplementary information

bioinformatics↗

DENIES: A deep learning based two-layer predictor for enhancing the identification of enhancers and their strength with DNA shape information

The identification of enhancers has always been an important task in bioinformatics owing to their major role in regulating gene expression. For this reason, many computational algorithms devoted to enhancer identification have been put forward over the years. To boost the performance of their methods, more features are extracted from the single DNA sequences and integrated to develop an ensemble classifier. Nevertheless, the sequence-derived features used in previous studies can hardly provide the 3D structure information of DNA sequences, which is regarded as an important factor affecting the binding preferences of transcription factors to regulatory elements like enhancers. Given that, we here propose SENIES, a DNA shape enhanced deep learning predictor, for the identification of enhancers and their strength. The predictor consists of two layers where the first layer is for enhancer and non-enhancer identification, and the second layer is for predicting the strength of enhancers. Besides utilizing two common sequence-derived features (i.e. one-hot and k-mer) as input, it introduces DNA shape for describing the 3D structures of DNA sequences. Performance comparison with state-of-the-art methods conducted on the same datasets demonstrates the effectiveness and robustness of our method. The code implementation of our predictor is publicly available at https://github.com/hlju-liye/SENIES.

bioinformatics↗