bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,801 records · Page 100Linked to original sources

Ultra-rapid somatic variant detection via real-time threshold sequencing

Molecular markers are becoming increasingly important for cancer diagnosis, proper clinical trial enrollment, and even surgical decision making, motivating ultra-rapid, intraoperative variant detection. Sequencing-based detection is considered the gold standard approach, but typically takes hours to perform. In this work, we present Threshold Sequencing, a methodology for designing protocols for targeted variant detection on real-time sequencers with a minimal time to result. Threshold Sequencing analytically identifies a time-optimal threshold to stop target amplification and begin sequencing. To further reduce diagnostic time, we explore targeted Loop-mediated Isothermal Amplification (LAMP) and design a LAMP-specific bioinformatics tool--LAMPrey--to process sequenced LAMP product. LAMPreys concatemer aware alignment algorithm is designed to maximize recovery of diagnostically relevant information leading to a more rapid detection versus standard read alignment approaches. Coupled with time-optimized DNA extraction and library preparation, we demonstrate confirmation of a hot-spot mutation (250x support) from tumor tissue in less than 30 minutes.

bioinformatics↗

Single nucleotide polymorphism induces divergent dynamic patterns in CYP3A5: a microsecond scale biomolecular simulation of variants identified in Sub-Saharan African populations

Pharmacogenomics aims to reveal variants associated with drug response phenotypes. Genes whose roles involve the absorption, distribution, metabolism, and excretion of drugs, are highly polymorphic between populations. High coverage whole genome sequencing showed that a large proportion of the variants for these genes are rare in African populations. This study investigates the impact of such variants on protein structure to assess their functional importance. We use genetic data of CYP3A5 from 458 individuals from sub-Saharan Africa to conduct a structural bioinformatics analysis. Five missense variants were modeled and microsecond scale molecular dynamics simulations were conducted for each, as well as for the CYP3A5 wildtype, and the Y53C variant, which has a known deleterious impact on enzyme activity. The binding of ritonavir and artemether to CYP3A5 variant structures was also evaluated. Our results showed different conformational characteristics between all the variants. No significant structural changes were noticed. However, the genetic variability acts on the plasticity of the protein. The impact on drug binding may be drug dependant. We conclude that rare variants hold relevance in determining the pharmacogenomics properties of populations. This could have a significant impact on precision medicine applications in sub-Saharan Africa.

bioinformatics↗

KmerKeys: a web resource for searching indexed genome assemblies and variants

K-mers are short DNA sequences that are used for genome sequence analysis. Applications that use k-mers include genome assembly and alignment. Despite these current applications, the wider bioinformatic use of k-mers in has challenges related to the massive scale of genomic sequence data. A single human genome assembly has billions of these short sequences. The sheer amount of computation for effective use of k-mer information is enormous, particularly when involving multiple genome assemblies. To address these issues, we developed a new k-mer indexing data structure based on a hash table tuned for the lookup of k-mer keys. This web application, referred to as KmerKeys (https://kmerkeys.dgi-stanford.org/), provides performant, rapid query speeds for cloud computation on genome assemblies. We enable fuzzy as well as exact k-mer-based searches of assemblies. To enable robust and speedy performance, the website implements cache-friendly hash tables, memory mapping and massive parallel processing. Our method employs a scalable and efficient data structure that can be used to jointly index and search a large collection of human genome assembly information. One can include variant databases and their associated metadata such as the gnomAD population variant catalog. This feature enables the incorporation of future genomic information into sequencing analysis.

bioinformatics↗

Fast processing of environmental DNA metabarcoding sequence data using convolutional neural networks

1The intensification of anthropogenic pressures have increased consequences on biodiversity and ultimately on the functioning of ecosystems. To monitor and better understand biodiversity responses to environmental changes using standardized and reproducible methods, novel high-throughput DNA sequencing is becoming a major tool. Indeed, organisms shed DNA traces in their environment and this "environmental DNA" (eDNA) can be collected and sequenced using eDNA metabarcoding. The processing of large volumes of eDNA metabarcoding data remains challenging, especially its transformation to relevant taxonomic lists that can be interpreted by experts. Speed and accuracy are two major bottlenecks in this critical step. Here, we investigate whether convolutional neural networks (CNN) can optimize the processing of short eDNA sequences. We tested whether the speed and accuracy of a CNN are comparable to that of the frequently used OBITools bioinformatic pipeline. We applied the methodology on a massive eDNA dataset collected in Tropical South America (French Guiana), where freshwater fishes were targeted using a small region (60pb) of the 12S ribosomal RNA mitochondrial gene. We found that the taxonomic assignments from the CNN were comparable to those of OBITools, with high correlation levels and a similar match to the regional fish fauna. The CNN allowed the processing of raw fastq files at a rate of approximately 1 million sequences per minute which was 150 times faster than with OBITools. Once trained, the application of CNN to new eDNA metabarcoding data can be automated, which promises fast and easy deployment on the cloud for future eDNA analyses.

bioinformatics↗

Theory of local k-mer selection with applications to long-read alignment

MotivationSelecting a subset of k-mers in a string in a local manner is a common task in bioinformatics tools for speeding up computation. Arguably the most well-known and common method is the minimizer technique, which selects the lowest-ordered k-mer in a sliding window. Recently, it has been shown that minimizers are a sub-optimal method for selecting subsets of k-mers when mutations are present. There is however a lack of understanding behind the theory of why certain methods perform well. ResultsWe first theoretically investigate the conservation metric for k-mer selection methods. We derive an exact expression for calculating the conservation of a k-mer selection method. This turns out to be tractable enough for us to prove closed-form expressions for a variety of methods, including (open and closed) syncmers, (, b, n)-words, and an upper bound for minimizers. As a demonstration of our results, we modified the minimap2 read aligner to use a more optimal k-mer selection method and demonstrate that there is up to an 8.2% relative increase in number of mapped reads. Availability and supplementary informationSimulations and supplementary methods available at https://github.com/bluenote-1577/local-kmer-selection-results. os-minimap2 is a modified version of minimap2 and available at https://github.com/bluenote-1577/os-minimap2. Contactjshaw@math.toronto.edu

bioinformatics↗

Tripeptide loop closure: a detailed study of reconstructions based on Ramachandran distributions

Tripeptide loop closure (TLC) is a standard procedure to reconstruct protein backbone conformations, by solving a zero dimensional polynomial system yielding up to 16 solutions. In this work, we first show that multiprecision is required in a TLC solver to guarantee the existence and the accuracy of solutions. We then compare solutions yielded by the TLC solver against tripeptides from the Protein Data Bank. We show that these solutions are geometrically diverse (up to 3[A] RMSD with respect to the data), and sound in terms of potential energy. Finally, we compare Ramachandran distributions of data and reconstructions for the three amino acids. The distribution of reconstructions in the second angular space ({varphi}2,{psi} 2) stands out, with a rather uniform distribution leaving a central void. We anticipate that these insights, coupled to our robust implementation in the Structural Bioinformatics Library (https://sbl.inria.fr/doc/Tripeptide_loop_closure-user-manual.html), will boost the interest of TLC for structural modeling in general, and the generation of conformations of flexible loops in particular.

bioinformatics↗

VP-Detector: A 3D convolutional neural network for automated macromolecule localization and classification in cryo-electron tomograms

MotivationCryo-electron tomography (Cryo-ET) with sub-tomogram averaging (STA) is indispensable when studying macromolecule structures and functions in their native environments. However, current tomographic reconstructions suffer the low signal-to-noise (SNR) ratio and the missing wedge artifacts. Hence, automatic and accurate macromolecule localization and classification become the bottleneck problem for structural determination by STA. Here, we propose a 3D multi-scale dense convolutional neural network (MSDNet) for voxel-wise annotations of tomograms. Weighted focal loss is adopted as a loss function to solve the class imbalance. The proposed network combines 3D hybrid dilated convolutions (HDC) and dense connectivity to ensure an accurate performance with relatively few trainable parameters. 3D HDC expands the receptive field without losing resolution or learning extra parameters. Dense connectivity facilitates the re-use of feature maps to generate fewer intermediate feature maps and trainable parameters. Then, we design a 3D MSDNet based approach for fully automatic macromolecule localization and classification, called VP-Detector (Voxel-wise Particle Detector). VP-Detector is efficient because classification performs on the pre-calculated coordinates instead of a sliding window. ResultsWe evaluated the VP-Detector on simulated tomograms. Compared to the state-of-the-art methods, our method achieved a competitive performance on localization with the highest F1-score. We also demonstrated that the weighted focal loss improves the classification of hard classes. We trained the network on a part of training sets to prove the availability of training on relatively small datasets. Moreover, the experiment shows that VP-Detector has a fast particle detection speed, which costs less than 14 minutes on a test tomogram. Contactzsh@amss.ac.cn, xfcui@email.sdu.edu.cn, zhangfa@ict.ac.cn Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Coral: a web-based visual analysis tool for creating and characterizing cohorts

SummaryA main task in computational cancer analysis is the identification of patient subgroups (i.e., cohorts) based on metadata attributes (patient stratification) or genomic markers of response (biomarkers). Coral is a web-based cohort analysis tool that is designed to support this task: Users can interactively create and refine cohorts, which can then be compared, characterized, and inspected down to the level of single items. Coral visualizes the evolution of cohorts and also provides intuitive access to prevalence information. Furthermore, findings can be stored, shared, and reproduced via the integrated session management. Coral is pre-loaded with data from over 128,000 samples from the AACR Project GENIE, The Cancer Genome Atlas, and the Cell Line Encyclopedia. Availability and ImplementationCoral is publicly available at https://coral.caleydoapp.org. The source code is released at https://github.com/Caleydo/coral. Contactthomas.zichner@boehringer-ingelheim.com Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

VIQoR: a web service for Visually supervised protein Inference and protein Quantification

MotivationIn quantitative bottom-up mass spectrometry (MS)-based proteomics the reliable estimation of protein concentration changes from peptide quantifications between different biological samples is essential. This estimation is not a single task but comprises the two processes of protein inference and protein abundance summarization. Furthermore, due to the high complexity of proteomics data and associated uncertainty about the performance of these processes, there is a demand for comprehensive visualization methods able to integrate protein with peptide quantitative data including their post-translational modifications. Hence, there is a lack of a suitable tool that provides post-identification quantitative analysis of proteins with simultaneous interactive visualization. ResultsIn this article, we present VIQoR, a user-friendly web service that accepts peptide quantitative data of both labeled and label-free experiments and accomplishes the processes for relative protein quantification, along with interactive visualization modules, including the novel VIQoR plot. We implemented two parsimonious algorithms to solve the protein inference problem, while protein summarization is facilitated by a well established factor analysis algorithm called fast-FARMS followed by a weighted average summarization function that minimizes the effect of missing values. In addition, summarization is optimized by the so-called Global Correlation Indicator (GCI). We test the tool on three publicly available ground truth datasets and demonstrate the ability of the protein inference algorithms to handle degenerate peptides. We furthermore show that GCI increases the accuracy of the quantitative analysis in data sets with replicated design. Availability and implementationVIQoR is accessible at: http://computproteomics.bmb.sdu.dk:8192/app_direct/VIQoR/ The source code is available at: https://bitbucket.org/vtsiamis/viqor/ Contactveits@bmb.sdu.dk Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

MutaFrame - an interpretative visualization framework for deleteriousness prediction of missense variants in the human exome

MotivationHigh-throughput experiments are generating ever increasing amounts of various -omics data, so shedding new light on the link between human disorders, their genetic causes, and the related impact on protein behavior and structure. While numerous bioinformatics tools now exist that predict which variants in the human exome cause diseases, few tools predict the reasons why they might do so. Yet, understanding the impact of variants at the molecular level is a prerequisite for the rational development of targeted drugs or personalized therapies. ResultsWe present the updated MutaFrame webserver, which aims to meet this need. It offers two deleteriousness prediction softwares, DEOGEN2 and SNPMuSiC, and is designed for bioinformaticians and medical researchers who want to gain insights into the origins of monogenic diseases. It contains information at two levels for each human protein: its amino acid sequence and its 3-dimensional structure; we used the experimental structures whenever available, and modeled structures otherwise. MutaFrame also includes higher-level information, such as protein essentiality and protein-protein interactions. It has a user-friendly interface for the interpretation of results and a convenient visualization system for protein structures, in which the variant positions introduced by the user and other structural information are shown. In this way, MutaFrame aids our understanding of the pathogenic processes caused by single-site mutations and their molecular and contextual interpretation. AvailabilityMutaframe webserver at http://mutaframe.com

bioinformatics↗

STENCIL: A web templating engine for visualizing and sharing life science datasets

The ability to aggregate experimental data analysis and results into a concise and interpretable format is a key step in evaluating the success of an experiment. This critical step determines baselines for reproducibility and is a key requirement for data dissemination. However, in practice it can be difficult to consolidate data analyses that encapsulates the broad range of datatypes available in the life sciences. We present STENCIL, a web templating engine designed to organize, visualize, and enable the sharing of interactive data visualizations. STENCIL leverages a flexible web framework for creating templates to render highly customizable visual front ends. This flexibility enables researchers to render small or large sets of experimental outcomes, producing high-quality downloadable and editable figures that retain their original relationship to the source data. REST API based back ends provide programmatic data access and supports easy data sharing. STENCIL is a lightweight tool that can stream data from Galaxy, a popular bioinformatic analysis web platform. STENCIL has been used to support the analysis and dissemination of two large scale genomic projects containing the complete data analysis for over 2,400 distinct datasets. Code and implementation details are available on GitHub: https://github.com/CEGRcode/stencil

bioinformatics↗

Predicted Coronavirus Nsp5 Protease Cleavage Sites in the Human Proteome: A Resource for SARS-CoV-2 Research

BackgroundThe coronavirus nonstructural protein 5 (Nsp5) is a cysteine protease required for processing the viral polyprotein and is therefore crucial for viral replication. Nsp5 from several coronaviruses have also been found to cleave host proteins, disrupting molecular pathways involved in innate immunity. Nsp5 from the recently emerged SARS-CoV-2 virus interacts with and can cleave human proteins, which may be relevant to the pathogenesis of COVID-19. Based on the continuing global pandemic, and emerging understanding of coronavirus Nsp5-human protein interactions, we set out to predict what human proteins are cleaved by the coronavirus Nsp5 protease using a bioinformatics approach. ResultsUsing a previously developed neural network trained on coronavirus Nsp5 cleavage sites (NetCorona), we made predictions of Nsp5 cleavage sites in all human proteins. Structures of human proteins in the Protein Data Bank containing a predicted Nsp5 cleavage site were then examined, generating a list of 92 human proteins with a highly predicted and accessible cleavage site. Of those, 48 are expected to be found in the same cellular compartment as Nsp5. Analysis of this targeted list of proteins revealed molecular pathways susceptible to Nsp5 cleavage and therefore relevant to coronavirus infection, including pathways involved in mRNA processing, cytokine response, cytoskeleton organization, and apoptosis. ConclusionsThis study combines predictions of Nsp5 cleavage sites in human proteins with protein structure information and protein network analysis. We predicted cleavage sites in proteins recently shown to be cleaved in vitro by SARS-CoV-2 Nsp5, and we discuss how other potentially cleaved proteins may be relevant to coronavirus mediated immune dysregulation. The data presented here will assist in the design of more targeted experiments, to determine the role of coronavirus Nsp5 cleavage of host proteins, which is relevant to understanding the molecular pathology of SARS-CoV-2 infection.

bioinformatics↗

ECCsplorer: a pipeline to detect extrachromosomal circular DNA (eccDNA) from next-generation sequencing data

MotivationExtrachromosomal circular DNAs (eccDNAs) are ring-like DNA structures physically separated from the chromosomes with 100 bp to several megabasepairs in size. Apart from carrying tandemly repeated DNA, eccDNAs may also harbor extra copies of genes or recently activated transposable elements. As eccDNAs occur in all eukaryotes investigated so far and likely play roles in stress, cancer, and aging, they have been prime targets in recent research - with their investigation limited by the scarcity of computational tools. ResultsHere, we present the ECCsplorer, a bioinformatics pipeline to detect eccDNAs in any kind of organism or tissue using next-generation sequencing techniques. Following Illumina-sequencing of amplified circular DNA (circSeq), the ECCsplorer enables an easy and automated discovery of eccDNA candidates. The data analysis encompasses two major procedures: First, read mapping to the reference genome allows the detection of informative read distributions including high coverage, discordant mapping, and split reads. Second, reference-free comparison of read clusters from amplified eccDNA against control sample data reveals specifically enriched DNA circles. Both software parts can be run separately or jointly, depending on the individual aim or data availability. To illustrate the wide applicability of our approach, we analyzed semiartificial and published circSeq data from the model organisms H. sapiens and A. thaliana, and generated circSeq reads from the non-model crop B. vulgaris. We clearly identified eccDNA candidates from all datasets, with and without reference genomes. The ECCsplorer pipeline specifically detected mitochondrial mini-circles and retrotransposon activation, showcasing the ECCsplorers sensitivity and specificity. The derived eccDNA targets are valuable for a wide range of downstream investigations - from analysis of cancer-related eccDNAs over organelle genomics to identification of active transposable elements. Availability and implementationThe ECCsplorer pipeline is available on GitHub at https://github.com/crimBubble/ECCsplorer under the GNU license. ContactTony Heitkam (tony.heitkam@tu-dresden.de) Supplementary informationSupplementary data are available online.

bioinformatics↗

Prediction of the effect of pH on the aggregation and conditional folding of intrinsically disordered proteins with SolupHred and DispHred.

Proteins microenvironments modulate their structures. Binding partners, organic molecules, or dissolved ions can alter the proteins compaction, inducing aggregation or order-disorder conformational transitions. Surprisingly, bioinformatic platforms often disregard the protein context in their modeling. In recent work, we proposed that modeling how pH affects protein net charge and hydrophobicity might allow us to forecast pH-dependent aggregation and conditional disorder in intrinsically disordered proteins (IDPs). As these approaches showed remarkable success in recapitulating the available bibliographical data, we made these prediction methods available for the scientific community as two user-friendly web servers. SolupHred is the first dedicated software to predict pH-dependent aggregation, and DispHred is the first pH-dependent predictor of protein disorder. Here we dissect the features of these two software applications to train and assist scientists in studying pH-dependent conformational changes in IDPs.

bioinformatics↗

Identifying proximal RNA interactions from cDNA-encoded crosslinks with ShapeJumper

SHAPE-JuMP is a concise strategy for identifying close-in-space interactions in RNA molecules. Nucleotides in close three-dimensional proximity are crosslinked with a bi-reactive reagent that covalently links the 2-hydroxyl groups of the ribose moieties. The identities of crosslinked nucleotides are determined using an engineered reverse transcriptase that jumps across crosslinked sites, resulting in a deletion in the cDNA that is detected using massively parallel sequencing. Here we introduce ShapeJumper, a bioinformatics pipeline to process SHAPE-JuMP sequencing data and to accurately identify through-space interactions. ShapeJumper identifies proximal interactions with near-nucleotide resolution using an alignment strategy that is optimized to tolerate the unique non-templated reverse-transcription profile of the engineered crosslink-traversing reverse-transcriptase. JuMP-inspired strategies are now poised to replace adapter-ligation for detecting RNA-RNA interactions in most crosslinking experiments.

bioinformatics↗

Delimitation of the Tick-Borne Flaviviruses. Resolving the Tick-Borne Encephalitis and Louping-Ill Virus Paraphyletic Taxa

The tick-borne flavivirus (TBFV) group contains at least 12 members where five of them are important pathogens of humans inducing diseases with varying severity (from mild fever forms to acute encephalitis). The taxonomy structure of TBFV is not fully clarified at present. In particular, there is a number of paraphyletic issues of tick-borne encephalitis virus (TBEV) and louping-ill virus (LIV). In this study, we aimed to apply different bioinformatic approaches to analyze all available complete genome amino acid sequences to delineate TBFV members at the species level. Results showed that the European subtype of TBEV (TBEV-E) is a distinct species unit. LIV, in turn, should be separated into two species. Additional analysis of the diversity of TBEV and LIV antigenic determinants also demonstrate that TBEV-E and LIV are significantly different from other TBEV subtypes. The analysis of available literature provided data on other virus phenotypic particularities that supported our hypothesis. So, within the TBEV+LIV paraphyletic group, we offer to assign four species to get a more accurate understanding of the TBFV interspecies structure according to the modern monophyletic conception.

bioinformatics↗

A comprehensive in silico investigation into the nsSNPs of Drd2 gene predicts significant functional consequences in dopamine signaling and pharmacotherapy

DRD2 is a neuronal cell surface protein involved in brain development and function. Variations in the Drd2 gene have clinical significance since DRD2 is a pharmacotherapeutic target for treating psychiatric disorders like ADHD and schizophrenia. Despite numerous studies on the disease association of single nucleotide polymorphisms (SNPs) in the intronic regions, investigation into the coding regions is surprisingly limited. In this study, we aimed at identifying potential functionally and pharmaco-therapeutically deleterious non-synonymous SNPs of Drd2. A wide array of bioinformatics tools was used to evaluate the impact of nsSNPs on protein structure and functionality. Out of 260 nsSNPs retrieved from the dbSNP database, initially 9 were predicted as deleterious by 15 tools. Upon further assessment of their domain association, conservation profile, homology models and inter-atomic interaction, the mutant F389V was considered as the most impactful. In-depth analysis of F389V through Molecular Docking and Dynamics Simulation revealed a decline in affinity for its native agonist dopamine and an increase in affinity for the antipsychotic drug risperidone. Remarkable alterations in binding interactions and stability of the protein-ligand complex in simulated physiological conditions were also noted. These findings will improve our understanding of the consequence of nsSNPs in disease-susceptibility and therapeutic efficacy.

bioinformatics↗

STRipy: a graphical application for enhanced genotyping of pathogenic short tandem repeats in sequencing data

Short tandem repeats (STRs) are highly polymorphic with high mutation rates and expansions of STRs have been implicated as the causal variant in diseases. The application of genome sequencing in patients has recently allowed many new discoveries with over 50 disease causing loci known to date. There are several tools which allow genotyping of STRs from high-throughput sequencing (HTS) data. However, running these tools out of the box only allow around half of the known disease-causing loci to be genotyped, with lengths often limited to either read or fragment length which is less than the pathogenic cut-off for some diseases. While analysis tools can be customised to genotype extra loci, this requires proficiency in bioinformatics to set up, use, and analyse the resulting data, limiting their widespread usage by other researchers and clinicians. To address these issues, we have created a new software called STRipy that has an intuitive graphical interface and requires no specific skills for usage, thus significantly simplifying detection of STRs expansions from human HTS data. STRipy is able to target all known disease-causing STRs with genotyping performed with an established tool, ExpansionHunter, that is incorporated into the software. We have created additional functionality into STRipy to work with long alleles exceeding the fragment length. STRipy was validated using over 60 thousand simulated samples and was shown to work on whole genome sequencing of biological samples with pathogenic variants. Finally, we have used STRipy to acquire genotypes of pathogenic loci for thousands of samples from various populations which are provided to the user along with the data from the literature to assist with results interpretation. We believe the simplicity and breadth of STRipy will increase the testing of STR diseases in current datasets resulting in further diagnoses of rare diseases caused by STRs expansions.

bioinformatics↗

Refine your search to explore more results.