bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

Accelerating COVID-19 research with graph mining and transformer-based learning

In 2020, the White House released the, "Call to Action to the Tech Community on New Machine Readable COVID-19 Dataset," wherein artificial intelligence experts are asked to collect data and develop text mining techniques that can help the science community answer high-priority scientific questions related to COVID-19. The Allen Institute for AI and collaborators announced the availability of a rapidly growing open dataset of publications, the COVID-19 Open Research Dataset (CORD-19). As the pace of research accelerates, biomedical scientists struggle to stay current. To expedite their investigations, scientists leverage hypothesis generation systems, which can automatically inspect published papers to discover novel implicit connections. We present an automated general purpose hypothesis generation systems AGATHA-C and AGATHA-GP for COVID-19 research. The systems are based on graph-mining and the transformer model. The systems are massively validated using retrospective information rediscovery and proactive analysis involving human-in-the-loop expert analysis. Both systems achieve high-quality predictions across domains (in some domains up to 0.97% ROC AUC) in fast computational time and are released to the broad scientific community to accelerate biomedical research. In addition, by performing the domain expert curated study, we show that the systems are able to discover on-going research findings such as the relationship between COVID-19 and oxytocin hormone. ReproducibilityAll code, details, and pre-trained models are available at https://github.com/IlyaTyagin/AGATHA-C-GP CCS CONCEPTS* Applied computing [->] Bioinformatics; Document management and text processing; * Computing methodologies [->] Learning latent representations; Neural networks; Information extraction; Semantic networks.

bioinformatics↗

Overexpression of ELOVL6 has been associated with poor prognosis in patients with head and neck squamous cell carcinoma

Head and neck squamous cell carcinoma (HNSCC) is a high mortality disease. Extension of long-chain fatty acid family member 6 (ELOVL6) is a key enzyme involved in fat formation that catalyzes the elongation of saturated and monounsaturated fatty acids. Overexpression of ELOVL6 has been associated with obesity-related malignancies, including hepatocellular carcinoma, breast, colon, prostate, and pancreatic cancer. The following study investigated the role of ELOVL6 in HNSCC patients. Gene expression and clinicopathological analysis, enrichment analysis, and immune infiltration analysis were based on the Gene Expression Omnibus (GEO) and the Cancer Genome Atlas (TCGA), with additional bioinformatics analyses. The statistical analysis was conducted in R, and TIMER was used to analyze the immune response of ELOVL6 expression in HNSCC. The expression of ELOVL6 was related to tumor grade. Survival analysis showed that patients with high expression of ELOVL6 had a poor prognosis. Moreover, the results of GSEA enrichment analysis showed that ELOVL6 affects the occurrence of HNSCC through fatty acid metabolism, biosynthesis of unsaturated fatty acids, and other pathways. Finally, ELOVL6 verified by the Human Protein Atlas (HPA) database were consistent with the mRNA levels in HNSCC samples. ELOVL6 is a new biomarker for HNSCC that may be used as a potential predictor of the prognosis of human HNSCC.

bioinformatics↗

Spatially Interacting Phosphorylation Sites and Mutations in Cancer

Advances in mass-spectrometry have generated increasingly large-scale proteomics datasets containing tens of thousands of phosphorylation sites (phosphosites) that require prioritization. We develop a bioinformatics tool called HotPho and systematically discover 3D co-clustering of phosphosites and cancer mutations on protein structures. HotPho identifies 474 such hybrid clusters containing 1,255 co-clustering phosphosites, including RET p.S904/Y928, the conserved HRAS/KRAS p.Y96, and IDH1 p.Y139/IDH2 p.Y179 that are adjacent to recurrent mutations on protein structures not found by linear proximity approaches. Hybrid clusters, enriched in histone and kinase domains, frequently include expression-associated mutations experimentally shown as activating and conferring genetic dependency. Approximately 300 co-clustering phosphosites are verified in patient samples of 5 cancer types or previously implicated in cancer, including CTNNB1 p.S29/Y30, EGFR p.S720, MAPK1 p.S142, and PTPN12 p.S275. In summary, systematic 3D clustering analysis highlights nearly 3,000 likely functional mutations and over 1,000 cancer phosphosites for downstream investigation and evaluation of potential clinical relevance.

bioinformatics↗

Comparative Analysis of common alignment tools for single cell RNA sequencing

With the rise of single cell RNA sequencing new bioinformatic tools became available to handle specific demands, such as quantifying unique molecular identifiers and correcting cell barcodes. Here, we analysed several datasets with the most common alignment tools for scRNA-seq data. We evaluated differences in the whitelisting, gene quantification, overall performance and potential variations in clustering or detection of differentially expressed genes. We compared the tools Cell Ranger 5, STARsolo, Kallisto and Alevin on three published datasets for human and mouse, sequenced with different versions of the 10X sequencing protocol. Striking differences have been observed in the overall runtime of the mappers. Besides that Kallisto and Alevin showed variances in the number of valid cells and detected genes per cell. Kallisto reported the highest number of cells, however, we observed an overrepresentation of cells with low gene content and unknown celtype. Conversely, Alevin rarely reported such low content cells. Further variations were detected in the set of expressed genes. While STARsolo, Cell Ranger 5 and Alevin released similar gene sets, Kallisto detected additional genes from the Vmn and Olfr gene family, which are likely mapping artifacts. We also observed differences in the mitochondrial content of the resulting cells when comparing a prefiltered annotation set to the full annotation set that includes pseudogenes and other biotypes. Overall, this study provides a detailed comparison of common scRNA-seq mappers and shows their specific properties on 10X Genomics data. Key messagesO_LIMapping and gene quantifications are the most resource and time intensive steps during the analysis of scRNA-Seq data. C_LIO_LIThe usage of alternative alignment tools reduces the time for analysing scRNA-Seq data. C_LIO_LIDifferent mapping strategies influence key properties of scRNA-SEQ e.g. total cell counts or genes per cell C_LIO_LIA better understanding of advantages and disadvantages for each mapping algorithm might improve analysis results. C_LI

bioinformatics↗

Supervised biomedical semantic similarity

BackgroundSemantic similarity between concepts in knowledge graphs is essential for several bioinformatics applications, including the prediction of protein-protein interactions and the discovery of associations between diseases and genes. Although knowledge graphs describe entities in terms of several perspectives (or semantic aspects), state-of-the-art semantic similarity measures are general-purpose. This can represent a challenge since different use cases for the application of semantic similarity may need different similarity perspectives and ultimately depend on expert knowledge for manual fine-tuning. ResultsWe present a new approach that uses supervised machine learning to tailor aspect-oriented semantic similarity measures to fit a particular view on biological similarity or relatedness. We implement and evaluate it using different combinations of representative semantic similarity measures and machine learning methods with four biological similarity views: protein-protein interaction, protein function similarity, protein sequence similarity and phenotype-based gene similarity. ConclusionsThe results demonstrate that our approach outperforms non-supervised methods, producing semantic similarity models that fit different biological perspectives significantly better than the commonly used manual combinations of semantic aspects.

bioinformatics↗

Identification of biomarkers and candidate inhibitors for multiple myeloma

Multiple myeloma (MM) is a plasma cell malignancy that causes the overabundance of monoclonal paraprotein (M protein) and organ damages. In our study, we aim to identify biological markers and processes of MM using a bioinformatics method to elucidate their potential pathogenesis. The gene expression profiles of the GSE153626 datasets were originally produced by using the high-throughput Illumina HiSeq 4000 (Mus musculus). The functional categories and biochemical pathways were identified and analyzed by the Kyoto Encyclopedia of Genes and Genomes pathway (KEGG), Gene Ontology (GO), and Reactom enrichment. KEGG and GO results showed the biological pathways related to immune dysfunction and signal transduction are mostly affected in the development of MM. Moreover, we identified several genes including Gngt2, Foxp3, and Cd3g were involved in the regulation of immune cells. We further predicted new inhibitors that have the ability to block the progression of MM based on the L1000fwd analysis. Therefore, this study provides further insights into the underlying pathogenesis of MM.

bioinformatics↗

Genome-wide survey of odorant-binding proteins in the dwarf honey bee Apis florea

Odorant binding proteins (OBPs) in insects bind to volatile chemical cue and help in their binding to odorant receptors. The odor coding hypothesis states that OBPs may bind with specificity to certain volatiles and aid the insect in various behaviours. Honeybees are eusocial insects with complex behaviour that requires olfactory inputs. Here, we have identified and annotated odorant binding proteins from the genome of the dwarf honey bee, Apis florea using an exhaustive homology-based bioinformatic pipeline and analyzed the evolutionary relationships between the OBP subfamilies. Our study suggests that Minus-C subfamily may have diverged from the Classic subfamily of odorant binding proteins in insects.

bioinformatics↗

Assessing heterogeneity in spatial data using the HTA index with applications to spatial transcriptomics and imaging

MotivationTumour heterogeneity is being increasingly recognised as an important characteristic of cancer and as a determinant of prognosis and treatment outcome. Emerging spatial transcriptomics data hold the potential to further our understanding of tumour heterogeneity and its implications. However, existing statistical tools are not sufficiently powerful to capture heterogeneity in the complex setting of spatial molecular biology. ResultsWe provide a statistical solution, the HeTerogeneity Average index (HTA), specifically designed to handle the multivariate nature of spatial transcriptomics. We prove that HTA has an approximately normal distribution, therefore lending itself to efficient statistical assessment and inference. We first demonstrate that HTA accurately reflects the level of heterogeneity in simulated data. We then use HTA to analyse heterogeneity in two cancer spatial transcriptomics datasets: spatial RNA sequencing by 10x Genomics and spatial transcriptomics inferred from H&E. Finally, we demonstrate that HTA also applies to 3D spatial data using brain MRI. In spatial RNA sequencing we use a known combination of molecular traits to assert that HTA aligns with the expected outcome for this combination. We also show that HTA captures immune-cell infiltration at multiple resolutions. In digital pathology we show how HTA can be used in survival analysis and demonstrate that high levels of heterogeneity may be linked to poor survival. In brain MRI we show that HTA differentiates between normal ageing, Alzheimers disease and two tumours. HTA also extends beyond molecular biology and medical imaging, and can be applied to many domains, including GIS. Availability . Contactlevyalona@gmail.com zohar.yakhini@gmail.com Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Open Natural Products Research: Curation and Dissemination of Biological Occurrences of Chemical Structures through Wikidata

Contemporary bioinformatic and chemoinformatic capabilities hold promise to reshape knowledge management, analysis and interpretation of data in natural products research. Currently, reliance on a disparate set of non-standardized, insular, and specialized databases presents a series of challenges for data access, both within the discipline and for integration and interoperability between related fields. The fundamental elements of exchange are referenced structure-organism pairs that establish relationships between distinct molecular structures and the living organisms from which they were identified. Consolidating and sharing such information via an open platform has strong transformative potential for natural products research and beyond. This is the ultimate goal of the newly established LOTUS initiative, which has now completed the first steps toward the harmonization, curation, validation and open dissemination of 750,000+ referenced structure-organism pairs. LOTUS data is hosted on Wikidata and regularly mirrored on https://lotus.naturalproducts.net. Data sharing within the Wikidata framework broadens data access and interoperability, opening new possibilities for community curation and evolving publication models. Furthermore, embedding LOTUS data into the vast Wikidata knowledge graph will facilitate new biological and chemical insights. The LOTUS initiative represents an important advancement in the design and deployment of a comprehensive and collaborative natural products knowledge base. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=61 SRC="FIGDIR/small/433265v3_ufig1.gif" ALT="Figure 1"> View larger version (14K): org.highwire.dtl.DTLVardef@184136borg.highwire.dtl.DTLVardef@16e0d6org.highwire.dtl.DTLVardef@30ae1org.highwire.dtl.DTLVardef@1bf4825_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Published Anti-SARS-CoV-2 In Vitro Hits Share Common Mechanisms of Action that Synergize with Antivirals

The global efforts in the past few months have led to the discovery of around 200 drug repurposing candidates for COVID-19. Although most of them only exhibited moderate anti- SARS-CoV-2 activity, gaining more insights into their mechanisms of action could facilitate a better understanding of infection and the development of therapeutics. Leveraging large-scale drug-induced gene expression profiles, we found 36% of the active compounds regulate genes related to cholesterol homeostasis and microtubule cytoskeleton organization. The expression change upon drug treatment was further experimentally confirmed in human lung primary small airway. Following bioinformatics analysis on COVID-19 patient data revealed that these genes are associated with COVID-19 patient severity. The expression level of these genes also has predicted power on anti-SARS-CoV-2 efficacy in vitro, which led to the discovery of monensin as an inhibitor of SARS-CoV-2 replication in Vero-E6 cells. The final survey of recent drug- combination data indicated that drugs co-targeting cholesterol homeostasis and microtubule cytoskeleton organization processes more likely present a synergistic effect with antivirals. Therefore, potential therapeutics should be centered around combinations of targeting these processes and viral proteins.

bioinformatics↗

Evotuning protocols for Transformer-based variant effect prediction on multi-domain proteins

Accurate variant effect prediction has broad impacts on protein engineering. Recent machine learning approaches toward this end are based on representation learning, by which feature vectors are learned and generated from unlabeled sequences. However, it is unclear how to effectively learn evolutionary properties of an engineering target protein from homologous sequences, taking into account the proteins sequence-level structure called domain architecture (DA). Additionally, no optimal protocols are established for incorporating such properties into Transformer, the neural network well-known to perform the best in natural language processing research. This article proposes DA-aware evolutionary fine-tuning, or "evotuning", protocols for Transformer-based variant effect prediction, considering various combinations of homology search, fine-tuning, and sequence vectorization strategies. We exhaustively evaluated our protocols on diverse proteins with different functions and DAs. The results indicated that our protocols achieved significantly better performances than previous DA-unaware ones. The visualizations of attention maps suggested that the structural information was incorporated by evotuning without direct supervision, possibly leading to better prediction accuracy. Availabilityhttps://github.com/dlnp2/evotuning_protocols_for_transformers Supplementary informationSupplementary data are available at Briefings in Bioinformatics online.

bioinformatics↗

Make Interactive Complex Heatmaps in R

Heatmap is a powerful visualization method on two-dimensional data to reveal patterns shared by subsets of rows and columns. In R, there are many packages that make heatmaps. Among them, ComplexHeatmap provides rich tools for constructing highly customizable heatmaps. It can easily establish connections between information from multiple sources by automatically concatenating and adjusting multiple heatmaps as well as complex annotations, which makes it widely applied in data analysis in various fields, especially in Bioinformatics. Nevertheless, the limit of ComplexHeatmap still exists. It only generates static plots which restricts deeper inspections on complex heatmaps, e.g., to look into a subset of rows and columns when a specific pattern of interest is observed from the heatmap. In this work, we described a new R/Bioconductor package InteractiveComplexHeatmap that brings interactivity to ComplexHeatmap. InteractiveComplexHeatmap is designed with an easy-to-use interface where static complex heatmaps can be directly exported to an interactive Shiny web application only with one extra line of code. The interactive application contains comprehensive tools for manipulating heatmaps. Besides common tools as supported in other interactive heatmap packages, InteractiveComplexHeatmap additionally supports, e.g., selecting over multiple heatmaps and searching heatmaps via row or column labels. Also, InteractiveComplexHeatmap provides methods for exporting static heatmaps from other popular heatmap functions, e.g., heatmap.2() or pheatmap(), to interactive heatmap applications. Finally, InteractiveComplexHeatmap provides flexible functionalities for integrating interactive heatmap widgets into other Shiny applications. InteractiveComplexHeatmap provides a user interface for self-defining response to the selection events on heatmaps, which helps to implement more complex Shiny web applications.

bioinformatics↗

CHIPS: A Snakemake pipeline for quality control and reproducible processing of chromatin profiling data

MotivationThe chromatin profile measured by ATAC-seq, ChIP-seq, or DNase-seq experiments can identify genomic regions critical in regulating gene expression and provide insights on biological processes such as diseases and development. However, quality control and processing chromatin profiling data involve many steps, and different bioinformatics tools are used at each step. It can be challenging to manage the analysis. ResultsWe developed a Snakemake pipeline called CHIPS (CHromatin enrichment Processor) to streamline the processing of ChIP-seq, ATAC-seq, and DNase-seq data. The pipeline supports single- and paired-end data and is flexible to start with FASTQ or BAM files. It includes basic steps such as read trimming, mapping, and peak calling. In addition, it calculates quality control metrics such as contamination profiles, PCR bottleneck coefficient, the fraction of reads in peaks, percentage of peaks overlapping with the union of public DNaseI hypersensitivity sites, and conservation profile of the peaks. For downstream analysis, it carries out peak annotations, motif finding, and regulatory potential calculation for all genes. The pipeline ensures that the processing is robust and reproducible. AvailabilityCHIPS is available at https://github.com/liulab-dfci/CHIPS

bioinformatics↗

Refget: standardised access to reference sequences

Reference sequences are essential in creating a baseline of knowledge for many common bioinformatics methods, especially those using genomic sequencing. We have created refget, a Global Alliance for Genomics and Health API specification to access reference sequences and sub-sequences using an identifier derived from the sequence itself. We present four reference implementations across in-house and cloud infrastructure, a compliance suite and a web report used to ensure specification conformity across implementations. https://w3id.org/ga4gh/refget.

bioinformatics↗

OCD.py - Characterizing immunoglobulin inter-domain orientations

SummaryInter-domain orientations between immunoglobulin domains are important for the modeling and engineering of novel antibody therapeutics. Previous tools to describe these orientations are applicable only to the variable domains of antibodies and T-cell receptors. We present the "Orientation of Cylindrical Domains (OCD)" tool, which employs a transferable approach to calculate inter-domain orientations for all immunoglobulin domains. Based on a reference structure, the OCD tool automatically builds a suitable reference coordinate system for each domain. Through alignment, the reference coordinate systems are transferred onto the sample to calculate six measures which fully characterize the inter-domain orientation. Availability and implementationThe OCD approach is implemented as a stand-alone Python script, OCD.py, which can handle multiple types of data input for the analysis of single structures and molecular dynamics trajectories alike. OCD.py is available at https://github.com/liedllab/OCD under MIT license. Supplementary InformationSupplementary information and data are available at Bioinformatics online.

bioinformatics↗

Immunoinformatic Approach for the identification of T Cell and B Cell Epitopes in the Surface Glycoprotein and Designing a Potent Multiepitope Vaccine Construct Against SARS-CoV-2 including the new UK variant

The emergence of a novel coronavirus in China in late 2019 has turned into a SARS-CoV-2 pandemic affecting several millions of people worldwide in a short span of time with high fatality. The crisis is further aggravated by the emergence and evolution of new variant SARS-CoV-2 strains in UK during December, 2020 followed by their transmission to other countries. A major concern is that prophylaxis and therapeutics are not available yet to control and prevent the virus which is spreading at an alarming rate, though several vaccine trials are in the final stage. As vaccines are developed through various strategies, their immunogenic potential may drastically vary and thus pose several challenges in offering both arms of immunity such as humoral and cell-mediated immune responses against the virus. In this study, we adopted an immunoinformatics-aided identification of B cell and T cell epitopes in the Spike protein, which is a surface glycoprotein of SARS-CoV-2, for developing a new Multiepitope vaccine construct (MEVC). MEVC has 575 amino acids and comprises adjuvants and various cytotoxic T-lymphocyte (CTL), helper T-lymphocyte (HTL), and B-cell epitopes that possess the highest affinity for the respective HLA alleles, assembled and joined by linkers. The computational data suggest that the MEVC is non-toxic, non-allergenic and thermostable with the capability to elicit both humoral and cell-mediated immune responses. The population coverage of various countries affected by COVID-19 with respect to the selected B and T cell epitopes in MEVC was also investigated. Subsequently, the biological activity of MEVC was assessed by bioinformatic tools using the interaction between the vaccine candidate and the innate immune system receptors TLR3 and TLR4. The epitopes of the construct were analyzed with that of the strains belonging to various clades including the new variant UK strain having multiple unique mutations in S protein. Due to the advantageous features, the MEVC can be tested in vitro for more practical validation and the study offers immense scope for developing a potential vaccine candidate against SARS-CoV-2 in view of the public health emergency associated with COVID-19 disease caused by SARS-CoV-2.

bioinformatics↗

RCytoGPS: An R Package for Reading and Visualizing Cytogenetics Data

SummaryCytogenetics data, or karyotypes, are among the most common clinically used forms of genetic data. Karyotypes are stored as standardized text strings using the International System for Human Cytogenomic Nomenclature (ISCN). Historically, these data have not been used in large-scale computational analyses due to limitations in the ISCN text format and structure. Recently developed computational tools such as CytoGPS have enabled large-scale computational analyses of karyotypes. To further enable such analyses, we have now developed RCytoGPS, an R package that takes JSON files generated from CytoGPS.org and converts them into objects in R. This conversion facilitates the analysis and visualizations of karyotype data. In effect this tool streamlines the process of performing large-scale karyotype analyses, thus advancing the field of computational cytogenetic pathology. Availability and ImplementationFreely available at https://CRAN.R-project.org/package=RCytoGPS Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Computational Evaluation of DNA Metabarcoding for Universal Diagnostics of Invasive Insect Pests

Appropriate design and selection of PCR primers plays a critical role in determining the sensitivity and specificity of a metabarcoding assay. Despite several studies applying metabarcoding to insect pest surveillance, the diagnostic performance of the short "mini-barcodes" required by high-throughput sequencing platforms has not been established across the broader taxonomic diversity of invasive insects. We address this by computationally evaluating the diagnostic sensitivity and predicted amplification bias for 68 published and novel cytochrome c oxidase subunit 1 (COI) primers on a curated database of 110,676 insect species, including 2,625 registered on global invasive species lists. We find that mini-barcodes between 125-257 bp can provide comparable resolution to the full-length barcode for both invasive insect pests and the broader Insecta, conditional upon the subregion of COI targeted and the genetic similarity threshold used to identify species. Taxa that could not be identified by any barcode lengths were phylogenetically clustered within problem groups, many arising through taxonomic inconsistencies rather than insufficient diagnostic information within the barcode itself. Substantial variation in predicted PCR bias was seen across published primers, with those including 4-5 degenerate nucleotide bases showing almost no mismatch to major insect orders. While not completely universal, a single COI mini-barcode can successfully differentiate the majority of pest and non-pest insects from their congenerics, even at the small amplicon size imposed by 2 x 150 bp sequencing. We provide a ranked summary of high-performing primers and discuss the bioinformatic steps required to curate reliable reference databases for metabarcoding studies.

bioinformatics↗