bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

powerEQTL: An R package and shiny application for sample size and power calculation of bulk tissue and single-cell eQTL analysis

SummaryGenome-wide association studies (GWAS) have revealed thousands of genetic loci for common diseases. One of the main challenges in the post-GWAS era is to understand the causality of the genetic variants. Expression quantitative trait locus (eQTL) analysis has been proven to be an effective way to address this question by examining the relationship between gene expression and genetic variation in a sufficiently powered cohort. However, it is often tricky to determine the sample size at which a variant with a specific allele frequency will be detected to associate with gene expression with sufficient power. This is particularly demanding with single-cell RNAseq studies. Therefore, a user-friendly tool to perform power analysis for eQTL at both bulk tissue and single-cell level will be critical. Here, we presented an R package called powerEQTL with flexible functions to calculate power, minimal sample size, or detectable minor allele frequency in both bulk tissue and single-cell eQTL analysis. A user-friendly, program-free web application is also provided, allowing customers to calculate and visualize the parameters interactively. Availability and implementationThe powerEQTL R package source code and online tutorial are freely available at CRAN: https://cran.r-project.org/web/packages/powerEQTL/. The R shiny application is publicly hosted at https://bwhbioinfo.shinyapps.io/powerEQTL/. ContactXianjun Dong (xdong@rics.bwh.harvard.edu), Weiliang Qiu (weiliang.qiu@sanofi.com) Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

TranSuite: a software suite for accurate translation and characterization of transcripts

Protein translation programs often select the longest open reading frame (ORF) in a transcript leading to numerous inaccurate and mis-annotated ORFs in databases. Unproductive transcript isoforms containing premature termination codons (PTCs) are potential substrates for nonsense-mediated decay (NMD). These transcripts often contain truncated ORFs but are incorrectly annotated due to selection of a long ORF beginning at an AUG downstream of the PTC despite the transcript containing the authentic translation start AUG. In gene expression and alternative splicing analyses, it is important to identify transcript isoforms which code for different protein variants and to distinguish these from potential NMD substrates. Here, we present TranSuite, a pipeline of bioinformatics tools that address these challenges by performing accurate translations, characterizing alternative ORFs and identifying NMD and other features of transcripts in newly assembled and existing transcriptomes. Directly comparing ORFs defined by TranSuite and TransDecoder for the Arabidopsis transcriptome AtRTD2 identified ORF mis-calling in over 16k (27%) of transcripts by TransDecoder.

bioinformatics↗

TIPS: Trajectory Inference of Pathway Significance through Pseudotime Comparison for Functional Assessment of single-cell RNAseq Data

Recent advances in bioinformatics analyses have led to the development of novel tools enabling the capture and trajectory mapping of single-cell RNA sequencing (scRNAseq) data. However, there is a lack of methods to assess the contributions of biological pathways and transcription factors to an overall developmental trajectory mapped from scRNAseq data. In this manuscript, we present a simplified approach for trajectory inference of pathway significance (TIPS) that leverages existing knowledgebases of functional pathways and transcription factor targets to enable further mechanistic insights into a biological process. TIPS returns both the key pathways whose changes are associated with the process of interest, as well as the individual genes that best reflect these changes. TIPS also provides insight into the relative timing of pathway changes, as well as a suite of visualizations to enable simplified data interpretation of scRNAseq libraries generated using a wide range of techniques. The TIPS package can be run through either a web server, or downloaded as a user-friendly GUI run in R, and may serve as a useful tool to help biologists perform deeper functional analyses and visualization of their single-cell and/or large cohort RNAseq data.

bioinformatics↗

Genomic Copy Number Signatures Based Classifiers for Subtype Identification in Cancer

Copy number aberrations (CNA) are one of the most important classes of genomic mutations related to oncogenetic effects. In the past three decades, a vast amount of CNA data has been generated by molecular-cytogenetic and genome sequencing based methods. While this data has been instrumental in the identification of cancer-related genes and promoted research into the relation between CNA and histo-pathologically defined cancer types, the heterogeneity of source data and derived CNV profiles pose great challenges for data integration and comparative analysis. Furthermore, a majority of existing studies have been focused on the association of CNA to pre-selected "driver" genes with limited application to rare drivers and other genomic elements. In this study, we developed a bioinformatics pipeline to integrate a collection of 44,988 high-quality CNA profiles of high diversity. Using a hybrid model of neural networks and attention algorithm, we generated the CNA signatures of 31 cancer subtypes, depicting the uniqueness of their respective CNA landscapes. Finally, we constructed a multi-label classifier to identify the cancer type and the organ of origin from copy number profiling data. The investigation of the signatures suggested common patterns, not only of physiologically related cancer types but also of clinico-pathologically distant cancer types such as different cancers originating from the neural crest. Further experiments of classification models confirmed the effectiveness of the signatures in distinguishing different cancer types and demonstrated their potential in tumor classification.

bioinformatics↗

CellVGAE: An unsupervised scRNA-seq analysis workflow with graph attention networks

AO_SCPLOWBSTRACTC_SCPLOWCurrently, single-cell RNA sequencing (scRNA-seq) allows high-resolution views of individual cells, for libraries of up to (tens of) thousands of samples. In this study, we introduce the use of graph neural networks (GNN) in the unsupervised study of scRNA-seq data, namely for dimensionality reduction and clustering. Motivated by the success of non-neural graph-based techniques in bioinformatics, as well as the now common feedforward neural networks being applied to scRNA-seq measurements, we develop an architecture based on a variational graph autoencoder with graph attention layers that works directly on the connectivity of cells. With the help of three case studies, we show that our model, named CellVGAE, can be effectively used for exploratory analysis, even on challenging datasets, by extracting meaningful features from the data and providing the means to visualise and interpret different aspects of the model. Furthermore, we evaluate the dimensionality reduction and clustering performance on 9 well-annotated datasets, where we compare with leading neural and non-neural techniques. CellVGAE outperforms competing methods in all 9 scenarios. Finally, we show that CellVGAE is more interpretable than existing architectures by analysing the graph attention coefficients. The software and code to generate all the figures are available at https://github.com/davidbuterez/CellVGAE.

bioinformatics↗

easyMF: A Web Platform for Matrix Factorization-based Biological Discovery from Large-scale Transcriptome Data

With the development of high-throughput experimental technologies, large-scale RNA sequencing (RNA-Seq) data have been and continue to be produced, but have led to challenges in extracting relevant biological knowledge hidden in the produced high-dimensional gene expression matrices. Here, we present easyMF, a user-friendly web platform that aims to facilitate biological discovery from large-scale transcriptome data through matrix factorization (MF). The easyMF platform enables users with little bioinformatics experience to streamline transcriptome analysis from raw reads to gene expression and to decompose expression matrix from thousands of genes to a handful of metagenes. easyMF also offers a series of functional modules for metagene-based exploratory analysis with an emphasis on functional gene discovery. As a modular, containerized and open-source platform, easyMF can be customized to satisfy users specific demands and deployed as a web server for broad applications. easyMF is freely available at https://github.com/cma2015/easyMF. We demonstrated the application of easyMF with four case studies using 940 RNA sequencing datasets from maize (Zea mays L.).

bioinformatics↗

Identification of viral-mediated pathogenic mechanisms in neurodegenerative diseases using network-based approaches

During the course of a viral infection, virus-host protein-protein interactions (PPIs) play a critical role in allowing viruses to evade host immune responses, replicate and hence survive within the host. These interspecies molecular interactions can lead to viral-mediated perturbations of the human interactome causing the generation of various complex diseases, from cancer to neurodegenerative diseases (NDs). There are evidences suggesting that viral-mediated perturbations are a possible pathogenic aetiology in several NDs such as Amyloid Later Sclerosis, Parkinsons disease, Alzheimers disease and Multiple Sclerosis (MS), as they can cause degeneration of neurons via both direct and/or indirect actions. These diseases share several common pathological mechanisms, as well as unique disease mechanisms that reflect disease phenotype. NDs are chronic degenerative diseases of the central nervous system and current therapeutic approaches provide only mild symptomatic relief rather than treating the disease at heart, therefore there is unmet need for the discovery of novel therapeutic targets and pharmacotherapies. In this paper we initially review databases and tools that can be utilized to investigate viral-mediated perturbations in complex NDs using network-based analysis by examining the interaction between the ND-related PPI disease networks and the virus-host PPI network. Afterwards we present our integrative network-based bioinformatics approach that accounts for pathogen-genes-disease related PPIs with the aim to identify viral-mediated pathogenic mechanisms focusing in MS disease. We identified 7 high centrality nodes that can act as disease communicator nodes and exert systemic effects in the MS enriched KEGG pathways network. In addition, we identified 12 KEGG pathways targeted by 67 viral proteins from 8 viral species that might exert viral-mediated pathogenic mechanisms in MS by interacting with the disease communicator nodes. Finally, our analysis highlighted the Th17 differentiation pathway, a hub-bottleneck disease communicator node and part of the 12 underlined KEGG pathways, as a key viral-mediated pathogenic mechanism and a possible therapeutic target for MS disease.

bioinformatics↗

Identification of Key Genes Potentially Related to Triple Receptor Negative Breast Cancer by Microarray Analysis

Triple receptor negative breast cancer (TNBC) is the type of gynecological cancer in the elderly women. This study is aimed to explore molecular mechanism of TNBC via bioinformatics analysis. The gene expression profiles of GSE88715 (including 38 TNBC and 38 normal control) was downloaded from the Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) were screened using the limma package in R software. Pathway and gene ontology (GO) enrichment analysis were performed based on various pathway dabases and GO database. Then, InnateDb interactome database, Cytoscape and PEWCC1 were applied to construct the protein-protein interaction (PPI) network and screen hub genes. Similarly, miRNet database, NetworkAnalyst database and Cytoscape were applied to construct the target gene - miRNA network and target gene - TF network, and screen targate genes. Pathway and GO enrichment analysis was further performed for hub genes, gene clusters identified via module analysis and targate genes. The expression of hub genes with prognostic values was validated on the UALCAN, cBio Portal, The Human Protein Atlas, receiver operator characteristic (ROC) curve analysis, RT-PCR analysis and immune infiltration analysis. A total of 949 DEGs were identified in TNBC (469 up regulated genes, and 480 down regulated genes), and they were mainly enriched in the terms of phospholipases, toxoplasmosis, immune response, cell surface, glycolysis, biosynthesis of amino acids, carboxylic acid metabolic process and organic substance catabolic process extracellular space. Hub genes including UBD, HLA-B, MYC and HSP90AB1 were identified via PPI network and modules, which were mainly enriched in immune response, antigen processing and presentation, cell cycle and pathways in cancer. Targate genes including CCDC80, PEG10, HOPX and CCNA2 were identified via target gene - miRNA network and target gene - TF network, which were mainly enriched in extracellular structure organization, validated targets of C-MYC transcriptional activation, ensemble of genes encoding core extracellular matrix including ECM glycoproteins and cell cycle. The top five significantly overexpressed mRNA (ADAM15, BATF, NOTCH3, ITGAX and SDC1) and the top five significantly underexpressed mRNA (RPL4, EEF1G, RPL3, RBMX and ABCC2) were selected for further validation in TNBCpatients and healthy controls. Analysis of the expression of genes in the various databases showed that ADAM15, BATF, NOTCH3, ITGAX, SDC1, RPL4, EEF1G, RPL3, RBMX and ABCC2 expressions have a cancer specific pattern in TNBC. Collectively, ADAM15, BATF, NOTCH3, ITGAX, SDC1, RPL4, EEF1G, RPL3, RBMX and ABCC2 may be useful candidate biomarkers for TNBC diagnosis, prognosis and theraputic targates.

bioinformatics↗

SLIDR and SLOPPR: Flexible identification of spliced leader trans-splicing and prediction of eukaryotic operons from RNA-Seq data

BackgroundSpliced leader (SL) trans-splicing replaces the 5 end of pre-mRNAs with the spliced leader, an exon derived from a specialised non-coding RNA originating from elsewhere in the genome. This process is essential for resolving polycistronic pre-mRNAs produced by eukaryotic operons into monocistronic transcripts. SL trans-splicing and operons may have independently evolved multiple times throughout Eukarya, yet our understanding of these phenomena is limited to only a few well-characterised organisms, most notably C. elegans and trypanosomes. The primary barrier to systematic discovery and characterisation of SL trans-splicing and operons is the lack of computational tools for exploiting the surge of transcriptomic and genomic resources for a wide range of eukaryotes. ResultsHere we present two novel pipelines that automate the discovery of SLs and the prediction of operons in eukaryotic genomes from RNA-Seq data. SLIDR assembles putative SLs from 5 read tails present after read alignment to a reference genome or transcriptome, which are then verified by interrogating corresponding SL RNA genes for sequence motifs expected in bona fide SL RNA molecules. SLOPPR identifies RNA-Seq reads that contain a given 5 SL sequence, quantifies genomewide SL trans-splicing events and predicts operons via distinct patterns of SL trans-splicing events across adjacent genes. We tested both pipelines with organisms known to carry out SL trans-splicing and organise their genes into operons, and demonstrate that 1) SLIDR correctly detects expected SLs and often discovers novel SL variants; 2) SLOPPR correctly identifies functionally specialised SLs, correctly predicts known operons and detects plausible novel operons. ConclusionsSLIDR and SLOPPR are flexible tools that will accelerate research into the evolutionary dynamics of SL trans-splicing and operons throughout Eukarya and improve gene discovery and annotation for a wide-range of eukaryotic genomes. Both pipelines are implemented in Bash and R and are built upon readily available software commonly installed on most bioinformatics servers. Biological insight can be gleaned even from sparse, low-coverage datasets, implying that an untapped wealth of information can be derived from existing RNA-Seq datasets as well as from novel full-isoform sequencing protocols as they become more widely available.

bioinformatics↗

Multi-class Cancer Classification and Biomarker Identification using Deep Learning

Genetic data is important for analysing cellular functions whose disruption gives rise to various kinds of cancer. The intricacies of gene interaction are captured in various kinds of data for cancer detection through sequencing technology, but diagnosis, prognosis and treatment are still hard. Advent of machine learning helped researchers in supervised and unsupervised learning tasks along with gene identification but resourcefulness has not been overtly satisfactory. This research revolves around multi-class cancer classification, feature extraction and relevant gene identification through deep learning methods for 12 different types of cancers using RNA-SEQ from The Cancer Genome Atlas. It has been constrained by hardware resource availability and within them the experiments that have been performed have shown promising results. Stacked De-noising Autoencoders were used for feature extraction and biomarker identification while 1D Convolutional Neural Networks for classification. Classification was performed with extracted features and relevant genes, which gave average performance of around 94% and 95% respectively. We were able to identify generic cancer-related pathways and their associated genes through Stacked De-noising Auto-encoders generated weight matrix and features. The common pathways include WNT Signalling Pathway, Angiogenesis. Moreover, across all pathways some recurrent genes were observed, namely: PIK3C2G, PCDHB8, WNT10A and these genes were found, in literature, to be involved in multiple types of cancer. The proposed approach shows superior performance and promise against traditional techniques used by bioinformatics community, in terms of accuracy and relevant gene identification.

bioinformatics↗

Fast end-to-end learning on protein surfaces

Proteins biological functions are defined by the geometric and chemical structure of their 3D molecular surfaces. Recent works have shown that geometric deep learning can be used on mesh-based representations of proteins to identify potential functional sites, such as binding targets for potential drugs. Unfortunately though, the use of meshes as the underlying representation for protein structure has multiple drawbacks including the need to pre-compute the input features and mesh connectivities. This becomes a bottleneck for many important tasks in protein science. In this paper, we present a new framework for deep learning on protein structures that addresses these limitations. Among the key advantages of our method are the computation and sampling of the molecular surface on-the-fly from the underlying atomic point cloud and a novel efficient geometric convolutional layer. As a result, we are able to process large collections of proteins in an end-to-end fashion, taking as the sole input the raw 3D coordinates and chemical types of their atoms, eliminating the need for any hand-crafted pre-computed features. To showcase the performance of our approach, we test it on two tasks in the field of protein structural bioinformatics: the identification of interaction sites and the prediction of protein-protein interactions. On both tasks, we achieve state-of-the-art performance with much faster run times and fewer parameters than previous models. These results will considerably ease the deployment of deep learning methods in protein science and open the door for end-to-end differentiable approaches in protein modeling tasks such as function prediction and design.

bioinformatics↗

Wide occurrence of putative mobilized colistin resistance genes in the human gut microbiome

BackgroundThe high incidence of bacterial genes that confer resistance to last-resort antibiotics, such as colistin caused by MCR genes, poses an unprecedented threat to our civilizations health. To understand the spread, evolution, and distribution of such genes among human populations, with the final goal of diminishing their occurrence in human environments should be a priority. To tackle this problem, we investigated the distribution and prevalence of potential mcr genes in the human gut microbiome we used a set of bioinformatics tools to screen the Unified Human Gastrointestinal Genome (UHGG) collection for the presence, synteny and phylogeny of putative mcr genes, and co-located antibiotic resistance genes. ResultsA total of 2,079 ARGs were classified as different MCR in 2,046 Metagenome assembled genomes (MAGs), present in 1,596 individuals from 41 countries, of which 215 MCRs were identified in plasmidial contigs. The genera that presented the largest number of MCR-like genes were Suterella and Parasuterella, prevalent human gut bacteria of which Suterella wadsworthensis is associated with autism. Other potential pathogens carrying MCR genes belonged to the genus Vibrio, Escherichia and Campylobacter. Finally, we identified a total of 22,746 ARGs belonging to 21 different classes in the same 2,046 MAGs, suggesting multi-resistance potential in the corresponding bacterial strains, increasing the concern of ARGs impact in the clinical settings. ConclusionThis study uncovers the diversity of MCR-like genes in the human gut microbiome. We showed the cosmopolitan distribution of these genes in individuals worldwide and the co-presence of other antibiotic resistance genes, including Extended-spectrum beta-lactamases (ESBL). Also, we described mcr-like genes fused to a PAP2-like domain in S. wadsworthensis. Although these novel sequences increase our knowledge about the diversity and evolution of mcr-like genes, their activity and a potential colistin resistance in the corresponding strains has to be experimentally validated.

bioinformatics↗

Identifying accurate metagenome and amplicon software via a meta-analysis of benchmarking studies

Environmental DNA sequencing has rapidly become a widely-used technique for investigating a range of questions, particularly related to health and environmental monitoring. There has also been a proliferation of bioinformatic tools for analysing metagenomic and amplicon datasets, which makes selecting adequate tools a significant challenge. A number of benchmark studies have been undertaken; however, these can present conflicting results. We have applied a robust Z-score ranking procedure and a network meta-analysis method to identify software tools that are generally accurate for mapping DNA sequences to taxonomic hierarchies. Based upon these results we have identified some tools and computational strategies that produce robust predictions.

bioinformatics↗

Predicting chemotherapy response using a variational autoencoder approach

MotivationMultiple studies have shown the utility of transcriptome-wide RNA-seq profiles as features for machine learning-based prediction of response to chemotherapy in cancer. While tumor transcriptome profiles are publicly available for thousands of tumors for many cancer types, a relatively modest number of tumor profiles are clinically annotated for response to chemotherapy. The paucity of labeled examples and high dimension of the feature data limit performance for predicting therapeutic response using fully-supervised classification methods. Recently, multiple studies have established the utility of a deep neural network approach, the variational autoencoder (VAE), for generating meaningful latent features from original data. Here, we report first study of a semi-supervised approach using VAE-encoded tumor transcriptome features and regularized gradient boosted decision trees (XGBoost) to predict chemotherapy drug response for five cancer types: colon adenocarcinoma, pancreatic adenocarcinoma, bladder carcinoma, sarcoma, and breast invasive carcinoma. ResultsWe found: (1) VAE-encoding of the tumor transcriptome preserves the cancer type identity of the tumor, suggesting preservation of biologically relevant information; and (2) as a feature-set for supervised classification to predict response-to-chemotherapy, the unsupervised VAE encoding of the tumors gene expression profile leads to better area under the receiver operating characteristic curve (AUROC) classification performance than either the original gene expression profile or the PCA principal components of the gene expression profile, in four out of five cancer types that we tested. Availabilitygithub.com/ATHED/VAE_for_chemotherapy_drug_response_prediction Contactramseyst@oregonstate.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Sample-wise unsupervised deconvolution of complex tissues

MotivationComplex biological tissues are often a heterogeneous mixture of several molecularly distinct cell or tissue subtypes. Both subtype compositions and expressions in individual samples can vary across different biological states or conditions. Computational deconvolution aims to dissect patterns of bulk gene expression data into subtype compositions and subtype-specific expressions. Typically, existing deconvolution methods can only estimate averaged subtype-specific expressions in a population, while detecting differential expressions or co-expression networks in particular subtypes requires unique subtype expression estimates in individual samples. Different from population-level deconvolution, however, individual-level deconvolution is mathematically an underdetermined problem because there are more variables than observations. ResultsWe report a sample-wise Convex Analysis of Mixtures (swCAM) method that can estimate subtype proportions and subtype-specific expressions in individual samples from bulk tissue transcriptomes. We extend our previous CAM framework to include a new term accounting for between-sample variations and formulate swCAM as a nuclear-norm and{ell} 2,1-norm regularized matrix factorization problem. We determine hyperparameter values using a cross-validation scheme with random entry exclusion and obtain a swCAM solution using an efficient alternating direction method of multipliers. The swCAM is implemented in open-source R scripts. Experimental results on realistic simulation data show that swCAM can accurately estimate subtype-specific expressions in individual samples and successfully extract co-expression networks in particular subtypes that are otherwise unobtainable using bulk expression data. Application of swCAM to bulk-tissue data of 320 samples from bipolar disorder patients and controls identified changes in cell proportions, expression and coexpression modules in patient neurons. Mitochondria related genes showed significant changes suggesting an important role of energy dysregulation in bipolar disorder. Availability and implementationThe R Scripts of swCAM is freely available at https://github.com/Lululuella/swCAM. A users guide and a vignette are provided. Contactyuewang@vt.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Learning association for single-cell transcriptomics by integrating profiling of gene expression and alternative polyadenylation

Single-cell RNA-sequencing (scRNA-seq) has enabled transcriptome-wide profiling of gene expressions in individual cells. A myriad of computational methods have been proposed to learn cell-cell similarities and/or cluster cells, however, high variability and dropout rate inherent in scRNA-seq confounds reliable quantification of cell-cell associations based on the gene expression profile alone. Lately bioinformatics studies have emerged to capture key transcriptome information on alternative polyadenylation (APA) from standard scRNA-seq and revealed APA dynamics among cell types, suggesting the possibility of discerning cell identities with the APA profile. Complementary information at both layers of APA isoforms and genes creates great potential to develop cost-efficient approaches to dissect cell types based on multiple modalities derived from existing scRNA-seq data without changing experimental technologies. We proposed a toolkit called scLAPA for learning association for single-cell transcriptomics by combing single-cell profiling of gene expression and alternative polyadenylation derived from the same scRNA-seq data. We compared scLAPA with seven similarity metrics and five clustering methods using diverse scRNA-seq datasets. Comparative results showed that scLAPA is more effective and robust for learning cell-cell similarities and clustering cell types than competing methods. Moreover, with scLAPA we found two hidden subpopulations of peripheral blood mononuclear cells that were undetectable using the gene expression data alone. As a comprehensive toolkit, scLAPA provides a unique strategy to learn cell-cell associations, improve cell type clustering and discover novel cell types by augmentation of gene expression profiles with polyadenylation information, which can be incorporated in most existing scRNA-seq pipelines. scLAPA is available at https://github.com/BMILAB/scLAPA.

bioinformatics↗

Auto-CORPus: Automated and Consistent Outputs from Research Publications

To analyse large corpora using machine learning and other Natural Language Processing (NLP) algorithms, the corpora need to be standardised. The BioC format is a community-driven simple data structure for sharing text and annotations, however there is limited access to biomedical literature in BioC format and a lack of bioinformatics tools to convert online publication HTML formats to BioC. We present Auto-CORPus (Automated pipeline for Consistent Outputs from Research Publications), a novel NLP tool for the standardisation and conversion of publication HTML and table image files to three convenient machine-interpretable outputs to support biomedical text analytics. Firstly, Auto-CORPus can be configured to convert HTML from various publication sources to BioC. To standardise the description of heterogenous publication sections, the Information Artifact Ontology is used to annotate each section within the BioC output. Secondly, Auto-CORPus transforms publication tables to a JSON format to store, exchange and annotate table data between text analytics systems. The BioC specification does not include a data structure for representing publication table data, so we present a JSON format for sharing table content and metadata. Inline tables within full-text HTML files and linked tables within separate HTML files are processed and converted to machine-interpretable table JSON format. Finally, Auto-CORPus extracts abbreviations declared within publication text and provides an abbreviations JSON output that relates an abbreviation with the full definition. This abbreviation collection supports text mining tasks such as named entity recognition by including abbreviations unique to individual publications that are not contained within standard bio-ontologies and dictionaries. AvailabilityThe Auto-CORPus package is freely available with detailed instructions from Github at https://github.com/omicsNLP/Auto-CORPus/.

bioinformatics↗

Fungal Ice2p has remote homology to SERINCs, restriction factors for HIV and other viruses

Ice2p is an integral endoplasmic reticulum (ER) membrane protein in budding yeast S. cerevisiae named ICE because it is required for Inheritance of Cortical ER. Ice2p has also been reported to be involved in an ER metabolic branch-point that regulates the flux of lipid either to be stored in lipid droplets or to be used as membrane components. Alternately, Ice2p has been proposed to act as a tether that physically bridges the ER at contact sites with both lipid droplets and the plasma membrane via a long loop on the proteins cytoplasmic face that contains multiple predicted amphipathic helices. Here we carried out a bioinformatic analysis to increase understanding of Ice2p. Firstly, regarding topology, we found that diverse members of the fungal Ice2 family have ten transmembrane helices, which places the long loop on the exofacial face of Ice2p, where it cannot form inter-organelle bridges. Secondly, we identified Ice2 as a full-length homologue of SERINC (serine incorporator), a family of proteins with ten transmembrane helices found universally in eukaryotes. Since SERINCs are potent restriction factors for HIV and other viruses, study of Ice2p may reveal functions or mechanisms that shed light on viral restriction by SERINCs.

bioinformatics↗