bioRxiv Science⌕ Search

Biology subjects

Schork, K.

Publications and source records attributed to Schork, K..

3 recordsLinked to original sources

GoMi - A new gold standard corpus for miRNA Named Entity Recognition to test dictionary, rule-based and machine-learning approaches.

Biomarkers have been the focus of research for more than 30 years [REF1]. Paone et al. were among the first scientists to use the term biomarker in the course of a comparative study dealing with breast carcinoma [REF2]. In recent years, in addition to proteins and genes, miRNA or micro RNAs, which play an essential role in gene expression, have gained increased interest as valuable biomarkers. As a result, more and more information on miRNA biomarkers can be extracted via text mining approaches from the increasing amount of scientific literature. In the late 1990s the recognition of specific terms in biomedical texts has become a focus of bioinformatic research to automatically extract knowledge out of the increasing number of publications. For this, amongst other methods, machine learning algorithms are applied. However, the recognition (classification) capability of terms by machine learning or rule based algorithms depends on their correct and reproducible training and development. In the case of machine learning-based algorithms the quality of the available training and test data is crucial. The algorithms have to be tested and trained with curated and trustable data sets, the so-called gold or silver standards. Gold standards are text corpora, which are annotated by expertes, whereby silver standards are curated automatically by other algorithms. Training and calibration of neural networks is based on such corpora. In the literature there are some silver standards with approx. 500,000 tokens [REF3]. Also there are already published gold standards for species, genes, proteins or diseases. However, there is no corpus that has been generated specifically for miRNA. To close this gap, we have generated GoMi, a novel and manually curated gold standard corpus for miRNA. GoMi can be directly used to train ML-methods to calibrate or test different algorithms based on the rule-based approach or dictionary-based approach. The GoMi gold standard corpus was created using publicly available PubMed abstracts. GoMi can be downloaded here: https://github.com/mpc-bioinformatics/mirnaGS---GoMi.

bioinformatics↗

Characterization of peptide-protein relationships in protein ambiguity groups via bipartite graphs

In bottom-up proteomics, proteins are enzymatically digested into peptides before measurement with mass spectrometry. The relationship between proteins and their corresponding peptides can be represented by bipartite graphs. We conduct a comprehensive analysis of bipartite graphs using quantified peptides from measured data sets as well as theoretical peptides from an in silico digestion of the corresponding complete taxonomic protein sequence databases. The aim of this study is to characterize and structure the different types of graphs that occur and to compare them between data sets. We observed a large influence of the accepted minimum peptide length during in silico digestion. When changing from theoretical peptides to measured ones, the graph structures are subject to two opposite effects. On the one hand, the graphs based on measured peptides are on average smaller and less complex compared to graphs using theoretical peptides. On the other hand, the proportion of protein nodes without unique peptides, which are a complicated case for protein inference and quantification, is considerably larger for measured data. Additionally, the proportion of graphs containing at least one protein node without unique peptides rises when going from database to quantitative level. The fraction of shared peptides and proteins without unique peptides as well as the complexity and size of the graphs highly depends on the data set and organism. Large differences between the structures of bipartite peptide-protein graphs have been observed between database and quantitative level as well as between analyzed species. In the analyzed measured data sets, the proportion of protein nodes without unique peptides ranged from 6.4% to 55.0%. This highlights the need for novel methods that can quantify proteins without unique peptides. The knowledge about the structure of the bipartite peptide-protein graphs gained in this study will be useful for the development of such algorithms.

bioinformatics↗

Differential interferon-α subtype immune signatures suppress SARS-CoV-2 infection

Type I interferons (IFN-I) exert pleiotropic biological effects during viral infections, balancing virus control versus immune-mediated pathologies and have been successfully employed for the treatment of viral diseases. Humans express twelve IFN-alpha () subtypes, which activate downstream signalling cascades and result in distinct patterns of immune responses and differential antiviral responses. Inborn errors in type I IFN immunity and the presence of anti-IFN autoantibodies account for very severe courses of COVID-19, therefore, early administration of type I IFNs may be protective against life-threatening disease. Here we comprehensively analysed the antiviral activity of all IFN subtypes against SARS-CoV-2 to identify the underlying immune signatures and explore their therapeutic potential. Prophylaxis of primary human airway epithelial cells (hAEC) with different IFN subtypes during SARS-CoV-2 infection uncovered distinct functional classes with high, intermediate and low antiviral IFNs. In particular IFN5 showed superior antiviral activity against SARS-CoV-2 infection. Dose-dependency studies further displayed additive effects upon co-administered with the broad antiviral drug remdesivir in cell culture. Transcriptomics of IFN-treated hAEC revealed different transcriptional signatures, uncovering distinct, intersecting and prototypical genes of individual IFN subtypes. Global proteomic analyses systematically assessed the abundance of specific antiviral key effector molecules which are involved in type I IFN signalling pathways, negative regulation of viral processes and immune effector processes for the potent antiviral IFN5. Taken together, our data provide a systemic, multi-modular definition of antiviral host responses mediated by defined type I IFNs. This knowledge shall support the development of novel therapeutic approaches against SARS-CoV-2.

cell biology↗