bioRxiv ScienceSearch

Biology subjects

Nielsen, M.

Publications and source records attributed to Nielsen, M..

8 recordsLinked to original sources

NetTCR: sequence-based prediction of TCR binding to peptide-MHC complexes using convolutional neural networks

Predicting epitopes recognized by cytotoxic T cells has been a long standing challenge within the field of immuno- and bioinformatics. While reliable predictions of peptide binding are available for most Major Histocompatibility Complex class I (MHCI) alleles, prediction models of T cell receptor (TCR) interactions with MHC class I-peptide complexes remain poor due to the limited amount of available training data. Recent next generation sequencing projects have however generated a considerable amount of data relating TCR sequences with their cognate HLA-peptide complex target. Here, we utilize such data to train a sequence-based predictor of the interaction between TCRs and peptides presented by the most common human MHCI allele, HLA-A*02:01. Our model is based on convolutional neural networks, which are especially designed to meet the challenges posed by the large length variations of TCRs. We show that such a sequence-based model allows for the identification of TCRs binding a given cognate peptide-MHC target out of a large pool of non-binding TCRs.

bioinformatics

NetSurfP-2.0: improved prediction of protein structural features by integrated deep learning

The ability to predict local structural features of a protein from the primary sequence is of paramount importance for unravelling its function in absence of experimental structural information. Two main factors affect the utility of potential prediction tools: their accuracy must enable extraction of reliable structural information on the proteins of interest, and their runtime must be low to keep pace with sequencing data being generated at a constantly increasing speed.\n\nHere, we present an updated and extended version of the NetSurfP tool (http://www.cbs.dtu.dk/services/NetSurfP-2.0/), that can predict the most important local structural features with unprecedented accuracy and runtime. NetSurfP-2.0 is sequence-based and uses an architecture composed of convolutional and long short-term memory neural networks trained on solved protein structures. Using a single integrated model, NetSurfP-2.0 predicts solvent accessibility, secondary structure, structural disorder, and backbone dihedral angles for each residue of the input sequences.\n\nWe assessed the accuracy of NetSurfP-2.0 on several independent test datasets and found it to consistently produce state-of-the-art predictions for each of its output features. We observe a correlation of 80% between predictions and experimental data for solvent accessibility, and a precision of 85% on secondary structure 3-class predictions. In addition to improved accuracy, the processing time has been optimized to allow predicting more than 1,000 proteins in less than 2 hours, and complete proteomes in less than 1 day.

bioinformatics

Footprints of antigen processing boost MHC class II natural ligand binding predictions

Major Histocompatibility complex class II (MHC-II) molecules present peptide fragments to T cells for immune recognition. Current predictors for peptide:MHC-II binding are trained on binding affinity data, generated in-vitro and therefore lacking information about antigen processing. For the first time, we here describe prediction models of peptide:MHC-II binding trained directly on naturally eluted peptides, and show that these, in addition to peptide binding to the MHC, incorporate identifiable rules of antigen processing. In fact, we observed detectable signals of protease cleavage at defined positions of the peptides. We also hypothesize a role of the length of the terminal ligand protrusions for trimming the peptide to the epitope presented. The results of integrating binding affinity and eluted ligand data in a combined model demonstrate improved performance for the prediction of MHC-II ligands, and foreshadow a new generation of improved peptide:MHC-II prediction tools of considerable importance for understanding and manipulating immune responses.

immunology

Passenger mutations in 2500 cancer genomes: Overall molecular functional impact and consequences

The Pan-cancer Analysis of Whole Genomes (PCAWG) project provides an unprecedented opportunity to comprehensively characterize a vast set of uniformly annotated coding and non-coding mutations present in thousands of cancer genomes. Classical models of cancer progression posit that only a small number of these mutations strongly drive tumor progression and that the remaining ones (termed \"putative passengers\") are inconsequential for tumorigenesis. In this study, we leveraged the comprehensive variant data from PCAWG to ascertain the molecular functional impact of each variant. The impact distribution of PCAWG mutations shows that, in addition to high- and low-impact mutations, there is a group of medium-impact putative passengers predicted to influence gene activity. Moreover, the predicted impact relates to the underlying mutational signature: different signatures confer divergent impact, differentially affecting distinct regulatory subsystems and gene categories. We also find that impact varies based on subclonal architecture (i.e., early vs. late mutations) and can be related to patient survival. Finally, we note that insufficient power due to limited cohort sizes precludes identification of weak drivers using standard recurrence-based approaches. To address this, we adapted an additive effects model derived from complex trait studies to show that aggregating the impact of putative passenger variants (i.e. including yet undetected weak drivers) provides significant predictability for cancer phenotypes beyond the PCAWG identified driver mutations (12.5% additive variance). Furthermore, this framework allowed us to estimate the frequency of potential weak driver mutations in the subset of PCAWG samples lacking well-characterized driver alterations.

genomics

Transcription-driven Chromatin Repression of Intragenic Promoters

Progression of RNA polymerase II (RNAPII) transcription relies on the appropriately positioned activities of elongation factors. The resulting profile of factors and chromatin signatures along transcription units provides a \"positional information system\" for transcribing RNAPII. Here, we investigate a chromatin-based mechanism that suppresses intragenic initiation of RNAPII transcription. We demonstrate that RNAPII transcription across gene promoters represses their function in plants. This repression is characterized by reduced promoter-specific molecular signatures and increased molecular signatures associated with RNAPII elongation. The FACT histone chaperone complex is required for this repression mechanism. Genome-wide mapping of Transcription Start Sites (TSSs) reveals thousands of discrete intragenic TSS positions in FACT mutants. Histone 3 lysine 4 mono-methylation poises exonic sites to initiate RNAPII transcription in FACT mutants. Uncovering the mechanism for intragenic TSS repression through the act of RNAPII elongation has important implications for understanding pervasive RNAPII transcription and the regulation of transcript isoform diversity.

genomics

Computational Tools for the Identification and Interpretation of Sequence Motifs in Immunopeptidomes

Recent advances in proteomics and mass-spectrometry have widely expanded the detectable peptide repertoire presented by major histocompatibility complex (MHC) molecules on the cell surface, collectively known as the immunopeptidome. Finely characterizing the immunopeptidome brings about important basic insights into the mechanisms of antigen presentation, but can also reveal promising targets for vaccine development and cancer immunotherapy. In this report, we describe a number of practical and efficient approaches to analyze immunopeptidomics data, discussing the identification of meaningful sequence motifs in various scenarios and considering current limitations. We address the issue of filtering false hits and contaminants, and the problem of motif deconvolution in cell lines expressing multiple MHC alleles, both for the MHC class I and class II systems. Finally, we demonstrate how machine learning can be readily employed by non-expert users to generate accurate prediction models directly from mass-spectrometry eluted ligand data sets.

bioinformatics

Improved prediction of Bovine Leucocyte Antigens (BoLA) presented ligands by use of MS eluted ligands and in-vitro binding data; impact for the identification T cell epitopes

Peptide binding to MHC class I molecules is the single most selective step in antigen presentation and the strongest single correlate to peptide cellular immunogenicity. The cost of experimentally characterizing the rules of peptide presentation for a given MHC-I molecule is extensive, and predictors of peptide-MHC interactions constitute an attractive alternative.\n\nRecently, an increasing amount of MHC presented peptides identified by mass spectrometry (MS ligands) has been published. Handling and interpretation of MS ligand data is in general challenging due to the poly-specificity nature of the data. We here outline a general pipeline for dealing with this challenge, and accurately annotate ligands to the relevant MHC-I molecule they were eluted from by use of GibbsClustering and binding motif information inferred from in-silico models. We illustrate the approach here in the context of MHCI molecules (BoLA) of cattle. Next, we demonstrate how such annotated BoLA MS ligand data can readily be integrated with in-vitro binding affinity data in a prediction model with very high and unprecedented performance for identification of BoLA-I restricted T cell epitopes.\n\nThe approach has here been applied to the BoLA-I system, but the pipeline is readily applicable to MHC systems in other species.

immunology

NetMHCpan 4.0: Improved peptide-MHC class I interaction predictions integrating eluted ligand and peptide binding affinity data

Cytotoxic T cells are of central importance in the immune systems response to disease. They recognize defective cells by binding to peptides presented on the cell surface by MHC (major histocompatibility complex) class I molecules. Peptide binding to MHC molecules is the single most selective step in the antigen presentation pathway. On the quest for T cell epitopes, the prediction of peptide binding to MHC molecules has therefore attracted large attention.\n\nIn the past, predictors of peptide-MHC interaction have in most cases been trained on binding affinity data. Recently an increasing amount of MHC presented peptides identified by mass spectrometry has been published containing information about peptide processing steps in the presentation pathway and the length distribution of naturally presented peptides. Here, we present NetMHCpan-4.0, a method trained on both binding affinity and eluted ligand data leveraging the information from both data types. Large-scale benchmarking of the method demonstrates an increased predictive performance compared to state-of-the-art when it comes to identification of naturally processed ligands, cancer neoantigens, and T cell epitopes.

bioinformatics