bioRxiv ScienceSearch

Biology subjects

Peters, B.

Publications and source records attributed to Peters, B..

8 recordsLinked to original sources

NetTCR: sequence-based prediction of TCR binding to peptide-MHC complexes using convolutional neural networks

Predicting epitopes recognized by cytotoxic T cells has been a long standing challenge within the field of immuno- and bioinformatics. While reliable predictions of peptide binding are available for most Major Histocompatibility Complex class I (MHCI) alleles, prediction models of T cell receptor (TCR) interactions with MHC class I-peptide complexes remain poor due to the limited amount of available training data. Recent next generation sequencing projects have however generated a considerable amount of data relating TCR sequences with their cognate HLA-peptide complex target. Here, we utilize such data to train a sequence-based predictor of the interaction between TCRs and peptides presented by the most common human MHCI allele, HLA-A*02:01. Our model is based on convolutional neural networks, which are especially designed to meet the challenges posed by the large length variations of TCRs. We show that such a sequence-based model allows for the identification of TCRs binding a given cognate peptide-MHC target out of a large pool of non-binding TCRs.

bioinformatics

Gene Regulatory Cross Networks: Inferring Gene Level Cell-to-Cell Communications of Immune Cells

BackgroundGene level cell-to-cell communications are crucial part of biology as they may be potential targets of drugs and vaccines against a disease condition of interest. Yet, there are only few studies that propose algorithms on this particularly important research field.\n\nResultsIn this study, we first overview the current literature and define two general terms for the types of approaches in general for gene level cell-to-cell communications: Gene Regulatory Cross Networks (GRCN) and Gene Co-Expression Cross Networks (GCCN). We then propose two algorithms for each type, named as GRCNone and GCCNone. We applied them to reveal communications among 8 different immune cell types and evaluate their performances mainly via membrane protein database. Also, we show the biological relevance of the predicted cross-networks with pathway enrichment analysis. We then provide an approach that prioritize the targets by ranking them before experimental validations.\n\nConclusionsWe establish two main approaches and propose algorithms for genome-wide scale gene level cell-to-cell communications between any two different cell-types. This study aims accelerating this relatively new avenue of research in cross-networks and points out the gap of it with the well-established single cell type gene networks. The proposed algorithms have the potential to reveal gene level interactions between normal and disease cell types. For instance, they might reveal the interaction of genes between tumor and normal cells, which are the potential drug-targets and thus can help finding new cures that might prevent the prevailing of tumor cells.

bioinformatics

No cell is an island: circulating T cell:monocyte complexes are markers of immune perturbations

Our results highlight for the first time that a significant proportion of cell doublets in flow cytometry, previously believed to be the result of technical artefacts and thus ignored in data acquisition and analysis, are the result of true biological interaction between immune cells. In particular, we show that cell:cell doublets pairing a T cell and a monocyte can be directly isolated from human blood, and high resolution microscopy shows polarized distribution of LFA1/ICAM1 in many doublets, suggesting in vivo formation. Intriguingly, T cell:monocyte complex frequency and phenotype fluctuate with the onset of immune perturbations such as infection or immunization, reflecting expected polarization of immune responses. Overall these data suggest that cell doublets reflecting T cell-monocyte in vivo immune interactions can be detected in human blood and that the common approach in flow cytometry to avoid studying cell:cell complexes should be revisited.

immunology

Attention samples objects held in working memory at a theta rhythm

Attention selects relevant information regardless of whether it is physically present or internally stored in working memory. Perceptual research has shown that attentional selection of external information is better conceived as rhythmic prioritization than as stable allocation. Here we tested this principle using information processing of internal representations held in working memory. Participants memorized four spatial positions that formed the endpoints of two objects. One of the positions was cued for a delayed match-non-match test. When uncued positions were probed, participants responded faster to uncued positions located on the same object as the cued position than to those located on the other object, revealing object-based attention in working memory. Manipulating the interval between cue and probe at a high temporal resolution revealed that reaction times oscillated at a theta rhythm of 6 Hz. Moreover, oscillations showed an anti-phase relationship between memorized but uncued positions on the same versus other object as the cued position, suggesting that attentional prioritization fluctuated rhythmically in an object-based manner. Our results demonstrate the highly rhythmic nature of attentional selection in working memory. Moreover, the striking similarity between rhythmic attentional selection of mental representations and perceptual information suggests that attentional oscillations are a general mechanism of information processing in human cognition. These findings have important implications for current, attention-based models of working memory.

neuroscience

3’ Branch Ligation: A Novel Method to Ligate Non-Complementary DNA to Recessed or Internal 3’OH Ends in DNA or RNA

Nucleic acid ligases are crucial enzymes that repair breaks in DNA or RNA during synthesis, repair and recombination. Various molecular tools have been developed using the diverse activities of DNA/RNA ligases. Herein, we demonstrate a non-conventional ability of T4 DNA ligase to join 5 phosphorylated blunt-end double-stranded DNA to DNA breaks at 3 recessive ends, gaps, or nicks to form a 3 branch structure. Therefore, this base pairing-independent ligation is termed 3 branch ligation (3BL). In an extensive study of optimal ligation conditions, similar to blunt-end ligation, the presence of 10% PEG-8000 in the ligation buffer significantly increased ligation efficiency. A low level of nucleotide preference was observed at the junction sites using different synthetic DNAs. Furthermore, we discovered that T4 DNA ligase efficiently ligated DNA to the 3 recessed end of RNA, not to that of DNA, in a DNA/RNA hybrid, whereas RNA ligases are less efficient in this reaction. These novel properties of T4 DNA ligase can be utilized as a broad molecular technique in many important applications. We performed a proof-of-concept study of a new directional tagmentation protocol for next generation sequencing (NGS) library construction that eliminates inverted adapters and allows sample barcode insertion adjacent to genomic DNA. 3BL after single transposon tagmentation can theoretically achieve 100% usable template, and our empirical data demonstrate that the new approach produced higher yield compared with traditional double transposon or Y transposon tagmentation. We further explore the potential use of 3BL for preparing targeted RNA NGS libraries with mitigated structure-based bias and adapter dimer problems.

molecular biology

Single tube bead-based DNA co-barcoding for cost effective and accurate sequencing, haplotyping, and assembly

Obtaining accurate sequences from long DNA molecules is very important for genome assembly and other applications. Here we describe single tube long fragment read (stLFR), a technology that enables this a low cost. It is based on adding the same barcode sequence to sub-fragments of the original long DNA molecule (DNA co-barcoding). To achieve this efficiently, stLFR uses the surface of microbeads to create millions of miniaturized barcoding reactions in a single tube. Using a combinatorial process up to 3.6 billion unique barcode sequences were generated on beads, enabling practically non-redundant co-barcoding with 50 million barcodes per sample. Using stLFR, we demonstrate efficient unique co-barcoding of over 8 million 20-300 kb genomic DNA fragments. Analysis of the genome of the human genome NA12878 with stLFR demonstrated high quality variant calling and phasing into contigs up to N50 34 Mb. We also demonstrate detection of complex structural variants and complete diploid de novo assembly of NA12878. These analyses were all performed using single stLFR libraries and their construction did not significantly add to the time or cost of whole genome sequencing (WGS) library preparation. stLFR represents an easily automatable solution that enables high quality sequencing, phasing, SV detection, scaffolding, cost-effective diploid de novo genome assembly, and other long DNA sequencing applications.

genomics

DAFi: A Directed Recursive Filtering and Clustering Approach to Data-Driven Identification of Cell Populations from Polychromatic Flow Cytometry Data

Computational methods for identification of cell populations from high-dimensional flow cytometry data are changing the paradigm of cytometry bioinformatics. Data clustering is the most common computational approach to unsupervised identification of cell populations from multidimensional cytometry data. We found that combining recursive filtering and clustering with constraints converted from the user manual gating strategy can effectively identify overlapping and rare cell populations from smeared data that would have been difficult to resolve by either a single run of data clustering or manual segregation. We named this new method DAFi: Directed Automated Filtering and Identification of cell populations. Design of DAFi preserves the data-driven characteristics of unsupervised clustering for identifying novel cell-based biomarkers, but also makes the results interpretable to experimental scientists as in supervised classification through mapping and merging the high-dimensional data clusters into the user-defined 2D gating hierarchy. By recursive data filtering before clustering, DAFi can uncover small local clusters which are otherwise difficult to identify due to the statistical interference of the irrelevant major clusters. Quantitative assessment of cell type specific characteristics demonstrates that the population proportions calculated by DAFi, while being highly consistent with those by expert centralized manual gating, have smaller technical variance than those from individual manual gating analysis. Visual examination of the dot plots showed that the boundaries of the DAFi-identified cell populations followed the natural shapes of the data distributions. To further exemplify the utility of DAFi, we show that DAFi can incorporate the FLOCK clustering method to identify novel cell-based biomarkers. Implementation of DAFi supports options including clustering, bisecting, slope-based gating, and reversed filtering to meet various auto-gating needs from different scientific use cases.

bioinformatics

NetMHCpan 4.0: Improved peptide-MHC class I interaction predictions integrating eluted ligand and peptide binding affinity data

Cytotoxic T cells are of central importance in the immune systems response to disease. They recognize defective cells by binding to peptides presented on the cell surface by MHC (major histocompatibility complex) class I molecules. Peptide binding to MHC molecules is the single most selective step in the antigen presentation pathway. On the quest for T cell epitopes, the prediction of peptide binding to MHC molecules has therefore attracted large attention.\n\nIn the past, predictors of peptide-MHC interaction have in most cases been trained on binding affinity data. Recently an increasing amount of MHC presented peptides identified by mass spectrometry has been published containing information about peptide processing steps in the presentation pathway and the length distribution of naturally presented peptides. Here, we present NetMHCpan-4.0, a method trained on both binding affinity and eluted ligand data leveraging the information from both data types. Large-scale benchmarking of the method demonstrates an increased predictive performance compared to state-of-the-art when it comes to identification of naturally processed ligands, cancer neoantigens, and T cell epitopes.

bioinformatics