bioRxiv Science⌕ Search

Biology subjects

Chudakov, D.

Publications and source records attributed to Chudakov, D..

2 recordsLinked to original sources

Comprehensive analysis of αβT-cell receptor repertoires reveals signatures of thymic selection

Thymic selection is a multi-stage process that establishes a T-cell immunity that is efficient in fighting diverse foreign pathogens, while not being self-reactive. During the process, T-cell receptor (TCR) and {beta} chains are rearranged to form a highly diverse set of heterodimers that are selected based on their affinity to peptides presented by major histocompatibility molecules (MHC) in the thymus. Here we employ high-throughput TCR sequencing data, theoretical model of TCR rearrangement and dedicated statistical methods to infer how the selection process affects human TCR repertoire on different scales. On a global scale, our results indicate differences in V(D)J gene usage, complementarity determining region 3 (CDR3) amino acid composition, CDR3 physicochemical properties, k-mer composition and differences in loop structure induced by the selection process. On a local scale, we were able to determine enriched TCR motifs and "holes" in repertoire induced by positive and negative selection and characterize their features. Finally, we demonstrated how TCR sequence composition affects lineage commitment via thymic selection and highlighted the effect of individual MHC haplotype. Our results can aid in identification of potentially self-reactive TCRs in donor repertoires in autoimmunity and immunotherapy studies. Graphical abstract (Figure 0)Study overview. A. Datasets used in the study and the way TCR sequences were analyzed. B. Main directions in which thymic selection shapes TCR repertoire. Selection results in TCR-CDR3 losing positively charged and large amino-acids while increasing its flexibility. Unlike CDR3 of the TCR chain, CDR3{beta} hydrophobicity is also increased by selection. CDR3s carrying Cysteines and glycosylation sites are unlikely to pass the selection. In contrast, CDR3 carrying poly-Glycine regions are more likely to be selected for both chains. T-cells committed to CD8+ lineage were more likely to feature bulged CDR3s compared to CD4+. C. CDR3s enriched after thymic selection guide lineage commitment according to single-cell RNA sequencing data analysis, e.g. the MAIT cells and CD8+ phenotypes. D. Enrichment and depletion of certain TCR motifs pinpoint fine-structure of post-selection repertoires, as inferred by sequence neighborhood analysis. E. Comparative analysis of enriched and depleted TCR clusters in monozygotic twins revealed that the selection process is shaped by HLA haplotype. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=195 SRC="FIGDIR/small/635277v2_ufig1.gif" ALT="Figure 1"> View larger version (44K): org.highwire.dtl.DTLVardef@148f9f1org.highwire.dtl.DTLVardef@117858org.highwire.dtl.DTLVardef@f3f837org.highwire.dtl.DTLVardef@12cf2c1_HPS_FORMAT_FIGEXP M_FIG C_FIG

immunology↗

Ultrasensitive allele inference from immune repertoire sequencing data with MiXCR

Allelic variability in the adaptive immune receptor loci, which harbor the gene segments that encode B cell and T cell receptors (BCR/TCR), has been shown to be of critical importance for immune responses to pathogens and vaccines. In recent years, B cell and T cell receptor repertoire sequencing (Rep-Seq) has become widespread in immunology research making it the most readily available source of information about allelic diversity in immunoglobulin (IG) and T cell receptor (TR) loci in different populations. Here we present a novel algorithm for extra-sensitive and specific variable (V) and joining (J) gene allele inference and genotyping allowing reconstruction of individual high-quality gene segment libraries. The approach can be applied for inferring allelic variants from peripheral blood lymphocyte BCR and TCR repertoire sequencing data, including hypermutated isotype-switched BCR sequences, thus allowing high-throughput genotyping and novel allele discovery from a wide variety of existing datasets. The developed algorithm is a part of the MiXCR software (https://mixcr.com) and can be incorporated into any pipeline utilizing upstream processing with MiXCR. We demonstrate the accuracy of this approach using Rep-Seq paired with long-read genomic sequencing data, comparing it to a widely used algorithm, TIgGER. We applied the algorithm to a large set of IG heavy chain (IGH) Rep-Seq data from 450 donors of ancestrally diverse population groups, and to the largest reported full-length TCR alpha and beta chain (TRA; TRB) Rep-Seq dataset, representing 134 individuals. This allowed us to assess the genetic diversity of genes within the IGH, TRA and TRB loci in different populations and demonstrate the connection between antibody repertoire gene usage and the number of allelic variants present in the population. Finally we established a database of allelic variants of V and J genes inferred from Rep-Seq data and their population frequencies with free public access at https://vdj.online.

immunology↗