bioRxiv ScienceSearch

Biology subjects

Kleinstein, S. H.

Publications and source records attributed to Kleinstein, S. H..

3 recordsLinked to original sources

Identification of subject-specific immunoglobulin alleles from expressed repertoire sequencing data

The adaptive immune receptor repertoire (AIRR) contains information on an individuals immune past, present and potential in the form of the evolving sequences that encode the B cell receptor (BCR) repertoire. AIRR sequencing (AIRR-seq) studies rely on databases of known BCR germline variable (V), diversity (D) and joining (J) genes to detect somatic mutations in AIRR-seq data via comparison to the best-aligning database alleles. However, it has been shown that these databases are far from complete, leading to systematic misidentification of mutated positions in subsets of sample sequences. We previously presented TIgGER, a computational method to identify subject-specific V gene genotypes, including the presence of novel V gene alleles, directly from AIRR-seq data. However, the original algorithm was unable to detect alleles that differed by more than 5 single nucleotide polymorphisms (SNPs) from a database allele. Here we present and apply an improved version of the TIgGER algorithm which can detect alleles that differ by any number of SNPs from the nearest database allele, and can construct subject-specific genotypes with minimal prior information. TIgGER predictions are validated both computationally (using a leave-one-out strategy) and experimentally (using genomic sequencing), resulting in the addition of three new immunoglobulin heavy chain V (IGHV) gene alleles to the IMGT repertoire. Finally, we develop a Bayesian strategy to provide a confidence estimate associated with genotype calls. All together, these methods allow for much higher accuracy in germline allele assignment, an essential step in AIRR-seq studies.

bioinformatics

Performance-optimized partitioning of clonotypes from high-throughput immunoglobulin repertoire sequencing data

MotivationDuring adaptive immune responses, activated B cells expand and undergo somatic hypermutation of their immunoglobulin (Ig) receptor, forming a clone of diversified cells that can be related back to a common ancestor. Identification of B cell clonotypes from high-throughput Adaptive Immune Receptor Repertoire sequencing (AIRR-seq) data relies on computational analysis. Recently, we proposed an automate method to partition sequences into clonal groups based on single-linkage clustering of the Ig receptor junction region with length-normalized hamming distance metric. This method could identify clonally-related sequences with high confidence on several benchmark experimental and simulated data sets. However, this approach was computationally expensive, and unable to provide estimates of accuracy for new data. Here, a new method is presented that address this computational bottleneck and also provides a study-specific estimation of performance, including sensitivity and specificity. The method uses a finite mixture modeling fitting procedure for learning the parameters of two univariate curves which fit the bimodal distributions of the distance vector between pairs of sequences. These distribution are used to estimate the performance of different threshold choices for partitioning sequences into clonotypes. These performance estimates are validated using simulated and experimental datasets. With this method, clonotypes can be identified from AIRR-seq data with sensitivity and specificity profiles that are user-defined based on the overall goals of the study.\n\nAvailabilitySource code is freely available at the Immcantation Portal: www.immcantation.com under the CC BY-SA 4.0 license.\n\nContactsteven.kleinstein@yale.edu

immunology

Tumor-infiltrating immune repertoires captured by single-cell barcoding in emulsion

Tumor-infiltrating lymphocytes (TILs) are critical to anti-cancer immune responses, but their diverse phenotypes and functions remain poorly understood and challenging to study. We therefore developed a single-cell barcoding technology for deep characterization of TILs without the need for cell-sorting or culture. Our emulsion-based method captures full-length, natively paired B-cell and T-cell receptor (BCR and TCR) sequences from lymphocytes among millions of input cells. We validated the method with 3 million B-cells from healthy human blood and 350,000 B-cells from an HIV elite controller, before processing 400,000 cells from an unsorted dissociated ovarian adenocarcinoma and recovering paired BCRs and TCRs from over 11,000 TILs. We then extended the barcoding method to detect DNA-labeled antibodies, allowing ultra-high throughput, simultaneous protein detection and RNA sequencing from single cells.

immunology