bioRxiv Science⌕ Search

Biology subjects

Shields, K.

Publications and source records attributed to Shields, K..

3 recordsLinked to original sources

Characterization of the immunoglobulin lambda chain locus from diverse populations reveals extensive genetic variation

Immunoglobulins (IGs), crucial components of the adaptive immune system, are encoded by three genomic loci. However, the complexity of the IG loci severely limits the effective use of short read sequencing, limiting our knowledge of population diversity in these loci. We leveraged existing long read whole-genome sequencing (WGS) data, fosmid technology, and IG targeted single-molecule, real-time (SMRT) long-read sequencing (IG-Cap) to create haplotype-resolved assemblies of the IG Lambda (IGL) locus from 6 ethnically diverse individuals. In addition, we generated 10 diploid assemblies of IGL from a diverse cohort of individuals utilizing IG-cap. From these 16 individuals, we identified significant allelic diversity, including 37 novel IGLV alleles. In addition, we observed highly elevated single nucleotide variation (SNV) in IGLV genes relative to IGL intergenic and genomic background SNV density. By comparing SNV calls between our high quality assemblies and existing short read datasets from the same individuals, we show a high propensity for false-positives in the short read datasets. Finally, for the first time, we nucleotide-resolved common 5-10 Kb duplications in the IGLC region that contain functional IGLJ and IGLC genes. Together these data represent a significant advancement in our understanding of genetic variation and population diversity in the IGL locus.

genetics↗

Antibody repertoire gene usage is explained by common genetic variants in the immunoglobulin heavy chain locus

Variation in the antibody response has been linked to differential outcomes in disease, and suboptimal vaccine and therapeutic responsiveness, the determinants of which have not been fully elucidated. Countering models that presume antibodies are generated largely by stochastic processes, we demonstrate that polymorphisms within the immunoglobulin heavy chain locus (IGH) significantly impact the naive and antigen-experienced antibody repertoire, indicating that genetics predisposes individuals to mount qualitatively and quantitatively different antibody responses. We pair recently developed long-read genomic sequencing methods with antibody repertoire profiling to comprehensively resolve IGH genetic variation, including novel structural variants, single nucleotide variants, and genes and alleles. We show that IGH germline variants determine the presence and frequency of antibody genes in the expressed repertoire, including those enriched in functional elements linked to V(D)J recombination, and overlapping disease-associated variants. These results illuminate the power of leveraging IGH genetics to better understand the regulation, function and dynamics of the antibody response in disease.

genomics↗

Targeted long-read sequencing facilitates phased diploid assembly and genotyping of the human T cell receptor alpha, delta and beta loci

T cell receptors (TCRs) recognize peptide fragments presented by the major histocompatibility complex (MHC) and are critical to T cell mediated immunity. Early studies demonstrated an enrichment of polymorphisms within TCR-encoding (TR) gene loci. However, more recent data indicate that variation in these loci are underexplored, limiting understanding of the impact of TR polymorphism on TCR function in disease, even though: (i) TCR repertoire signatures are heritable and (ii) associate with disease phenotypes. TR variant discovery and curation has been difficult using standard high-throughput methods. To address this, we expanded our published targeted long-read sequencing approach to generate highly accurate haplotype resolved assemblies of the human TR beta (TRB) and alpha/delta (TRA/D) loci, facilitating the detection and genotyping of single nucleotide polymorphisms (SNPs), insertion-deletions (indels), structural variants (SVs) and TR genes. We validate our approach using two mother-father-child trios and 5 unrelated donors representing multiple populations. Comparisons of long-read derived variants to short-read datasets revealed improved genotyping accuracy, and TR gene annotation led to the discovery of 79 previously undocumented V, D, and J alleles. This demonstrates the utility of this framework to resolve the TR loci, and ultimately our understanding of TCR function in disease.

genomics↗