bioRxiv ScienceSearch

Biology subjects

Lange, V.

Publications and source records attributed to Lange, V..

3 recordsLinked to original sources

High-throughput Interpretation of Killer-cell Immunoglobulin-like Receptor Short-read Sequencing Data with PING

The killer-cell immunoglobulin-like receptor (KIR) complex on chromosome 19 encodes receptors that modulate the activity of natural killer cells, and variation in these genes has been linked to infectious and autoimmune disease, as well as having bearing on pregnancy and transplant outcomes. The medical relevance and high variability of KIR genes makes short-read sequencing an attractive technology for interrogating the region, providing a high-throughput, high-fidelity sequencing method that is cost-effective. However, because this gene complex is characterized by extensive nucleotide polymorphism, structural variation including gene fusions and deletions, and a high level of homology between genes, its interrogation at high resolution has been thwarted by bioinformatic challenges, with most studies limited to examining presence or absence of specific genes. Here, we present the PING (Pushing Immunogenetics to the Next Generation) pipeline, which incorporates empirical data, novel alignment strategies and a custom alignment processing workflow to enable high-throughput KIR sequence analysis from short-read data. PING provides KIR gene copy number classification functionality for all KIR genes through use of a comprehensive alignment reference. The gene copy number determined per individual enables an innovative genotype determination workflow using genotype-matched references. Together, these methods address the challenges imposed by the structural complexity and overall homology of the KIR complex. To determine copy number and genotype determination accuracy, we applied PING to European and African validation cohorts and a synthetic dataset. PING demonstrated exceptional copy number determination performance across all datasets and robust genotype determination performance. Finally, an investigation into discordant genotypes for the synthetic dataset provides insight into misaligned reads, advancing our understanding in interpretation of short-read sequencing data in complex genomic regions. PING promises to support a new era of studies of KIR polymorphism, delivering high-resolution KIR genotypes that are highly accurate, enabling high-quality, high-throughput KIR genotyping for disease and population studies. Author summaryKiller cell immunoglobulin-like receptors (KIR) serve a critical role in regulating natural killer cell function. They are encoded by highly polymorphic genes within a complex genomic region that has proven difficult to interrogate owing to structural variation and extensive sequence homology. While methods for sequencing KIR genes have matured, there is a lack of bioinformatic support to accurately interpret KIR short-read sequencing data. The extensive structural variation of KIR, both the small-scale nucleotide insertions and deletions and the large-scale gene duplications and deletions, coupled with the extensive sequence similarity among KIR genes presents considerable challenges to bioinformatic analyses. PING addressed these issues through a highly-dynamic alignment workflow, which constructs individualized references that reflect the determined copy number and genotype makeup of a sample. This alignment workflow is enabled by a custom alignment processing pipeline, which scaffolds reads aligned to all reference sequences from the same gene into an overall gene alignment, enabling processing of these alignments as if a single reference sequence was used regardless of the number of sequences or of any insertions or deletions present in the component sequences. Together, these methods provide a novel and robust workflow for the accurate interpretation of KIR short-read sequencing data.

immunology

DR2S: An Integrated Algorithm Providing Reference-Grade Haplotype Sequences from Heterozygous Samples

BackgroundHigh resolution HLA genotyping of donors and recipients is a crucially important prerequisite for haematopoetic stem-cell transplantation and relies heavily on the quality and completeness of immunogenetic reference sequence databases of allelic variation. ResultsHere, we report on DR2S, an R package that leverages the strengths of two sequencing technologies - the accuracy of next-generation sequencing with the read length of third-generation sequencing technologies like PacBios SMRT sequencing or ONTs nanopore sequencing - to reconstruct fully-phased high-quality full-length haplotype sequences. Although optimised for HLA and KIR genes, DR2S is applicable to all loci with known reference sequences provided that full-length sequencing data is available for analysis. In addition, DR2S integrates supporting tools for easy visualisation and quality control of the reconstructed haplotype to ensure suitability for submission to public allele databases. ConclusionsDR2S is a largely automated workflow designed to create high-quality fully-phased reference allele sequences for highly polymorphic gene regions such as HLA or KIR. It has been used by biologists to successfully characterise and submit more than 500 HLA alleles and more than 500 KIR alleles to the IPD-IMGT/HLA and IPD-KIR databases.

bioinformatics

HLAssign 2.0: An advanced Graphical User Interface for the analysis of short and long read Human Leukocyte Antigen-typing data

Next Generation Sequencing (NGS) based Human Leukocyte Antigen (HLA) typing has been a challenge due to the polymorphism of the HLA region. Nevertheless, the methods accuracy has increased during the last years and it is now routinely used by many large centers including bone marrow registries. However, challenging HLA genotype compositions exist, which hinder a fully automated analysis. Therefore, HLA typing results are still visually inspected in diagnostics, i.e. the underlying read mappings and phasing information is controlled. Here, we present HLAssign 2.0 that now includes a strict workflow, improved tools for visual inspection and read phasing analysis in the automatic caller. In collaboration with interface design researchers, biologists and informaticians we developed an elaborate graphical user interface for visual evaluation of automated HLA calls for Illumina NGS reads. We also provide tools to preprocess 10x Genomics and PacBio sequencing reads for HLAssign analysis. We benchmarked our automatic caller against STC-seq and xHLA, showing comparable automatic call rates. Additional manual inspection of the automatic results in our GUI assists the user to assign the correct HLA calls and to achieve diagnostic accuracy. HLAssign 2.0 is free for research and commercial use and is available for Windows and MacOS.

bioinformatics