bioRxiv Science⌕ Search

Biology subjects

Hung, T.-K.

Publications and source records attributed to Hung, T.-K..

5 recordsLinked to original sources

deCYPher: Star Allele-Resolution Computational Framework of Pharmacogenes for Haplotype-Resolved Long-Read Assemblies

Although existing next-generation sequencing (NGS) tools, such as Aldy and Cyrius, have been applied for allele typing, they cannot achieve complete accuracy due to various genomic challenges including pseudogenes, structural variations, hybrid genes, copy number variations, and gene deletions. These complexities make accurate pharmacogene interpretation more challenging, despite the crucial role pharmacogenomics plays in precision medicine. We developed deCYPher, a tool that generates personalized pharmacogenomic reports from haplotype-resolved assemblies. The tool enables analysis of all PharmVar 1A level genes, such as CYP2B6, CYP2C9, CYP2C19, CYP2D6, CYP3A5, CYP4F2, DPYD, NUDT15, and SLCO1B1. Applied to all HPRC haplotypes (including both release 1 and release 2 data), deCYPher demonstrated high accuracy in resolving complex gene structures. In the case of CYP2D6, release 1 identified 6% gene multiplications, 6% full gene deletions, and 4% CYP2D6/CYP2D7 hybrids. By contrast, release 2 demonstrated an increased prevalence of multiplications (14%) and hybrids (11%), while the frequency of full gene deletions remained comparable at 5%. Comparison with pb-StarPhase revealed discrepancies in 12 of 94 assemblies in the release 1 dataset. For instance, in sample HG02257, Aldy, Cyrius, and deCYPher consistently identified the genotype as *2/*35, whereas pb-StarPhase reported *2/*2. Notably, the *35-defining variants were present in the BAM and VCF files in the pb-StarPhase pipeline, but the local read depth over the *35-specific region was only 5x in HG02257-p, suggesting that the misclassification likely resulted from insufficient coverage - a known limitation of pb-StarPhase under low-depth conditions.

bioinformatics↗

gAIRR-wgs: An Algorithm to Genotype T Cell Receptor Alleles Using Whole Genome Sequencing Data

T cell receptor (TR) genes, including variable (TR_V), diversity (TR_D), and joining (TR_J) segments, exhibit allelic diversity that is critical to adaptive immunity. Growing evidence has identified associations between TR genes and immune-related diseases. Germline variants may influence TR gene function and subsequent usage, highlighting the importance of accurate TR allele profiling. However, accurately identifying germline TR from standard WGS data remains challenging due to short read lengths, limited depth, and high sequence similarity. To address these challenges, we developed gAIRR-wgs, for WGS-based TR allele typing. By incorporating novel alleles from HPRC individuals, gAIRR-wgs exhibited excellent performance in allele calling, with F1 scores of 100.0% for TR_D, 99.8% for TR_J, and 98.3% for TR_V. Applying this pipeline to 1,492 individuals from the Taiwan Biobank (TWB), we identified 449 novel TR alleles, 277 of which overlapped with HPRC release 1 data of mixed ethnicity and are absent in the IMGT database. Further population comparison analysis revealed significant TR allele distribution differences across global populations, showing population-specific patterns and diversity variations between ethnic groups. We also discovered TWB-specific deletion polymorphisms affecting contiguous TRGV and TRBV genes, which are not recorded in the gnomAD database and undetected by standard structural variant callers, highlighting the need for tailored approaches to resolve complex immune gene regions. In conclusion, gAIRR-wgs enables accurate TR allele calling from standard WGS data using feasible computational resources and reveals substantial immunogenetic diversity in population cohorts.

bioinformatics↗

TypeAssembly: Copy number estimation and allele typing for haplotype assemblies

Accurately annotating complex genes in the human genome, particularly from haplotype assemblies, remains a significant challenge. To overcome this, we developed TypeAssembly, a local alignment-based framework for copy number estimation and allele typing. Operating in two modes, mode-FASTA and mode-VCF, TypeAssembly can define alleles by either sequence or variant information. We successfully applied it to annotate 41 genes in the MHC locus, 17 KIR genes, and, for the first time, 15 pharmacogenes across 466 haplotype assemblies. This study establishes TypeAssembly as a robust method for accurately annotating complex genomic regions and provides an evaluation of existing gene annotations and callers.

bioinformatics↗

Graph-KIR: Graph-based KIR Copy Number Estimation and Allele Calling Using Short-read Sequencing Data

MotivationThe Killer-cell Immunoglobulin-like Receptor (KIR) is a highly polymorphic region in the human genome, associated with autoimmune diseases and organ transplantation. The sequences of KIR genes are highly similar among star alleles as well as in between individual genes, with the copy number of each KIR gene typically ranging from 0 to 4. In this study, we introduce a tool Graph-KIR that aims to estimate the copy number of genes and to call full-resolution (7-digit) KIR alleles from a whole genome sequencing (WGS) sample. ResultsGraph-KIR, unlike most KIR tools, is capable of independently typing KIR alleles per sample with no reliance on the distribution of any framework gene in a cohort. In a set of 100 simulated samples, Graph-KIR demonstrated 100% accuracy in copy number estimation and high accuracy of allele typing: 91.2% at 7-digit resolution, 97.0% at 5-digit resolution, 97.2% at 3-digit resolution, and 99.6% at gene-level resolution. Graph-KIR outperforms existing tools such as PINGs WGS version (91.9% accuracy) and T1K (84.6% accuracy) at 5-digit resolution. By analyzing the results on 44 HPRC samples, Graph-KIR achieves an accuracy of 85.0%, better than PINGs WGS version (75.5% accuracy) at 5-digit resolution. The release of Graph-KIR adds another valuable tool to assist users in accurately estimating copy numbers and calling alleles of KIR genes from WGS samples, ensuring reliable performance. AvailabilityThe Graph-KIR and paper-related pipeline codes are available at https://github.com/linnil1/KIR_graph.

bioinformatics↗

Genetic Diversity and Structural Complexity of the Killer-Cell Immunoglobulin-Like Receptor Gene Complex: A Comprehensive Analysis using Human Pangenome Assemblies

The killer-cell immunoglobulin-like receptor (KIR) gene complex, a highly polymorphic region of the human genome that encodes proteins involved in immune responses, poses strong challenges in genotyping due to its remarkable genetic diversity and structural intricacy. Accurate analysis of KIR alleles, including their structural variations, is crucial for understanding their roles in various immune responses. Leveraging the high-quality genome assemblies from the Human Pangenome Reference Consortium (HPRC), we present a novel bioinformatic tool, the Structural KIR annoTator (SKIRT), to investigate gene diversity and facilitate precise KIR allele analysis. We applied SKIRT on 47 HPRC-phased assemblies and identified a recurrent novel KIR2DS4/3DL1 fusion gene in the paternal haplotype of HG02630 and maternal haplotype of NA19240. Additionally, SKIRT accurately identifies eight structural variants and 17 novel nonsynonymous alleles, all of which were independently validated using short-read data or quantitative polymerase chain reaction. Our study has discovered a total of 570 novel alleles, among which eight haplotypes harbor at least one KIR gene duplication, six haplotypes have lost at least one framework gene, and 75 out of 94 haplotypes (79.8%) carry at least five novel alleles, thus confirming KIR genetic diversity. These findings are pivotal in providing insights into KIR gene diversity and serve as a solid foundation for understanding the functional consequences of KIR structural variations. High-resolution genome assemblies offer unprecedented opportunities to explore polymorphic regions that are challenging to investigate using short-read sequencing methods. The SKIRT pipeline emerges as a highly efficient tool, enabling the comprehensive detection of the complete spectrum of KIR alleles within human genome assemblies.

genomics↗