bioRxiv Science⌕ Search

Biology subjects

Woroncow, M.

Publications and source records attributed to Woroncow, M..

4 recordsLinked to original sources

HLA alleles and haplotype distribution across Russian population groups

HLA loci are highly polymorphic genome regions, with allele frequencies varying significantly across different populations. Population HLA frequency databases may contain biases and make cross-study comparison complicated due to varying data curation protocols, genotyping methodologies, resolution, and inconsistencies in the selection criteria for population samples. This study presents HLA allele frequencies of class I (HLA-A, -B, -C) and class II (HLA-DRB1, -DQB1, -DQA1) as well as their combined haplotypes obtained from over 18,000 whole genome sequencing samples of the Russian population. Cohort was stratified based on PCA and admixture components providing frequencies for 14 different ethnic groups. For 12 groups cohort size allowed us to reach average saturation of 96% of allele frequencies in groups. Moreover, we demonstrated the utility of composed statistics for disease populational study using type 1 diabetes (T1D) as an example. Populations with similar aggregated genetic risk for T1D demonstrated substantial differences in frequencies of risk and protective HLA alleles. Obtained frequency data was made publicly available through the Allele Frequency Net Database improving previously sparse coverage in HLA frequencies data for east Europe and north Asia regions.

genetics↗

Systematic analysis of insertions signature in gnomAD revealed large set of novel processed pseudogenes

Pseudogenes are non-functional copies of protein-coding genes that arise through genomic duplication or retrotransposition. Processed pseudogenes (PPs) is the most abundant class of pseudogenes, which is generated via mRNA reverse transcription and subsequent cDNA integration. Presence of PPs complicates the analysis of short read sequencing data due to high similarity with parental gene and frequent absence from reference genome. Here we demonstrate that the presence of non-reference (absent from reference genome) PPs leads to the very distinctive artefact of germline variant calling - long insertions on exon-intron boundaries, which sequences could be mapped to other exons of the same gene. We showed that by detecting these artifacts it is possible to identify non-reference PPs existence based on the cohort summary statistics without analysing sample-level data. We used identified signature of PPs presence to systematically mine the gnomAD database which currently contains over 70,000 whole-genome and over 700,000 exome samples to describe novel non-reference PPs. Our approach uncovered 1498 non-reference PPs of which 1268 were novel and absent in the latest GENCODE release. This resource enhances the accuracy of variant interpretation and contributes to a deeper understanding of pseudogenes diversity across human populations.

genomics↗

Systematic search for new HLA alleles in 4195 human 30x WGS samples

HLA (Human Leukocyte Antigens) is a highly polymorphic locus in the human genome which also has a high clinical significance. New alleles of HLA genes are constantly being discovered but mostly through the efforts of laboratories which primarily focus on HLA typing and are using field-specific experimental and data processing techniques, like enrichment of HLA region in high-throughput sequencing data. Nevertheless, a vast amount of whole genome sequencing (WGS) data was accumulated over the past years and continues to expand rapidly. Therefore it is an appealing possibility to identify new HLA alleles and refine the information on known alleles from already available WGS data. Currently there are many tools designed for HLA typing, e.g. assigning known alleles, from non HLA enriched WGS data, but none of them specifically tailored towards identification and immediate thorough description of new HLA alleles. Here we are presenting a pipeline HLAchecker, which is specifically designed to identify potentially new HLA alleles based on discrepancies between predicted HLA types, made by any other dedicated tool, and underlying raw 30x WGS data. HLAchecker reports structured in a way which simplifies further validation of potentially new HLA alleles and streamlines submission of alleles to appropriate databases. We validated this tool on 4195 30x WGS samples typed by HLA-HD, discovered 17 potentially new HLA alleles with substitutions in exonic regions and validated five randomly chosen alleles by Sanger sequencing.

bioinformatics↗

Sanger validation of WGS variants - when to?

With the development of Next-Generation Sequencing (NGS) technologies it became possible to simultaneously analyze millions of variants. Despite the quality improvement it is generally still required to confirm the variants before reporting. However, in recent years the dominant idea is that one could define the quality thresholds for "high quality" variants which do not require orthogonal validation. Despite that, no works to date report the concordance between variants from whole genome sequencing and their gold-standard Sanger validation. In this study we analyzed the concordance for 1756 WGS variants in order to establish the appropriate thresholds for high-quality variants filtering. Resulting thresholds allowed us to drastically reduce the number of variants which require validation, to 5,6% and 1.2% of the initial set for caller-agnostic thresholds and caller-dependent QUAL threshold respectively.

genomics↗