bioRxiv Science⌕ Search

Biology subjects

Volchkov, P. Y.

Publications and source records attributed to Volchkov, P. Y..

3 recordsLinked to original sources

Systematic search for new HLA alleles in 4195 human 30x WGS samples

HLA (Human Leukocyte Antigens) is a highly polymorphic locus in the human genome which also has a high clinical significance. New alleles of HLA genes are constantly being discovered but mostly through the efforts of laboratories which primarily focus on HLA typing and are using field-specific experimental and data processing techniques, like enrichment of HLA region in high-throughput sequencing data. Nevertheless, a vast amount of whole genome sequencing (WGS) data was accumulated over the past years and continues to expand rapidly. Therefore it is an appealing possibility to identify new HLA alleles and refine the information on known alleles from already available WGS data. Currently there are many tools designed for HLA typing, e.g. assigning known alleles, from non HLA enriched WGS data, but none of them specifically tailored towards identification and immediate thorough description of new HLA alleles. Here we are presenting a pipeline HLAchecker, which is specifically designed to identify potentially new HLA alleles based on discrepancies between predicted HLA types, made by any other dedicated tool, and underlying raw 30x WGS data. HLAchecker reports structured in a way which simplifies further validation of potentially new HLA alleles and streamlines submission of alleles to appropriate databases. We validated this tool on 4195 30x WGS samples typed by HLA-HD, discovered 17 potentially new HLA alleles with substitutions in exonic regions and validated five randomly chosen alleles by Sanger sequencing.

bioinformatics↗

Sanger validation of WGS variants - when to?

With the development of Next-Generation Sequencing (NGS) technologies it became possible to simultaneously analyze millions of variants. Despite the quality improvement it is generally still required to confirm the variants before reporting. However, in recent years the dominant idea is that one could define the quality thresholds for "high quality" variants which do not require orthogonal validation. Despite that, no works to date report the concordance between variants from whole genome sequencing and their gold-standard Sanger validation. In this study we analyzed the concordance for 1756 WGS variants in order to establish the appropriate thresholds for high-quality variants filtering. Resulting thresholds allowed us to drastically reduce the number of variants which require validation, to 5,6% and 1.2% of the initial set for caller-agnostic thresholds and caller-dependent QUAL threshold respectively.

genomics↗

Concatenation of segmented viral genomes for reassortment analysis

Most reassortment identification methods are based on searching for phylogenetic discrepancies between phylogenetic trees for different segments. Other methods use pairwise genetic distances or compare the position of individual genome components in a tree relative to a reference component of the viral genome. However, such approaches are labour-intensive and hardly scalable. Recent advances in the availability of viral sequencing technologies have led to the sequencing of large numbers of pathogen genomes, making manual processing of this large data difficult. At the same time, recombination analysis methods can process almost any number of sequences simultaneously. Such approaches are not suitable for the simultaneous analysis of multiple segments and are therefore not used to search for reassortment events. However, in the case of sequential concatenation of all segments, the methods of recombination analysis can be used to detect traces of reassortment events. The code is available at https://github.com/melibrun/Concatenation-of-segmented-viral-genomes-for-reassortment-analysis. The service is implemented as a web application https://melibrun.shinyapps.io/viralsegmentconcatenator1/. It concatenates segmented viral genomes for reassortment analysis. The tool accepts files in GenBank format as input and generates a set of sequences in fasta format that are sequentially concatenated sequences of viral segments named in accordance with the "strain" field of the GenBank record annotation. In order to use recombination search algorithms in the study of reassortment events, we have developed a method (Virus Segment Concatenator, VSC) to automatically concatenate the sequences of all segments of a virus into a single sequence. The applicability of VSC for automated searches for reassortment events was demonstrated using CCHFV, an H5N5 subtype of influenza virus.

bioinformatics↗