bioRxiv ScienceSearch

Biology subjects

Maurer-Stroh, S.

Publications and source records attributed to Maurer-Stroh, S..

2 recordsLinked to original sources

Phylogenetic Clustering by Linear Integer Programming (PhyCLIP)

Sub-species nomenclature systems of pathogens are increasingly based on sequence data. The use of phylogenetics to identify and differentiate between clusters of genetically similar pathogens is particularly prevalent in virology from the nomenclature of human papillomaviruses to highly pathogenic avian influenza (HPAI) H5Nx viruses. These nomenclature systems rely on absolute genetic distance thresholds to define the maximum genetic divergence tolerated between viruses designated as closely related. However, the phylogenetic clustering methods used in these nomenclature systems are limited by the arbitrariness of setting intra- and inter-cluster diversity thresholds. The lack of a consensus ground truth to define well-delineated, meaningful phylogenetic subpopulations amplifies the difficulties in identifying an informative distance threshold. Consequently, phylogenetic clustering often becomes an exploratory, ad-hoc exercise.\n\nPhylogenetic Clustering by Linear Integer Programming (PhyCLIP) was developed to provide a statistically-principled phylogenetic clustering framework that negates the need for an arbitrarily-defined distance threshold. Using the pairwise patristic distance distributions of an input phylogeny, PhyCLIP parameterises the intra- and inter-cluster divergence limits as statistical bounds in an integer linear programming model which is subsequently optimised to cluster as many sequences as possible. When applied to the haemagglutinin phylogeny of HPAI H5Nx viruses, PhyCLIP was not only able to recapitulate the current WHO/OIE/FAO H5 nomenclature system but also further delineated informative higher resolution clusters that capture geographically-distinct subpopulations of viruses. PhyCLIP is pathogen-agnostic and can be generalised to a wide variety of research questions concerning the identification of biologically informative clusters in pathogen phylogenies. PhyCLIP is freely available at http://github.com/alvinxhan/PhyCLIP.

bioinformatics

A Novel Method for the Capture-based Purification of Whole Viral Native RNA Genomes

Current technologies for targeted characterization and manipulation of viral RNA either involve amplification or ultracentrifugation with isopycnic gradients of viral particles to decrease host RNA background. The former strategy is non-compatible for characterizing properties innate to RNA strands such as secondary structure, RNA-RNA interactions, and also for nanopore direct RNA sequencing involving the sequencing of native RNA strands. The latter strategy, ultracentrifugation, causes loss in genomic information due to its inability to retrieve unassembled viral RNA. We developed a novel nucleic acid manipulation technique involving the capture of whole viral native RNA genomes for downstream RNA assays using hybridization baits in solution to circumvent these problems. This technique involves hybridization of biotinylated baits at 500 nucleotides (nt) intervals, stringent washes and release of free native RNA strands using DNase I treatment, with a turnaround time of about 6 h 15 min. Proof of concept was primarily done using RT-qPCR with dengue virus infected Huh-7 cells. We report that this protocol was able to purify viral RNA (561-791 fold). We also describe a successful application of our capture-based purification method to direct RNA sequencing, with a 77.47% of reads mapping to the target viral genome. We observed a reduction in human host RNA background by 1580 fold, a 99.91% recovery of viral genome with at least 15x coverage, and a mean coverage across the genome of 120x. This report is, to the best of our knowledge, the first description of a capture-based purification method for whole viral RNA genomes. The fundamental advantages of using our capture-based purification method makes it a superior alternative to conventional viral purification methods and would potentially pave a new path for the direct characterization and sequencing of native RNA molecules.

genomics