bioRxiv ScienceSearch

Biology subjects

Poon, A.

Publications and source records attributed to Poon, A..

4 recordsLinked to original sources

An open-source k-mer based machine learning tool for fast and accurate subtyping of HIV-1 genomes

For many disease-causing virus species, global diversity is clustered into a taxonomy of subtypes with clinical significance. In particular, the classification of infections among the subtypes of human immunodeficiency virus type 1 (HIV-1) is a routine component of clinical management, and there are now many classification algorithms available for this purpose. Although several of these algorithms are similar in accuracy and speed, the majority are proprietary and require laboratories to transmit HIV-1 sequence data over the network to remote servers. This potentially exposes sensitive patient data to unauthorized access, and makes it impossible to determine how classifications are made and to maintain the data provenance of clinical bioinformatic workflows. We propose an open-source supervised and alignment-free subtyping method (KO_SCPCAPAMERISC_SCPCAP) that operates on k-mer frequencies in HIV-1 sequences. We performed a detailed study of the accuracy and performance of subtype classification in comparison to four state-of-the-art programs. Based on our testing data set of manually curated real-world HIV-1 sequences (n = 2, 784), Kameris obtained an overall accuracy of 97%, which matches or exceeds all other tested software, with a processing rate of over 1,500 sequences per second. Furthermore, our fully standalone general-purpose software provides key advantages in terms of data security and privacy, transparency and reproducibility. Finally, we show that our method is readily adaptable to subtype classification of other viruses including dengue, influenza A, and hepatitis B and C virus.

bioinformatics

Food web networks shift across a precipitation gradient due to changes in community composition and species interactions.

Ecological networks change across spatial and environmental gradients due to (i) changes in species composition or (ii) changes in the frequency or strength of interactions. Here we use the communities of aquatic invertebrates inhabiting clusters of bromeliad phytotelms along the Brazilian coast as a model system for examining turnover in the properties of ecological networks. We first document the variation in the species pools of sites across a geographical climate gradient. Using the same sites, we also explored the geographic variation in species interaction strength using a newly developed Markov network approach. We found that community composition differed along a gradient of water volume within bromeliads due to the turnover of some species. From the Markov network analysis, we found that the top-down effects of certain predators differed geographically, which could also be explained by geographic differences in bromeliad water volumes. Overall, this study illustrates how a network can change across an environmental gradient through both changes in both species and their interactions.

ecology

A model-based clustering method to detect infectious disease transmission outbreaks from sequence variation

Clustering infections by genetic similarity is a popular technique for identifying potential outbreaks of infectious disease, in part because sequences are now routinely collected for clinical management of many infections. A diverse number of nonparametric clustering methods have been developed for this purpose. These methods are generally intuitive, rapid to compute, and readily scale with large data sets. However, we have found that nonparametric clustering methods can be biased towards identifying clusters of diagnosis -- where individuals are sampled sooner post-infection -- rather than the clusters of rapid transmission that are meant to be potential foci for public health efforts. We develop a fundamentally new approach to genetic clustering based on fitting a Markov-modulated Poisson process (MMPP), which represents the evolution of transmission rates along the tree relating different infections. We evaluated this model-based method alongside five nonparametric clustering methods using both simulated and actual HIV sequence data sets. For simulated clusters of rapid transmission, the MMPP clustering method obtained higher mean sensitivity (85%) and specificity (91%) than the nonparametric methods. When we applied these clustering methods to published HIV-1 sequences from a study cohort of men who have sex with men in Seattle, USA, we found that the MMPP method categorized about half (46%) as many individuals to clusters compared to the other methods, and that the MMPP clusters were more consistent with transmission outbreaks. This new approach to genetic clustering has significant implications for the application of pathogen sequence analysis to public health, where it is critical to robustly and accurately identify clusters for the most cost-effective deployment of resources.

epidemiology

OMSV enables accurate and comprehensive identification of large structural variations from nanochannel-based single-molecule optical maps

Human genomes contain structural variations (SVs) that are associated with various phenotypic variations and diseases. SV detection by sequencing is incomplete due to limited read length. Nanochannel-based optical mapping (OM) allows direct observation of SVs up to hundreds of kilo-bases in size on individual DNA molecules, making it a promising alternative technology for identifying large SVs. SV detection from optical maps is non-trivial due to complex types of error present in OM data, and no existing methods can simultaneously handle all these complex errors and the wide spectrum of SV types. Here we present a novel method, OMSV, for accurate and comprehensive identification of SVs from optical maps. OMSV detects both homozygous and heterozygous SVs, SVs of various types and sizes, and SVs with and without creating/destroying restriction sites. In an extensive series of tests based on real and simulated data, OMSV achieved both high sensitivity and specificity, with clear performance gains over the latest existing method. Applying OMSV to a human cell line, we identified hundreds of SVs >2kbp, with 65% of them missed by sequencing-based callers. Independent experimental validations confirmed the high accuracy of these SVs. We also demonstrate how OMSV can incorporate sequencing data to determine precise SV break points and novel sequences in the SVs not contained in the reference. We provide OMSV as open-source software to facilitate systematic studies of large SVs.

bioinformatics