bioRxiv Science⌕ Search

Biology subjects

Husmeier, D.

Publications and source records attributed to Husmeier, D..

2 recordsLinked to original sources

Do identification guides hold the key to speciesmisclassification by citizen scientists?

1O_LICitizen science data often contain high levels of species misclassification that can bias inference and conservation decisions. Current approaches to address mislabelling rely on expert taxonomists validating every record. This approach makes intensive use of a scarce resource and reduces the role of the citizen scientist. C_LIO_LISpecies, however, are not confused at random. If two species appear more similar, it is probable they will be more easily confused than two highly distinctive species. Identification guides are intended to use these patterns to aid correct classification, but misclassifications still occur due to user-error and imperfect guidebook design. Statistical models should be able to exploit this non-randomness to learn confusion patterns from small validation datasets provided by expert taxonomists, yielding a much-needed reduction in expert workload. Here, we use a variety of Bayesian hierarchical models to probabilistically classify species based on the species-label provided by the citizen scientist. We also explore the utility of guidebooks provided by the citizen science schemes as a prior for species similarity, and hence draw conclusions for their future improvement. C_LIO_LIWe find that the species-label assigned to a record by a citizen scientist, even when incorrect, contains useful information about the true species-identity. The citizen scientists correctly identify the species in around 58% of records. Using models trained on only 10% of these records (validated by experts), we can correctly predict species-identity for 69 (90%CI: 64-73)% of records when the guidebook is used, vs 64 (58-69)% for models that do not use the guidebook. The fact that misclassifications can be predicted systematically indicates that improvements could be made to the guidebook to reduce misclassification. C_LIO_LIBy using Bayesian, hierarchical models we can greatly reduce the workload for experts by providing a probabilistic correction to citizen science records, rather than requiring manual review. This is increasingly important as the number of citizen science schemes grows and the relative number of taxonomists shrinks. By learning confusion patterns statistically, we open up future avenues of research to identify what causes these confusions and how to better address them. C_LI

ecology↗

A Bayesian approach to incorporate structural data into the mapping of genotype to antigenic phenotype of influenza A(H3N2) viruses

Surface antigens of pathogens are commonly targeted by vaccine-elicited antibodies but antigenic variability, notably in RNA viruses such as influenza, HIV and SARS-CoV-2, pose challenges for control by vaccination. For example, influenza A(H3N2) entered the human population in 1968 causing a pandemic and has since been monitored, along with other seasonal influenza viruses, for the emergence of antigenic drift variants through intensive global surveillance and laboratory characterisation. Statistical models of the relationship between genetic differences among viruses and their antigenic similarity provide useful information to inform vaccine development though accurate identification of causative mutations is complicated by highly correlated genetic signals that arise due to the evolutionary process. Here, using a sparse hierarchical Bayesian analogue of an experimentally validated model for integrating genetic and antigenic data, we identify the genetic changes in influenza A(H3N2) virus that underpin antigenic drift. We show that incorporating protein structural data into variable selection helps resolve ambiguities arising due to correlated signals, with the proportion of variables representing haemagglutinin positions decisively included, or excluded, increased from 59.8% to 72.4%. The accuracy of variable selection judged by proximity to experimentally determined antigenic sites was improved simultaneously. Structure-guided variable selection thus improves confidence in the identification of genetic explanations of antigenic variation and we also show that prioritising the identification of causative mutations is not detrimental to the predictive capability of the analysis. Indeed, incorporating structural information into variable selection resulted in a model that could more accurately predict antigenic assay titres for phenotypically-uncharactrised virus from genetic sequence. Combined, these analyses have the potential to inform choices of reference viruses, the targeting of laboratory assays, and predictions of the evolutionary success of different genotypes, and can therefore be used to inform vaccine selection processes.

microbiology↗