bioRxiv Science⌕ Search

Biology subjects

Nickols, W. A.

Publications and source records attributed to Nickols, W. A..

3 recordsLinked to original sources

Distinguishing new from persistent infections at the strain level using longitudinal genotyping data

MotivationLongitudinal pathogen genotyping data from individual hosts can uncover strain-specific infection dynamics and their relationships to disease and intervention, especially in the malaria field. An important use case involves distinguishing newly incident from pre-existing (persistent) strains, but implementation faces statistical challenges relating to individual samples containing multiple strains, strains sharing alleles, and markers dropping out stochastically during the genotyping process. Current approaches to distinguish new versus persistent strains therefore rely primarily on simple rules that consider only the time since alleles were last observed. ResultsWe developed DINEMITES (Distinguishing New Malaria Infections in Time Series), a set of statistical methods to estimate, from longitudinal genotyping data, the probability each sequenced allele represents a new infection harboring that allele, the total molecular force of infection (molFOI, the cumulative number of newly acquired strains over time) for each individual, and the total number of new infection events for each individual. DINEMITES can handle time points with missing sequencing data, incorporate treatment history and covariates affecting the rate of new or persistent infections, and can scale to studies with thousands of samples sequenced across multiple loci containing hundreds of possible alleles. In synthetic evaluations, the DINEMITES Bayesian model, which generally outperformed an alternative clustering-based model also developed in this work, accurately estimated key clinical parameters such as molFOI (bias 2.5, compared to -12.2 for a typical simple rule). When applied to three real longitudinal genotyping datasets, the model detected 33%, 112%, and 359% more average infections per participant than would have been detected by applying a typical simple rule to the equivalent datasets without sequencing. Availability and implementationDINEMITES is freely available as an R package, along with documentation, tutorials, and example data, at https://github.com/WillNickols/dinemites.

genomics↗

MaAsLin 3: Refining and extending generalized multivariable linear models for meta-omic association discovery

A key question in microbial community analysis is determining which microbial features are associated with community properties such as environmental or health phenotypes. This statistical task is impeded by characteristics of typical microbial community profiling technologies, including sparsity (which can be either technical or biological) and the compositionality imposed by most nucleotide sequencing approaches. Many models have been proposed that focus on how the relative abundance of a feature (e.g. taxon or pathway) relates to one or more covariates. Few of these, however, simultaneously control false discovery rates, achieve reasonable power, incorporate complex modeling terms such as random effects, and also permit assessment of prevalence (presence/absence) associations and absolute abundance associations (when appropriate measurements are available, e.g. qPCR or spike-ins). Here, we introduce MaAsLin 3 (Microbiome Multivariable Associations with Linear Models), a modeling framework that simultaneously identifies both abundance and prevalence relationships in microbiome studies with modern, potentially complex designs. MaAsLin 3 also newly accounts for compositionality with experimental (spike-ins and total microbial load estimation) or computational techniques, and it expands the space of biological hypotheses that can be tested with inference for new covariate types. On a variety of synthetic and real datasets, MaAsLin 3 outperformed current state-of-the-art differential abundance methods in testing and inferring associations from compositional data. When applied to the Inflammatory Bowel Disease Multi-omics Database, MaAsLin 3 corroborated many previously reported microbial associations with the inflammatory bowel diseases, but notably 77% of associations were with feature prevalence rather than abundance. In summary, MaAsLin 3 enables researchers to identify microbiome associations with higher accuracy and more specific association types, especially in complex datasets with multiple covariates and repeated measures.

microbiology↗

Evaluating metagenomic analyses for undercharacterized environments: what's needed to light up the microbial dark matter?

Non-human-associated microbial communities play important biological roles, but they remain less understood than human-associated communities. Here, we assess the impact of key environmental sample properties on a variety of state-of-the-art metagenomic analysis methods. In simulated datasets, all methods performed similarly at high taxonomic ranks, but newer marker-based methods incorporating metagenomic assembled genomes outperformed others at lower taxonomic levels. In real environmental data, taxonomic profiles assigned to the same sample by different methods showed little agreement at lower taxonomic levels, but the methods agreed better on community diversity estimates and estimates of the relationships between environmental parameters and microbial profiles.

microbiology↗