bioRxiv ScienceSearch

Biology subjects

David, L. A.

Publications and source records attributed to David, L. A..

3 recordsLinked to original sources

Statistical Considerations in the Design and Analysis of Longitudinal Microbiome Studies

Longitudinal studies of microbial communities have emphasized that host-associated microbiota are highly dynamic as well as underscoring the potential biomedical relevance of understanding these dynamics. Despite this increasing appreciation, statistical challenges in the design and analysis of longitudinal microbiome studies such as sequence counting, technical variation, signal aliasing, contamination, sparsity, missing data, and algorithmic scalability remain. In this review we discuss these challenges and highlight current progress in the field. Where possible, we try to provide guidelines for best practices as well as discuss how to tailor design and analysis to the hypothesis and ecosystem under study. Overall, this review is intended to serve as an introduction to longitudinal microbiome studies for both statisticians new to the microbiome field as well as biologists with little prior experience with longitudinal study design and analysis.

microbiology

Dynamic linear models guide design and analysis of microbiota studies within artificial human guts

Artificial gut models provide unique opportunities to study human-associated microbiota. Outstanding questions for these models fundamental biology include the timescales on which microbiota vary and the factors that drive such change. Answering these questions though requires overcoming analytical obstacles like estimating the effects of technical variation on observed microbiota dynamics, as well as the lack of appropriate benchmark datasets. To address these obstacles, we created a modeling framework based on multinomial logistic-normal dynamic linear models (MALLARDs) and performed dense longitudinal sampling of replicate artificial human guts over the course of 1 month. The resulting analyses revealed that when observed on an hourly basis, 76% of community variation could be ascribed to technical noise from sample processing, which could also skew the observed covariation between taxa. Our analyses also supported hypotheses that human gut microbiota fluctuate on sub-daily timescales in the absence of a host and that microbiota can follow replicable trajectories in the presence of environmental driving forces. Finally, multiple aspects of our approach are generalizable and could ultimately be used to facilitate the design and analysis of longitudinal microbiota studies in vivo.

genomics

Phylofactorization - a graph partitioning algorithm to identify phylogenetic scales of ecological data

The problem of pattern and scale is a central challenge in ecology. The problem of scale is central to community ecology, where functional ecological groups are aggregated and treated as a unit underlying an ecological pattern, such as aggregation of \"nitrogen fixing trees\" into a total abundance of a trait underlying ecosystem physiology. With the emergence of massive community ecological datasets, from microbiomes to breeding bird surveys, there is a need to objectively identify the scales of organization pertaining to well-defined patterns in community ecological data.\n\nThe phylogeny is a scaffold for identifying key phylogenetic scales associated with macroscopic patterns. Phylofactorization was developed to objectively identify phylogenetic scales underlying patterns in relative abundance data. However, many ecological data, such as presence-absences and counts, are not relative abundances, yet it is still desireable and informative to identify phylogenetic scales underlying a pattern of interest. Here, we generalize phylofactorization beyond relative abundances to a graph-partitioning algorithm for any community ecological data.\n\nGeneralizing phylofactorization connects many tools from data analysis to phylogenetically-informe analysis of community ecological data. Two-sample tests identify three phylogenetic factors of mammalian body mass which arose during the K-Pg extinction event, consistent with other analyses of mammalian body mass evolution. Projection of data onto coordinates defined by the phylogeny yield a phylogenetic principal components analysis which refines our understanding of the major sources of variation in the human gut microbiome. These same coordinates allow generalized additive modeling of microbes in Central Park soils and confirm that a large clade of Acidobacteria thrive in neutral soils. Generalized linear and additive modeling of exponential family random variables can be performed by phylogenetically-constrained reduced-rank regression or stepwise factor contrasts. We finish with a discussion of how phylofac-torization produces an ecological species concept with a phylogenetic constraint. All of these tools can be implemented with a new R package available online.

ecology