bioRxiv Science⌕ Search

Biology subjects

Korenek, A.

Publications and source records attributed to Korenek, A..

2 recordsLinked to original sources

Evaluation of statistical approaches for differential metaproteomics

BackgroundMetaproteomics characterizes and compares molecular phenotypes of organisms in communities by comprehensively analyzing their protein expression profiles using statistical methods. However, not all statistical methods are suitable for determining differentially abundant protein groups in metaproteomic analyses. Statistical challenges in metaproteomics include: data sparsity, non-normality, compositionality, and large between-sample variability. These challenges can potentially be addressed with several data processing steps, including imputation, normalization, transformation, and selection of the appropriate statistical tests. The potential combinations of different processing methods create a complex matrix of analysis options and it is currently unclear how these combinations impact the results of statistical tests on metaproteomic data. ResultsTo determine what data processing methods and statistical tests are best for identifying differentially abundant proteins in metaproteomics datasets, we generated a set of thirteen metaproteomic samples with known compositions, known differences, and differing levels of complexity. These defined metaproteomes address the general challenges outlined above, using various scenarios in metaproteomic data analyses. We compared over 110 different statistical analysis combination options, including regression-based tools, general statistics inference, and machine learning techniques. We found that several combinations within the frameworks of limma, edgeR, MaAslin2, custom linear and Bayesian linear models, and random forests all offer suitable evaluation options. ConclusionsWe highlight key recommendations for differential expression analysis in metaproteomics. Our work enables improved assessment of statistical methods for metaproteomics by establishing a framework for testing statistical approaches, including comprehensive raw mass spectrometry data and reproducible benchmarking code.

molecular biology↗

Metaproteomics and DNA metabarcoding as tools to assess dietary intake in humans

Objective biomarkers of food intake are a sought-after goal in nutrition research. Most biomarker development to date has focused on metabolites detected in blood, urine, skin or hair, but detection of consumed foods in stool has also been shown to be possible via DNA sequencing. An additional food macromolecule in stool that harbors sequence information is protein. However, the use of protein as an intake biomarker has only been explored to a very limited extent. Here, we evaluate and compare measurement of residual food-derived DNA and protein in stool as potential biomarkers of intake. We performed a pilot study of DNA sequencing-based metabarcoding (FoodSeq) and mass spectrometry-based metaproteomics in five individuals stool sampled in short, longitudinal bursts accompanied by detailed diet records (n=27 total samples). Dietary data provided by stool DNA, stool protein, and written diet record independently identified a strong within-person dietary signature, identified similar food taxa, and had significantly similar global structure in two of the three pairwise comparisons between measurement techniques (DNA-to-protein and DNA-to-diet record). Metaproteomics identified proteins including myosin, ovalbumin, and beta-lactoglobulin that differentiated food tissue types like beef from dairy and chicken from egg, distinctions that were not possible by DNA alone. Overall, our results lay the groundwork for development of targeted metaproteomic assays for dietary assessment and demonstrate that diverse molecular components of food can be leveraged to study food intake using stool samples.

molecular biology↗