bioRxiv Science⌕ Search

Biology subjects

Smilde, A.

Publications and source records attributed to Smilde, A..

5 recordsLinked to original sources

Longitudinal metabolomics data analysis informed bymechanistic models

MotivationMetabolomics measurements are noisy, often characterized by a small sample size and missing entries. While data-driven methods have shown promise in terms of analyzing metabolomics data, e.g., revealing biomarkers of various phenotypes, metabolomics data analysis can significantly benefit from incorporating prior information about metabolic mechanisms. In this paper, we introduce a novel data analysis approach where data-driven methods are guided by prior information through joint analysis of simulated data generated using a human metabolic model and real metabolomics measurements. ResultsWe arrange time-resolved metabolomics measurements of plasma samples collected during a meal challenge test from the COPSAC2000 cohort as a third-order tensor: subjects by metabolites by time samples. Simulated challenge test data generated using a human whole-body metabolic model is also arranged as a third-order tensor: virtual subjects by metabolites by time samples. Real and simulated data sets are coupled in the metabolites mode and jointly analyzed using coupled tensor factorizations to reveal the underlying patterns. Our experiments demonstrate that joint analysis of simulated and real data has a better performance in terms of pattern discovery achieving higher correlations with a BMI (body mass index)-related phenotype compared to the analysis of only real data in males while in females, the performance is comparable. We also demonstrate the advantages of such a joint analysis approach in the presence of incomplete measurements and its limitations in the presence of wrong prior information. AvailabilityThe code for joint analysis of real and simulated metabolomics data sets is released as a GitHub repository. Simulated data can also be accessed using the GitHub repo. Real measurements of plasma samples are not publicly available. Data may be shared by COPSAC through a collaboration agreement. Data access requests should be directed to Morten A. Rasmussen (morten.arendt@dbac.dk).

bioinformatics↗

parafac4microbiome: Exploratory analysis of longitudinal microbiome data using Parallel Factor Analysis

Studies investigating microbial temporal dynamics are increasingly common, leveraging longitudinal designs that collect microbial abundance data across multiple time points from the same subjects. Traditional exploratory approaches like Principal Component Analysis (PCA) fail to fully utilize this structure. By organizing data as a three-way array--subjects as rows, microbial abundances as columns, and time points as the third dimension--multi-way methods such as Parallel Factor Analysis (PARAFAC) can better capture temporal and structural patterns. This study demonstrates Parallel Factor Analysis (PARAFAC) as a method to explore longitudinal microbiome data using three exemplary studies. In the first example, a long time series of in vitro microbiomes, PARAFAC identifies primary time-resolved variations. The second example, a longitudinal infant gut microbiome study, shows that PARAFAC can distinguish subject groups and enhance comparative analysis, even with moderate missing data. In the third example, a gingivitis intervention study of the oral microbiome, PARAFAC enables the identification of microbial subcommunities of interest through post-hoc clustering. These examples highlight PARAFACs broad applicability for analysing longitudinal microbiome data across diverse environments. The approach is implemented in the R package parafac4microbiome, available on CRAN, providing researchers with accessible tools for similar analyses. ImportanceUnderstanding how microbiomes change over time can give us valuable insights into their role in health and disease. Many traditional methods like Principal Component Analysis (PCA) miss important patterns in data collected over time, but Parallel Factor Analysis (PARAFAC) helps uncover these trends in a much clearer way. Using this approach, we were able to identify key changes in microbiomes across different settings, like lab experiments, the infant gut, and the mouth. PARAFAC also works well even when some data is missing, which is a common issue. To make this tool accessible, we have included it in a user-friendly R package, enabling other researchers to analyse microbiome dynamics in their own studies and explore how these changes might influence health and treatments.

microbiology↗

Multi-way modelling of oral microbial dynamics and host-microbiome interactions during induced gingivitis

Gingivitis - the inflammation of the gums - is a reversible stage of periodontal disease. It is caused by dental plaque formation due to poor oral hygiene. However, gingivitis susceptibility involves a complex set of interactions between the oral microbiome, oral metabolome and the host. In this study, we investigated the dynamics of the oral microbiome and its interactions with the salivary metabolome during experimental gingivitis in a cohort of 41 systemically healthy participants. We use Parallel Factor Analysis (PARAFAC), which is a multi-way generalization of Principal Component Analysis (PCA) that can model the variability in the response due to subjects, variables and time. Using the modelled responses, we identified microbial subcommunities with similar dynamics that connect to the magnitude of the gingivitis response. By performing high level integration of the predicted metabolic functions of the microbiome and salivary metabolome, we identified pathways of interest that describe the changing proportions of Gram-positive and Gram-negative microbiota, variation in anaerobic bacteria, biofilm formation and virulence.

systems biology↗

MASCARA: coexpression analysis in data from designed experiments

Experiments in plant transcriptomics are usually designed to induce variation in a pathway of interest. Harsh experimental conditions can cause widespread transcriptional changes between groups. Discovering coexpression within a pathway of interest (here the strigolactone pathway) in this context is hampered by the dominant variance induced by the design. Minor changes in experimental conditions not controlled for may affect the plants, leading to small coordinated differences in genes within pathways of interest and related pathways between replicate plants in the same controlled experimental condition. These systematic differences are usually averaged out, but we argue here that they can be used to improve the detection of genes that co-express. We introduce a novel framework "MASCARA" which combines ANOVA simultaneous component analysis and partial least squares to remove the experimentally induced variance and investigate multivariate relationships in the non-designed variance. MASCARA is tested against a selection of competitors on simulated data, created to mimic a designed transcriptome study, where its benefit is demonstrated. In a coexpression analysis of a real dataset MASCARA detects several uncharacterised but relevant transcripts. Our results indicate that there is sufficient structure left in a typical dataset after correcting for experimental variance and that this residual information is useful to investigate coexpression. Author SummaryExperiments in the life sciences usually purposefully induce significant variance between different treatments, in order to activate or repress certain mechanisms of interest. Whilst this is necessary it can make it challenging to detect meaningful relationships within pathways of interest, particularly when the experimental conditions are drastically different. Instead of focusing on the drastic changes in response due to the different treatment, MASCARA uses the systematic synchronous variances between replicates to find related features within the pathway of interest. Through simulation studies and application to a real dataset, we demonstrate the effectiveness of MASCARA in detecting relevant transcripts and extracting coexpression patterns from gene expression data.

bioinformatics↗

Corruption of the Pearson correlation coefficient by measurement error: estimation, bias, and correction under different error models

Correlation coefficients are abundantly used in the life sciences. Their use can be limited to simple exploratory analysis or to construct association networks for visualization but they are also basic ingredients for sophisticated multivariate data analysis methods. It is therefore important to have reliable estimates for correlation coefficients. In modern life sciences, comprehensive measurement techniques are used to measure metabolites, proteins, gene-expressions and other types of data. All these measurement techniques have errors. Whereas in the old days, with simple measurements, the errors were also simple, that is not the case anymore. Errors are heterogeneous, non-constant and not independent. This hampers the quality of the estimated correlation coefficients seriously. We will discuss the different types of errors as present in modern comprehensive life science data and show with theory, simulations and real-life data how these affect the correlation coefficients. We will briefly discuss ways to improve the estimation of such coefficients.

bioinformatics↗