bioRxiv Science⌕ Search

Biology subjects

Fleischmann, F.

Publications and source records attributed to Fleischmann, F..

2 recordsLinked to original sources

Individuality and information content of infrared molecular profiles: insights from a large longitudinal health-profiling study

In this study, we investigate the individuality and information content of infrared molecular profiles derived from blood samples in a large, longitudinal health-profiling cohort and compare them to a standard clinical laboratory panel. Using Fourier-transform infrared spectroscopy, we obtained comprehensive molecular fingerprints from 4,704 self-reported healthy individuals over five visits spanning 1.5 years, alongside routine clinical laboratory measurements. We show that infrared profiles are highly individual-specific and remarkably stable over time, with intra-individual variability significantly lower than inter-individual differences--paralleling the characteristics observed in clinical laboratory data. To quantify and compare the information content of these molecular datasets, we employ individual identification as a proxy for Shannon entropy. In this framework, higher identification accuracy reflects a higher amount of information. Infrared profiles outperform the clinical laboratory panel in identifying individuals at scale, suggesting higher intrinsic information content. Furthermore, combining infrared and clinical laboratory data substantially improves identification performance (the identification of less than 3000 individuals by the clinical laboratory panel is boosted to more than 4000 by incorporating the infrared spectroscopic markers), highlighting the value of integrating complementary data modalities. These findings suggest a practical framework, rooted in information theory, for comparing molecular profiling approaches and emphasize the potential of infrared spectroscopy as a complementary tool in personalized medicine.

biophysics↗

CODI: Enhancing machine learning-based molecular profiling through contextual out-of-distribution integration

Molecular analytics increasingly utilize machine learning (ML) for predictive modeling based on data acquired through molecular profiling technologies. However, developing robust models that accurately capture physiological phenotypes is challenged by a multitude of factors. These include the dynamics inherent to biological systems, variability stemming from analytical procedures, and the resource-intensive nature of obtaining sufficiently representative datasets. Here, we propose and evaluate a new method: Contextual Out-of-Distribution Integration (CODI). Based on experimental observations, CODI generates synthetic data that integrate unrepresented sources of variation encountered in real-world applications into a given molecular fingerprint dataset. By augmenting a dataset with out-of-distribution variance, CODI enables an ML model to better generalize to samples beyond the initial training data. Using three independent longitudinal clinical studies and a case-control study, we demonstrate CODIs application to several classification scenarios involving vibrational spectroscopy of human blood. We showcase our approachs ability to enable personalized fingerprinting for multi-year longitudinal molecular monitoring and enhance the robustness of trained ML models for improved disease detection. Our comparative analyses revealed that incorporating CODI into the classification workflow consistently led to significantly improved classification accuracy while minimizing the requirement of collecting extensive experimental observations. SIGNIFICANCE STATEMENTAnalyzing molecular fingerprint data is challenging due to multiple sources of biological and analytical variability. This variability hinders the capacity to collect sufficiently large and representative datasets that encompass realistic data distributions. Consequently, the development of machine learning models that generalize to unseen, independently collected samples is often compromised. Here, we introduce CODI, a versatile framework that enhances traditional classifier training methodologies. CODI is a general framework that incorporates information about possible out-of-distribution variations into a given training dataset, augmenting it with simulated samples that better capture the true distribution of the data. This allows the classification to achieve improved predictive performance on samples beyond the original distribution of the training data.

bioinformatics↗