bioRxiv Science⌕ Search

Biology subjects

Anyaegbunam, U. A.

Publications and source records attributed to Anyaegbunam, U. A..

2 recordsLinked to original sources

Cross-Domain Transfer Learning from Peptides to Lipids Using a Multi-Property Fine-Tuned LLM

Accurate knowledge of liquid chromatography retention time (RT) is essential for confident compound identification in metabolomics and lipidomics. Yet, it is often constrained by the scarcity of experimental data for many molecular classes. Current workflows depend on experimental RT libraries, which are time-consuming to build and limited to previously observed compounds. Here, we present a transfer learning pipeline that leverages large, publicly available peptide datasets to enable accurate lipid RT prediction in data-sparse scenarios. We first show that a ChemBERTa language model, when pre-trained on peptides with a multi-task objective (predicting both RT and fundamental RDKit molecular descriptors), learns a more robust and generalizable chemical representation than a single-task (RT-only) model. This multi-property pre-training yielded superior generalization in lipids, achieving test R2 values of 0.842 against 0.814 (RT-only). Crucially, transferring this peptide-based model to lipid data provided a pronounced advantage in data-sparse scenarios. When fine-tuned on only 5% of available lipid data, the transferred model improved the median test R2 by +0.234 over a model trained from scratch. Significant benefits persisted at intermediate data scales (50-75%), with performance converging only when 100% of the lipid data was used. Notably, the pre-trained model never underperformed the baseline, exhibiting more stable training across all data scales. These results demonstrate that multi-property pre-training guides language models towards chemically meaningful representations that support better RT prediction in different molecular domains. Furthermore, peptide-based pre-training facilitates cross-domain transfer of chemical properties to lipid. Our work provides a practical, scalable strategy to mitigate data scarcity in lipidomics by transferring knowledge from data-rich peptide databases, offering a computational alternative to extensive experimental library generation and enabling more confident identification in small-scale omics studies.

bioinformatics↗

Evaluating Genetic Regulators of MicroRNAs Using Machine Learning Models

This study explores the genetic regulators of microRNAs (miRNAs) using an ensemble of machine learning models to predict miRNA expression levels from gene expression data. Employing ridge regression, we accurately predicted the expression of 353 human miRNAs (R2 > 0.5), revealing robust miRNA-gene regulatory relationships. By analyzing the coefficients of these predictive models, we identified genetic regulators for each miRNA and highlighted the multifactorial nature of miRNA regulation. Further network analysis uncovered that miRNAs with higher predictive accuracy are more densely connected to their top predictive genes, reflecting strong regulatory control within miRNA-gene networks. To refine these insights, we filtered the gene-miRNA interaction networks to identify miRNAs specifically associated with enriched pathways, such as synaptic function and cardiovascular processes. From this pathway-centric analysis, we present a curated list of miRNAs and their genetic regulators, pinpointing their activity within distinct biological contexts. Additionally, our study provides a comprehensive set of metrics and coefficients for the genes most predictive of miRNA expression, along with a filtered subnetwork of miRNAs linked to specific pathways and phenotypes. By integrating miRNA expression predictors with network analysis and pathway enrichment, this work advances our understanding of miRNA regulatory mechanisms and their roles across distinct biological systems.

bioinformatics↗