bioRxiv Science⌕ Search

Biology subjects

Abdelmessih, M.

Publications and source records attributed to Abdelmessih, M..

3 recordsLinked to original sources

Representation Learning of Human Disease Mechanisms for a Foundation Model in Rare and Common Diseases

A fundamental challenge in translational medicine is the computational modeling of complex human diseases to accelerate therapeutic development. Representation learning provides a powerful framework to address this, yet creating models that capture deep biological mechanisms remains a critical need. To this end, we propose a novel strategy that partitions the disease landscape into rare and non-rare categories, enabling systematic knowledge repurposing both within and between these groups. Here, we introduce Dis2Vec (Disease to Vector), a representation learning framework designed to operationalize this concept. Dis2Vec generates biologically grounded disease embeddings by learning from human genetic and phenotypic data, forming the foundation for Disease-Disease Association Learning (DDAL) and unsupervised disease clustering. We evaluate Dis2Vec representations in two downstream applications. First, we assess DDAL performance on a transfer learning benchmark designed to predict therapeutic transferability, using real-world drug repurposing investment decisions made in clinical trials. Second, unsupervised clustering analyses reveal shared biological mechanisms across diseases. By modeling the disease landscape in this way, Dis2Vec enhances translational research efficiency across both rare and non-rare diseases, advancing the development of foundational models for therapeutic science. Furthermore, Dis2Vec establishes a biologically grounded disease-representation and benchmarking layer that paves the way for trustworthy agentic biomedical AI systems in rare-disease indication expansion.

systems biology↗

Novel cell states arise in embryonic cells devoid of key reprogramming factors

The capacity for embryonic cells to differentiate relies on a large-scale reprogramming of the oocyte and sperm nucleus into a transient totipotent state. In zebrafish, this reprogramming step is achieved by the pioneer factors Nanog, Pou5f3, and Sox19b (NPS). Yet, it remains unclear whether cells lacking this reprogramming step are directed towards wild type states or towards novel developmental canals in the Waddington landscape of embryonic development. Here we investigate the developmental fate of embryonic cells mutant for NPS by analyzing their single-cell gene expression profiles. We find that cells lacking the first developmental reprogramming steps can acquire distinct cell states. These states are manifested by gene expression modules that result from a failure of nuclear reprogramming, the persistence of the maternal program, and the activation of somatic compensatory programs. As a result, most mutant cells follow new developmental canals and acquire new mixed cell states in development. In contrast, a group of mutant cells acquire primordial germ cell-like states, suggesting that NPS-dependent reprogramming is dispensable for these cell states. Together, these results demonstrate that developmental reprogramming after fertilization is required to differentiate most canonical developmental programs, and loss of the transient totipotent state canalizes embryonic cells into new developmental states in vivo.

developmental biology↗

Topology-Driven Negative Sampling Enhances Generalizability in Protein-Protein Interaction Prediction

Unraveling the human interactome to uncover disease-specific patterns and discover drug targets hinges on accurate protein-protein interaction (PPI) predictions. However, challenges persist in machine learning (ML) models due to a scarcity of quality hard negative samples, shortcut learning, and limited generalizability to novel proteins. Here, we introduce a novel approach for strategic sampling of protein-protein non-interactions (PPNIs) by leveraging higher-order network characteristics that capture the inherent complementarity-driven mechanisms of PPIs. Next, we introduce UPNA-PPI (Unsupervised Pre-training of Node Attributes tuned for PPI), a high throughput sequence-to-function ML pipeline, integrating unsupervised pretraining in protein representation learning with topological PPNI samples, capable of efficiently screening billions of interactions. UPNA-PPI improves PPI prediction generalizability and interpretability, particularly in identifying potential binding sites locations on amino acid sequences, strengthening the prioritization of screening assays and facilitating the transferability of ML predictions across protein families and homodimers. UPNA-PPI establishes the foundation for a fundamental negative sampling methodology in graph machine learning by integrating insights from network topology.

bioinformatics↗