bioRxiv Science⌕ Search

Biology subjects

Curabaz, N. N.

Publications and source records attributed to Curabaz, N. N..

2 recordsLinked to original sources

PrimeKG-Plus: a literature-derived expansion of a multimodal precision medicine knowledge graph

BackgroundDisease-centered knowledge graphs (KGs) support drug repurposing and precision medicine research, yet many remain static after release while primary databases and literature continue to expand. PrimeKG (Precision Medicine Knowledge Graph) is a widely used multimodal KG providing a holistic view of diseases. However, its public release reflects a June 2021 data cutoff, omitting several years of subsequent data growth. This lag is especially consequential for rare diseases, where mechanistic and therapeutic evidence often remains scattered across publications rather than structured resources. FindingWe present PrimeKG-Plus, a refreshed, rare-disease-enriched release of PrimeKG, rebuilt from all twenty original data resources updated to their December 2025 releases and three additional resources: OpenTargets, RepurposeDrugs, and nSIDES. Beyond synchronizing biomedical databases, PrimeKG-Plus captures approximately five years of previously unavailable rare-disease knowledge from the biomedical literature, curated from 637 PubMed abstracts and PubMed Central full-text articles using a language-model-assisted workflow focused on four rare neurological disorders: Canavan disease, Niemann-Pick disease type C, Tay-Sachs disease, and Batten disease. Extracted relations were refined through entity normalization, UMLS synonym mapping, embedding-based similarity ranking, and human expert review. Network topology analysis showed improved indirect drug-disease connectivity across three to six hops and added 447,288 drug-protein-disease paths linking previously unreachable drug-disease pairs. Temporal validation using drug approval records identified 55 molecular entities approved after the original PrimeKG June 2021 cutoff, 46 of which were absent from the original graph. ConclusionPrimeKG-Plus restores the temporal relevance of PrimeKG, providing an updated resource for drug repurposing, rare-disease research, and downstream machine-learning applications.

bioinformatics↗

Data-driven strategies for drug repurposing

Drug discovery is a complex, time-intensive, and costly process, often requiring more than a decade and substantial financial investment to bring a single therapeutic to market. Drug repurposing, the systematic identification of new indications for existing approved drugs, offers a cost-effective and expedited alternative to traditional pipelines, with the potential to address unmet clinical needs. In this study, we present a comparative analysis of drug-target interaction data from three extensively curated resources: ChEMBL, BindingDB, and GtoPdb, evaluating their release histories, curation methodologies, and coverage of approved and investigational compounds and targets. To facilitate therapeutic interpretation, we manually classified ChEMBL targets into 12 high-level biological families and mapped 817 clinically approved drug indications into 28 broader therapeutic groups. This structured framework enabled a systematic profiling of physicochemical properties among approved drugs across therapeutic categories. Our analyses revealed associations between physicochemical characteristics and therapeutic groups, providing practical guidance for indication-specific compound prioritization and refining the repurposing studies. We also examined cross-indication drug approvals to identify areas with high repurposing potential. Finally, we implemented a pathway-based computational pipeline to predict repositioning opportunities for FDA-approved drugs across ten major cancer types, demonstrating its adaptability to other disease contexts. Overall, this work consolidates drug-target data and computational repurposing into a data-driven framework that advances drug discovery and translational applications.

molecular biology↗