Overcoming Topology Bias and Cold-Start Limita-tions in Drug Repurposing: A Clinical-Outcome-Aligned LLM Framework
Graph Neural Networks (GNNs) in drug repurposing suffer from two limitations: transductive failure in zero-shot (cold-start) scenarios and popularity bias that misidentifies high-degree nodes as effective drugs. We propose a framework that shifts the optimization objective from graph topology to clinical utility, integrating Knowledge Graph RAG (KG-RAG), Supervised Fine-Tuning (SFT), and Kahneman-Tversky Optimization (KTO) using Phase III clinical trial outcomes as rewards. We evaluated our approach on a rigorous 1:10 negative sampling benchmark derived from MiRAGE, covering Standard, Cold-Start, and Degree-Matched settings. In Cold-Start scenarios where topological signals are absent, traditional GNNs (including TxGNN) collapse (Top-10 Precision < 0.30), whereas DR-SFT model achieves 0.80, demonstrating robust semantic generalization for novel compounds. Crucially, in Degree-Matched tests dominated by "popular but ineffective" decoys, the DR-KTO acts as a clinical gatekeeper, achieving 0.90 Top-10 Precision--significantly outperforming DR-SFT (0.70) and GNNs (0.2-0.4) by effectively penalizing hard negatives. Beyond repurposing accuracy, the model achieves state-of-the-art performance on BioASQ, and increased ability in Chemprot. Orthogonal validation via DrugReAlign confirms physical plausibility, yielding significantly lower docking binding energies for KTO-recommended candidates. SPR experiments further corroborate these findings. By aligning LLM reasoning with clinical evidence, our framework successfully bridges the gap between semantic inference, topological structure, and clinical reality.