bioRxiv Science⌕ Search

Biology subjects

Urbaniak, K.

Publications and source records attributed to Urbaniak, K..

3 recordsLinked to original sources

Integrating computational chemistry and machine learning to predict KRAS mutation-induced resistance

Mutation-induced drug resistance is a major contributor to the failure of targeted cancer therapies, particularly in tumors driven by mutations in the KRAS oncogene. Although covalent inhibitors effectively target KRAS G12C, secondary mutations such as G12C/Y96C, G12C/Y96S, and G12C/Y96D lead to resistance despite leaving the covalent attachment site intact. To predict these resistance outcomes, we developed a computational framework that integrates molecular dynamics-derived structural, energetic, thermodynamic, and contact-based descriptors with machine learning. Features extracted from simulations of treatment-sensitive and treatment-resistant KRAS mutants were used to train logistic regression, random forest, support vector machine, and Bayesian Network classifiers, achieving average accuracies above 90%. Solvent-accessible surface area variability, Lennard-Jones 1,4 energy, mean square displacement, and root mean square fluctuation emerged as the most discriminatory features. Residues G10, E62, and H95 showed the highest predictive value. This approach highlights conformational and solvent-exposure changes as central drivers of KRAS drug resistance and provides a generalizable workflow for other clinically relevant mutant targets. Author SummaryMutation-induced resistance is a common challenge across many cancer types and is often associated with aggressive tumor progression and poor therapeutic response. Investigating the dynamic properties of proteins harboring such mutations provides valuable insights into the structural and functional consequences of these alterations, thereby helping to elucidate the mechanisms of drug resistance. Machine learning algorithms are particularly effective at uncovering complex patterns within high-dimensional data, such as molecular dynamics simulation trajectories. Integrating these algorithms with analysis of protein dynamics holds significant potential to aid in drug discovery challenges by reducing both time and resource demands while increasing the likelihood of identifying effective therapeutic candidates. As a proof of concept, we developed a computational framework that integrates molecular dynamics-derived molecular features with machine learning to distinguish treatment-sensitive from treatment-resistant KRAS mutants. KRAS is known for drug resistance arising from secondary mutations that disrupt inhibitor binding despite intact covalent attachment sites. The models achieved over 90% accuracy and identified solvent-exposure and conformational changes at residues G10, E62, and H95 as key predictors of treatment resistance. This workflow offers a generalizable strategy for understanding and forecasting mutation-induced resistance.

biophysics↗

Ligand Discrimination in Immune Cells: Signal Processing Insights into Immune Dysfunction in ER+ Breast Cancer

Prior studies have shown that approximately 40% of estrogen receptor positive (ER+) breast cancer (BC) patients harbor immune signaling defects in their blood at diagnosis, and the presence of these defects predicts overall survival. Therefore, it is of interest to quantitatively characterize and measure signaling errors in immune signaling systems in these patients. Here we propose a novel approach combining communication theory and signal processing concepts to model ligand discrimination in immune cells in the peripheral blood. We use the model to measure the specificity of ligand discrimination in the presence of molecular noise by estimating the probability of error, which is the probability of making a wrong ligand identification. We apply our model to the JAK/STAT signaling pathway using high dimensional spectral flow cytometry measurements of transcription factors, including phosphorylated STATs and SMADs, in immune cells stimulated with several cytokines (IFN{gamma}, IL-2, IL-6, IL-4, and IL-10) from 19 ER+ breast cancer patients and 32 healthy controls. In addition, we apply our model to 10 healthy donor samples treated with a clinically approved JAK1/2 inhibitor. Our results show reduced ligand identification accuracy and higher levels of molecular noise in BC patients as compared to healthy controls, which may indicate altered immune signaling and the potential for immune cell dysfunction in these patients. Moreover, the inhibition of JAK1/2 produces ligand misidentification and molecular noise rates similar to, or even greater than, those observed in breast cancer. These results suggest a means to improve the use of signaling kinase inhibitor therapies by identifying patients with favorable ligand discrimination specificity profiles in their immune cells. One Sentence SummaryWe use a communication model to measure ligand discrimination errors in immune cells and molecular noise in cytokine signaling in ER+ breast cancer patients as compared to healthy controls.

cancer biology↗

BaNDyT: Bayesian Network modeling of molecular Dynamics Trajectories

Bayesian network modeling (BN modeling, or BNM) is an interpretable machine learning method for constructing probabilistic graphical models from the data. In recent years, it has been extensively applied to diverse types of biomedical datasets. Concurrently, our ability to perform long-timescale molecular dynamics (MD) simulations on proteins and other materials has increased exponentially. However, the analysis of MD simulation trajectories has not been data-driven but rather dependent on the users prior knowledge of the systems, thus limiting the scope and utility of the MD simulations. Recently, we pioneered using BNM for analyzing the MD trajectories of protein complexes. The resulting BN models yield novel fully data-driven insights into the functional importance of the amino acid residues that modulate proteins function. In this report, we describe the BaNDyT software package that implements the BNM specifically attuned to the MD simulation trajectories data. We believe that BaNDyT is the first software package to include specialized and advanced features for analyzing MD simulation trajectories using a probabilistic graphical network model. We describe here the softwares uses, the methods associated with it, and a comprehensive Python interface to the underlying generalist BNM code. This provides a powerful and versatile mechanism for users to control the workflow. As an application example, we have utilized this methodology and associated software to study how membrane proteins, specifically the G protein-coupled receptors, selectively couple to G proteins. The software can be used for analyzing MD trajectories of any protein as well as polymeric materials.

bioinformatics↗