bioRxiv Science⌕ Search

Biology subjects

de Sa, A. G. C.

Publications and source records attributed to de Sa, A. G. C..

3 recordsLinked to original sources

Transfer learning across molecular graphs for predicting protein-ligand affinities and their changes upon mutations

Predicting protein-ligand binding affinity and mutation-induced affinity changes ({Delta}{Delta}G) remains challenging due to limited data and complex interaction mechanisms. Here we present DDMuffin, a deep learning framework that integrates structural, sequence, and interaction graph features, employing transfer learning and stringent dataset partitioning to achieve reliable generalization. DDMuffin demonstrates strong predictive accuracy on the rigorous LP-PDBBind benchmark (Pearson r up to 0.70 after excluding top 10% outliers; RMSE = 1.48 kcal mol-{superscript 1}). In evaluating mutation-induced affinity changes for clinically relevant kinase inhibitors, DDMuffin achieves competitive average performance (mean Spearman{rho} = 0.39), outperforming or matching several established methods. The approach provides valuable insights into mutation-driven drug resistance and interpretable ligand binding mechanisms, particularly advantageous for guiding inhibitor design and personalized therapeutic strategies. To facilitate broad application, we deployed DDMuffin as an accessible web server at https://biosig.lab.uq.edu.au/ddmuffin/, promoting systematic exploration of protein-ligand interactions in drug discovery research.

bioinformatics↗

AI-m6ARS: Machine learning-driven m6A RNA methylation site discovery with integrated sequence, conservation, and geographical descriptors

N6-Methyladenosine (m6A) is a predominant type of human RNA methylation, regulating diverse biochemical processes and being associated with the development of several diseases. Despite its significance, an extensive experimental examination across diverse cellular and transcriptome contexts is still lacking due to time and cost constraints. Computational models have been proposed to prioritise potential m6A methylation sites, although having limited predictive performance due to inadequate characterisation and modelling of m6A sites. This work presents AI-m6ARS, a novel model that utilises integrated sequence, conservation, and geographical descriptive features to predict human m6A methylation sites. The model was trained using the Light Gradient Boosting Machine (LightGBM) algorithm, which was coupled with comprehensive feature selection to improve the data quality. AI-m6RS demonstrates strong predictive capabilities, achieving an impressive area under the receiver operating characteristic curve of 0.87 on cross-validation. Consistent results on unseen transcripts in a blind test highlight the AI-m6ARS generalisability. AI-m6ARS also demonstrates comparable performance to state-of-the-art models, but offers two significant benefits: the model interpretability and the availability of a user-friendly web server. The AI-m6ARS web server offers valuable insights into the distribution of m6A sites within the human genome, thereby facilitating progress in medical applications. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=77 SRC="FIGDIR/small/599439v1_ufig1.gif" ALT="Figure 1"> View larger version (13K): org.highwire.dtl.DTLVardef@12d7502org.highwire.dtl.DTLVardef@15cf6b5org.highwire.dtl.DTLVardef@490699org.highwire.dtl.DTLVardef@5046c1_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

From Mutations to Disease: Computational Analysis and Interpretation of GPCR-Associated Pathogenicity

G protein-coupled receptors (GPCRs) perform critical roles in numerous physiological processes and their mutations are, therefore, associated with various human diseases. Hence, understanding the molecular consequences of pathogenic mutations in GPCRs is essential for elucidating disease mechanisms and developing effective therapeutic strategies. In this study, we employed computational approaches to explore the impact of mutations on GPCRs using two distinct datasets: ClinVar and MutHTP. We first evaluated the performance of available pathogenicity predictors. Beyond that, we used statistical analysis to identify key characteristics of mutations in GPCRs leading to diseases. We first evaluated available computational predictors, such as SIFT, PolyPhen-2, PROVEAN, ESM1b, and AlphaMissense in classifying GPCR mutations. During this task, we observed that all predictors performed with reliability when assessing GPCR mutations leading to diseases in the ClinVar dataset. On the other hand, when dealing with the MutHTP dataset, all predictors demonstrated poor performance, emphasising the importance of dataset characteristics and the need for comprehensive evaluation when selecting mutation predictive tools for GPCR analysis. The statistical analysis of mutations on GPCRs and disease development suggests that mutations occurring in conserved regions or regions with stronger intermolecular interactions are more likely to disrupt protein function and contribute to disease pathogenesis. Additionally, regarding our analysis, we also obtained insights into the importance of hydrophobic interactions and hydrogen bonding patterns in mutations in GPCRs and pathogenicity. Overall, our study enhances our understanding of the molecular mechanisms underlying GPCR-associated diseases and provides valuable insights for future research and clinical diagnostics in this field.

bioinformatics↗