bioRxiv Science⌕ Search

Biology subjects

Zhang, J. Z. H.

Publications and source records attributed to Zhang, J. Z. H..

5 recordsLinked to original sources

Revolutionizing GPCR-Ligand Predictions: DeepGPCR with experimental Validation for High-Precision Drug Discovery

G-protein coupled receptors (GPCRs), crucial in various diseases, are targeted of over 40% of approved drugs. However, the reliable acquisition of experimental GPCRs structures is hindered by their lipid-embedded conformations. Traditional protein-ligand interaction models falter in GPCR-drug interactions, caused by limited and low-quality structures. Generalized models, trained on soluble protein-ligand pairs, are also inadequate. To address these issues, we developed two models, DeepGPCR_BC for binary classification and DeepGPCR_RG for affinity prediction. These models use non-structural GPCR-ligand interaction data, leveraging graph convolutional networks (GCN) and mol2vec techniques to represent binding pockets and ligands as graphs. This approach significantly speeds up predictions while preserving critical physical-chemical and spatial information. In independent tests, DeepGPCR_BC surpassed Autodock Vina and Schrodinger Dock with an AUC of 0.72, accuracy of 0.68, and TPR of 0.73, whereas DeepGPCR_RG demonstrated a Pearson correlation of 0.39 and RMSE of 1.34. We applied these models to screen drug candidates for GPR35 (Q9HC97), yielding promising results with 3 (F545-1970, K297-0698, S948-0241) out of 8 candidates. Furthermore, we also successfully obtained 6 active inhibitors for GLP-1R. Our GPCR-specific models pave the way for efficient and accurate large-scale virtual screening, potentially revolutionizing drug discovery in the GPCR field.

bioinformatics↗

ChemXTree: A Tree-enhanced Classification Approach to Small-molecule Drug Discovery

The rapid advancement of machine learning, particularly deep learning, has propelled significant strides in drug discovery, offering novel methodologies for molecular property prediction. However, despite these advancements, existing approaches often face challenges in effectively extracting and selecting relevant features from molecular data, which is crucial for accurate predictions. Our work introduces ChemXTree, a novel graph-based model that integrates tree-based algorithms to address these challenges. By incorporating a Gate Modulation Feature Unit (GMFU) for refined feature selection and a differentiable decision tree in the output layer. Extensive evaluations on benchmark datasets, including MoleculeNet and eight additional drug databases, have demonstrated ChemXTrees superior performance, particularly in feature optimization. Permutation experiments and ablation studies further validate the effectiveness of GMFU, positioning ChemXTree as a significant advancement in molecular informatics, capable of rivaling state-of-the-art models.

bioinformatics↗

DeepBindGCN: Integrating Molecular Vector Representation with Graph Convolutional Neural Networks for Accurate Protein-Ligand Interaction Prediction

The core of large-scale drug virtual screening is to accurately and efficiently select the binders with high affinity from large libraries of small molecules in which nonbinders are usually dominant. The protein pocket, ligand spatial information, and residue types/atom types play a pivotal role in binding affinity. Here we used the pocket residues or ligand atoms as nodes and constructed edges with the neighboring information to comprehensively represent the protein pocket or ligand information. Moreover, we find that the model with pre-trained molecular vectors performs better than the onehot representation. The main advantage of DeepBindGCN is that it is non-dependent on docking conformation and concisely keeps the spatial information and physical-chemical feature. Notably, the DeepBindGCN_BC has high precision in many DUD.E datasets, and DeepBindGCN_RG achieve a very low RMSE value in most DUD.E datasets. Using TIPE3 and PD-L1 dimer as proof-of-concept examples, we proposed a screening pipeline by integrating DeepBindGCN_BC, DeepBindGCN_RG, and other methods to identify strong binding affinity compounds. In addition, a DeepBindGCN_RG_x model has been used for comparing performance with other methods in PDBbind v.2016 and v.2013 core set. It is the first time that a non-complex dependent model achieves an RMSE value of 1.3843 and Pearson-R value of 0.7719 in the PDBbind v.2016 core set, showing comparable prediction power with the state-of-the-art affinity prediction models that rely upon the 3D complex. Our DeepBindGCN provides a powerful tool to predict the protein-ligand interaction and can be used in many important large-scale virtual screening application scenarios.

bioinformatics↗

Deep-learning based bioactive peptides generation and screening against Xanthine oxidase

In our previous work, we have developed LSTM_Pep to generate de novo potential active peptides by finetuning with known active peptides and developed DeepPep to effectively identify protein-peptide interaction. Here, we have combined LSTM_Pep and DeepPep to successfully obtained an active de novo peptide (ARG-ALA-PRO-GLU) of Xanthine oxidase (XOD) with IC50 value of 3.76mg/mL, and XOD inhibitory activity of 64.32%. Consistent with the experiment result, the peptide ARG-ALA-PRO-GLU has the highest DeepPep score, this strongly supports that we can generate de novo potential active peptides by finetune training LSTM_Pep over some known active peptides and identify those active peptides by DeepPep effectively. Our work sheds light on the development of deep learning-based methods and pipelines to effectively generate and obtain bioactive peptides with a specific therapeutic effect and showcases how artificial intelligence can help discover de novo bioactive peptides that can bind to a particular target.

bioinformatics↗

Deep-learning based bioactive therapeutic peptides generation and screening

Many bioactive peptides demonstrated therapeutic effects over-complicated diseases, such as antiviral, antibacterial, anticancer, etc. Similar to the generating de novo chemical compounds, with the accumulated bioactive peptides as a training set, it is possible to generate abundant potential bioactive peptides with deep learning. Such techniques would be significant for drug development since peptides are much easier and cheaper to synthesize than compounds. However, there are very few deep learning-based peptide generating models. Here, we have created an LSTM model (named LSTM_Pep) to generate de novo peptides and finetune learning to generate de novo peptides with certain potential therapeutic effects. Remarkably, the Antimicrobial Peptide Database has fully utilized in this work to generate various kinds of potential active de novo peptide. We proposed a pipeline for screening those generated peptides for a given target, and use Main protease of SARS-COV-2 as concept-of-proof example. Moreover, we have developed a deep learning-based protein-peptide prediction model (named DeepPep) for fast screening the generated peptides for the given targets. Together with the generating model, we have demonstrated iteratively finetune training, generating and screening peptides for higher predicted binding affinity peptides can be achieved. Our work sheds light on to the development of deep learning-based methods and pipelines to effectively generating and getting bioactive peptides with a specific therapeutic effect, and showcases how artificial intelligence can help discover de novo bioactive peptides that can bind to a particular target.

bioinformatics↗