bioRxiv Science⌕ Search

Biology subjects

Menghan, J.

Publications and source records attributed to Menghan, J..

2 recordsLinked to original sources

ACENet: a graph neural network for predicting enzyme pHmin using structure and sequence insights

The acid activity of enzymes, characterized by the minimum pH at which enzymes remain active (pHmin), is crucial for industrial applications in acidic environments. However, the rational design of acid-active enzymes remains challenging due to limited understanding of sequence-structure- activity relationships under acidic conditions. Here, we propose ACENet, a graph neural network that predicts enzyme pHmin by integrating surface features of protein structures with evolutionary representations derived from the large-scale protein language model ESM-2. ACENet achieved a Pearson correlation coefficient of 0.85 on the test dataset, significantly outperforming other deep learning baseline models and maintains stable pHmin predictions under various conditions. Even on a subset of the dataset with less than 20% homology, the PCC remains above 0.5, with an RMSE (Root mean square error) less than 1.4. ACENet also present excellent performance in the annotation of pHmin for homologous proteins and the predictive screening of minimal active pH in protein mutants. Remarkably, ACENet could identify the catalytic region as key determinants of acid activity through residue-level interpretability analysis. Overall, ACENet accelerates the development of highly efficient biocatalysts for diverse applications where acidic conditions predominate.

bioinformatics↗

HPClas: A data-driven approach for identifying halophilic proteins based on catBoost

Halophilic proteins possess unique structural properties and exhibit high stability under extreme conditions. Such distinct characteristic makes them invaluable for applications in various aspects such as bioenergy, pharmaceuticals, environmental clean-up and energy production. Generally, halophilic proteins are discovered and characterized through labor-intensive and time-consuming wetlab experiments. Here, we introduced HPClas, a machine learning-based classifier developed using the catBoost ensemble learning technique to identify halophilic proteins. Extensive in silico calculations were conducted on a large public data set of 12574 samples and an independent test set of 200 sample pairs, on which HPClas achieved an AUROC of 0.877 and 0.845, respectively. The source code and curated data set of HPClas are publicly available at https://github.com/Showmake2/HPClas. In conclusion, HPClas can be explored as a promising tool to aid in the identification of halophilic proteins and accelerate their applications in different fields. Impact StatementIn this study, we used a method based on prediction of proteins secreted by extreme halophilic bacteria to successfully extract a large number of halophilic proteins. Using this data, we have trained an accurate halophilic protein classifier that could determine whether an input protein is halophilic with a high accuracy of 84.5%. This research could not only promote the exploration and mining of halophilic proteins in nature, but also provide guidance for the generation of mutant halophilic enzymes.

bioinformatics↗