bioRxiv Science⌕ Search

Biology subjects

Ma, W.-Y.

Publications and source records attributed to Ma, W.-Y..

3 recordsLinked to original sources

Peptide Design through Binding Interface Mimicry

Peptides offer distinct advantages for targeted therapy, including oral bioavailability, cellular permeability, and high specificity, which set them apart from conventional small molecules and biologics. In this work, we developed an AI algorithm, named PepMimic, to transform a known protein receptor or an existing antibody of a target into a short peptide drug by mimicking the binding interfaces between targets and known binders. The structural root mean square deviation and interface DockQ with reference binders were 61% and 75% better than the best existing methods on the PepBench datasets. We then applied this novel peptide-design methodology to five drug targets: PD-L1, CD38, BCMA, HER2, and CD4. SPRi results show that 8% of the peptides exhibited dissociation constant (KD) values at the 10-8M level, and 26 peptides achieving KD values as low as 10-9M. This success rate was 20,000 times higher than that observed in a random library screening conducted under identical conditions. PepMimic was applied to target proteins lacking available binders by first utilizing AI algorithms to design protein binders, followed by the generation of peptides through simulation of these artificial interfaces. The top-ranked peptides underwent extensive cellular validation and in vivo testing through tail vein injections in breast, myeloma, and lung tumor mouse models. Experimental results demonstrated effective membrane binding and highlighted the strong potential of these peptides for clinical diagnostic imaging and targeted therapeutic applications.

biochemistry↗

MOL-AE: Auto-Encoder Based Molecular Representation Learning With 3D Cloze Test Objective

3D molecular representation learning has gained tremendous interest and achieved promising performance in various downstream tasks. A series of recent approaches follow a prevalent framework: an encoder-only model coupled with a coordinate denoising objective. However, through a series of analytical experiments, we prove that the encoderonly model with coordinate denoising objective exhibits inconsistency between pre-training and downstream objectives, as well as issues with disrupted atomic identifiers. To address these two issues, we propose MO_SCPLOWOLC_SCPLOW-AE for molecular representation learning, an auto-encoder model using positional encoding as atomic identifiers. We also propose a new training objective named 3D Cloze Test to make the model learn better atom spatial relationships from real molecular substructures. Empirical results demonstrate that MO_SCPLOWOLC_SCPLOW-AE achieves a large margin performance gain compared to the current state-of-the-art 3D molecular modeling approach. The source codes of MO_SCPLOWOLC_SCPLOW-AE are publicly available at https://github.com/yjwtheonly/MolAE.

bioinformatics↗

Multi-Scale Protein Language Model for Unified Molecular Modeling

Protein language models have demonstrated significant potential in the field of protein engineering. However, current protein language models primarily operate at the residue scale, which limits their ability to provide information at the atom level. This limitation prevents us from fully exploiting the capabilities of protein language models for applications involving both proteins and small molecules. In this paper, we propose ESM-AA (ESM All-Atom), a novel approach that enables atom-scale and residue-scale unified molecular modeling. ESM-AA achieves this by pretraining on multi-scale code-switch protein sequences and utilizing a multi-scale position encoding to capture relationships among residues and atoms. Experimental results indicate that ESM-AA surpasses previous methods in proteinmolecule tasks, demonstrating the full utilization of protein language models. Further investigations reveal that through unified molecular modeling, ESM-AA not only gains molecular knowledge but also retains its understanding of proteins.1

bioinformatics↗