bioRxiv Science⌕ Search

Biology subjects

Qing, R.

Publications and source records attributed to Qing, R..

2 recordsLinked to original sources

Recursive Cleaning for Large-scale Protein Data via Multimodal Learning

AO_SCPLOWBSTRACTC_SCPLOWReliable datasets and high-performance models work together to drive significant advancements in protein representation learning in the era of Artificial Intelligence. The size of protein models and datasets has grown exponentially in recent years. However, the quality of protein knowledge and model training has suffered from the lack of accurate and efficient data annotation and cleaning methods. To address this challenge, we introduce ProtAC, which corrects large Protein datasets with a scalable Automatic Cleaning framework that leverages both sequence and functional information through multimodal learning. To fulfill data cleaning, we propose the Sequence-Annotation Matching (SAM) module in the model, which filters the functional annotations that are more suitable for the corresponding sequences. Our approach is a cyclic process consisting of three stages: first pretraining the model on a large noisy dataset, then finetuning the model on a small manually annotated dataset, and finally cleaning the noisy dataset using the finetuned model. Through multiple rounds of "train-finetune-clean" cycles, we observe progressive improvement in protein function prediction and sequenceannotation matching. As a result, we achieve (1) a state-of-the-art (SOTA) model that outperforms competitors with fewer than 100M parameters, evaluated on multiple function-related downstream tasks, and (2) a cleaned UniRef50 dataset containing [~]50M proteins with well-annotated functions. Performing extensive biological analysis on a cleaned protein dataset, we demonstrate that our model is able to understand the relationships between different functional annotations in proteins and that proposed functional annotation revisions are reasonable.

bioengineering↗

PRO-LDM: Protein Sequence Generation with Conditional Latent Diffusion Models

AbstractsDeep learning-driven protein design holds enormous potential despite the complexities in sequences and structures. Recent developments in diffusion models yielded success in structure design, but awaits progress in sequence design and are computationally demanding. Here we present PRO-LDM: an efficient framework combining design fidelity and computational efficiency, utilizing the diffusion model in latent space to design proteins with property tuning. The model employs a joint autoencoder to capture latent variable distributions and generate meaningful embeddings from sequences. PRO-LDM (1) learns representations from biological features in natural proteins at both amino-acid and sequence level; (2) generates native-like new sequences with enhanced diversity; and (3) conditionally designs new proteins with tailored properties or functions. The out-of-distribution design enables sampling notably different sequences by adjusting classifier guidance strength. Our model presents a feasible pathway and an integratable tool to extract physicochemical and evolutionary information embedded within primary sequences, for protein design and optimization.

bioengineering↗