bioRxiv Science⌕ Search

Biology subjects

Madsen, N. G.

Publications and source records attributed to Madsen, N. G..

3 recordsLinked to original sources

ProteusAI: An Open-Source and User-Friendly Platform for Machine Learning-Guided Protein Design and Engineering

AO_SCPLOWBSTRACTC_SCPLOWProtein design and engineering are crucial for advancements in biotechnology, medicine, and sustainability. Machine learning (ML) models are used to design or enhance protein properties such as stability, catalytic activity, and selectivity. However, many existing ML tools require specialized expertise or lack open-source availability, limiting broader use and further development. To address this, we developed ProteusAI, a user-friendly and open-source ML platform to streamline protein engineering and design tasks. ProteusAI offers modules to support researchers in various stages of the design-build-test-learn (DBTL) cycle, including protein discovery, structure-based design, zero-shot predictions, and ML-guided directed evolution (MLDE). Our benchmarking results demonstrate ProteusAIs efficiency in improving proteins and enyzmes within a few DBTL-cycle iterations. ProteusAI democratizes access to ML-guided protein engineering and is freely available for academic and commercial use. Future work aims to expand and integrate novel methods in computational protein and enzyme design to further develop ProteusAI.

bioinformatics↗

Harnessing Chemical Space Neural Networks to Systematically Annotate GPCR ligands

Machine learning (ML) has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors (hGPCRs) in FDA-approved drugs, exhaustive in-distribution drug-target interaction (DTI) testing across all pairs of hGPCRs and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution (OOD) exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks (CSNN) that leverages network homophily and training-free graph neural networks (GNNs) with Labels as Features (LaF). We show that CSNNs ability to make accurate predictions strongly correlates with network homophily. Thus, LaFs strongly increase a ML models capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 DTIs, 539 compounds, 7 hGPCRs) to discover novel DTIs for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

bioinformatics↗

De novo Design of a Polycarbonate Hydrolase

Enzymatic degradation of plastics is currently limited to the use of engineered natural enzymes. As of yet, all engineering approaches applied to plastic degrading enzymes retain the natural /{beta} -fold. While mutations can be used to increase thermostability, an inherent maximum likely exists for the /{beta} -fold. It is thus of interest to introduce catalytic activity toward plastics in a different protein fold to escape the sequence space of plastic degrading enzymes. Here, a method for designing highly thermostable enzymes that can degrade plastics is described. This has been used to design an enzyme that can catalyze the hydrolysis of polycarbonate, which no known natural enzymes can degrade. Rosetta enzyme design is used to introduce a catalytic triad into a set of thermostable scaffolds. Through computational evaluation, a potential PCase was selected and produced recombinantly in E. coli. CD spectroscopy suggests that the design has a melting temperature of >95{degrees}C. Activity towards a commercially used polycarbonate (Makrolon 2808) was confirmed using AFM, which showed that a PCase had been designed successfully. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=58 SRC="FIGDIR/small/532063v1_ufig1.gif" ALT="Figure 1"> View larger version (15K): org.highwire.dtl.DTLVardef@1369436org.highwire.dtl.DTLVardef@3c87b3org.highwire.dtl.DTLVardef@1f1138forg.highwire.dtl.DTLVardef@3b4bba_HPS_FORMAT_FIGEXP M_FIG C_FIG

synthetic biology↗