bioRxiv Science⌕ Search

Biology subjects

Gasser, H.-C.

Publications and source records attributed to Gasser, H.-C..

3 recordsLinked to original sources

A novel decoding strategy for ProteinMPNN to design with less MHC Class I immune-visibility

Due to their versatility and diverse production methods, proteins have attracted a lot of interest for industrial as well as therapeutic applications. Designing new therapeutics requires careful consideration of immune responses, particularly the cytotoxic T-lymphocyte (CTL) reaction to intra-cellular proteins. In this study, we introduce CAPE-Beam, a novel decoding strategy for the established ProteinMPNN protein design model. Our approach minimizes CTL immunogenicity risk by limiting designs to only consist of kmers that are either predicted not to be presented to CTLs or are subject to central tolerance. We compare CAPE-Beam to greedily sampling from ProteinMPNN and CAPE-MPNN. We find that our novel decoding strategy can produce structurally similar proteins while incorporating more human like kmers. This significantly lowers CTL immunogenicity risk in precision medicine, and represents a key step towards reducing this risk in protein therapeutics targeting a wider patient population.

bioinformatics↗

Integrating MHC Class I visibility targets into the ProteinMPNN protein design process

ProteinMPNN is crucial in many protein design pipelines, identifying amino acid (AA) sequences that fold into given 3D protein backbone structures. We explore ProteinMPNN in the context of designing therapeutic proteins that need to avoid triggering unwanted immune reactions. More specifically, we focus on intra-cellular proteins that face the challenge of evading detection by Cytotoxic T-lymphocytes (CTLs) that detect their presence via the MHC Class I (MHC-I) pathway. To reduce visibility of the designed proteins to this immune-system component, we develop a framework that uses the large language model (LLM) tuning method, Direct Preference Optimization (DPO), to guide ProteinMPNN in minimizing the number of predicted MHC-I epitopes in its designs. Our goal is to design proteins with low MHC-I immune-visibility while preserving the original structure and function. For our assessment, we first use AlphaFold to predict the 3D structures of designed protein sequences. We then use TM-score, that measures the structural alignment between the predicted design and original protein, to evaluate fidelity to the original protein structure. We find our LLM-based tuning method for constraining MHC-I visibility is able to effectively reduce visibility without compromising structural similarity to the original protein.

bioinformatics↗

Utility of language model and physics-based approaches in modifying MHC Class-I immune-visibility for the design of vaccines and therapeutics

Proteins have an arsenal of medical applications that include disrupting protein interactions, acting as potent vaccines, and replacing genetically deficient proteins. While therapeutics must avoid triggering unwanted immune-responses, vaccines should support a robust immune-reaction targeting a broad range of pathogen variants. Therefore, computational methods modifying proteins immunogenicity without disrupting function are needed. While many components of the immune-system can be involved in a reaction, we focus on Cytotoxic T-lymphocytes (CTLs). These target short peptides presented via the MHC Class I (MHC-I) pathway. To explore the limits of modifying the visibility of those peptides to CTLs within the distribution of naturally occurring sequences, we developed a novel machine learning technique, CAPE-XVAE. It combines a language model with reinforcement learning to modify a proteins immune-visibility. Our results show that CAPE-XVAE effectively modifies the visibility of the HIV Nef protein to CTLs. We contrast CAPE-XVAE to CAPE-Packer, a physics-based method we also developed. Compared to CAPE-Packer, the machine learning approach suggests sequences that draw upon local sequence similarities in the training set. This is beneficial for vaccine development, where the sequence should be representative of the real viral population. Additionally, the language model approach holds promise for preserving both known and unknown functional constraints, which is essential for the immune-modulation of therapeutic proteins. In contrast, CAPE-Packer, emphasizes preserving the proteins overall fold and can reach greater extremes of immune-visibility, but falls short of capturing the sequence diversity of viral variants available to learn from. Source code: https://github.com/hcgasser/CAPE (Tag: CAPE 1.1)

bioinformatics↗