bioRxiv Science⌕ Search

Biology subjects

Shrimpton-Phoenix, E.

Publications and source records attributed to Shrimpton-Phoenix, E..

4 recordsLinked to original sources

MHChron: diversity-balanced dataset design for robust peptide-MHC binding prediction across MHC class I and II

Accurate prediction of peptide-MHC (pMHC) binding is central to immunogenicity assessment, yet many existing predictors are trained and evaluated on narrow allele sets and restricted peptide-lengths. Here, we present MHChron, a unified pMHC binding prediction framework predicated on systematic data curation, meticulous engineering of dataset balance and diversity, and rigorous evaluation through careful splits controlling for data leakage. We assemble one of the most diverse pMHC training dataset reported to date, integrating publicly available binding data across a broad allele coverage (class I n=214, class II n=98) and peptide length range (from 8 to 36 residues). Using a focused and carefully sampled subset of this dataset, we train complementary sequence-based and structure-aware models and test them under increasingly stringent generalisation regimes. Both models achieve consistently strong performance, outperforming the evaluated state-of-the-art predictors despite being trained on numerically fewer data points. Notably, the structure-aware model did not consistently surpass the sequence-based model, except under the most demanding setting of extrapolation to unseen allele clusters, suggesting that performance gains stem primarily from dataset diversity and rigorous evaluation rather than architectural complexity. Sequence-based MHChron is released with reproducible installation and an automated whole-protein screening pipeline, enabling broad and practical use.

bioinformatics↗

drFrankenstein: An Automated Pipeline for the Parameterisation of Non-Canonical Amino Acids

The incorporation of non-canonical amino acids (ncAAs) is a powerful strategy for introducing novel chemical functions into proteins. Molecular dynamics (MD) simulations are essential for understanding the structural and dynamic effects of these modifications, yet the creation of accurate force field parameters for ncAAs remains a significant bottleneck. Current parameterisation methods are often inaccurate or computationally expensive. To address this, we present drFrankenstein, an automated pipeline for generating AMBER force field parameters for ncAAs. drFrankenstein is a robust and accessible tool that streamlines the parameterisation workflow, enabling the routine use of MD simulations to study the behaviour of ncAA-containing proteins.

bioinformatics↗

Triplet Quenching by Active Site Cysteine Residues Improves Photostability in Fatty Acid Photodecarboxylase

Enzyme photobiocatalysis uses light to drive high-energy transformations but is limited by the rarity of photoenzymes. Fatty acid photodecarboxylase (FAP), a recently discovered photoenzyme, enables fatty acid conversion to alkanes/alkenes via excitation of an FAD cofactor, though its poor photostability and photoinactivation has hindered industrial applications. Here, we combine protein engineering approaches with biocatalytic and biophysical techniques, as well as computational chemistry, to demonstrate that additional active site cysteine residues can suppress oxygen-mediated inactivation processes that are driven by the FAD triplet-excited state. We identify a number of positions close to the FAD for cysteine residues that lead to a significant enhancement in activity as a result of an increase in the number of catalytic turnovers and improved photostability. The additional cysteine residues quench the triplet excited state of the FAD cofactor via a proposed proton-coupled electron transfer mechanism, resulting in lower levels of harmful reactive oxygen species. Our study highlights promising routes to mitigate non-productive, photoinactivation pathways in FAP and informs the rational design of new flavin-based photoenzymes.

biochemistry↗

drMD: Molecular Dynamics for Experimentalists

Molecular dynamics (MD) simulations can be used by protein scientists to investigate a wide array of biologically relevant properties such as the effects of mutations on a proteins structure and activity, or probing intermolecular interactions with small molecule substrates or other macromolecules. Within the world of computational structural biology, several programs have become popular for running these simulations, but each of these programs requires a significant time investment from the researcher to run even simple simulations. Even after learning how to run and analyse simulations, many elements remain a "black box." This greatly limits the accessibility of molecular dynamics simulations for non-experts. Here we present drMD, an automated pipeline for running molecular dynamics simulations using the OpenMM molecular mechanics toolkit. We have created drMD with non-experts in computational biology in mind. The drMD codebase has several functions that automatically handle routine procedures associated with running molecular dynamics simulations. This greatly reduces the expertise required to run MD simulations. We have also introduced a series of quality-of-life features to make the process of running MD simulations both easier and more pleasant. Finally, drMD explains the steps it is taking interactively and, where useful, provides relevant references so the user can learn more. All these features make drMD an effective tool for learning molecular dynamics while running publication-quality simulations. drMD is open source and can be found on GitHub: https://github.com/wells-wood-research/drMD.

bioinformatics↗