bioRxiv Science⌕ Search

bioRxiv · 10.1101/2024.08.06.606224

VUStruct: a compute pipeline for high throughput and personalized structural biology

Abstract

Effective diagnosis and treatment of rare genetic disorders requires the interpretation of a patients genetic variants of unknown significance (VUSs). Today, clinical decision-making is primarily guided by gene-phenotype association databases and DNA-based scoring methods. Our web-accessible variant analysis pipeline, VUStruct, supplements these established approaches by deeply analyzing the downstream molecular impact of variation in context of 3D protein structure. VUStructs growing impact is fueled by the co-proliferation of protein 3D structural models, gene sequencing, compute power, and artificial intelligence. Contextualizing VUSs in protein 3D structural models also illuminates longitudinal genomics studies and biochemical bench research focused on VUS, and we created VUStruct for clinicians and researchers alike. We now introduce VUStruct to the broad scientific community as a mature, web-facing, extensible, High-Performance Computing (HPC) software pipeline. VUStruct maps missense variants onto automatically selected protein structures and launches a broad range of analyses. These include energy-based assessments of protein folding and stability, pathogenicity prediction through spatial clustering analysis, and machine learning (ML) predictors of binding surface disruptions and nearby post-translational modification sites. The pipeline also considers the entire input set of VUS and identifies genes potentially involved in digenic disease. VUStructs utility in clinical rare disease genome interpretation has been demonstrated through its analysis of over 175 Undiagnosed Disease Network (UDN) Patient cases. VUStruct-leveraged hypotheses have often informed clinicians in their consideration of additional patient testing, and we report here details from two cases where VUStruct was key to their solution. We also note successes with academic research collaborators, for whom VUStruct has informed research directions in both computational genomics and wet lab studies.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Moth, C. W., Sheehan, J. H., Mamun, A. A., Sivley, R. M., Gulsevin, A., Rinker, D., Capra, J. A., Meiler, J.. 2024-08-07. VUStruct: a compute pipeline for high throughput and personalized structural biology. https://doi.org/10.1101/2024.08.06.606224

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

MgATP/MgADP-dependent conformational dynamics and intrinsically disordered regions of vascular KATP channels revealed by cryoEM

Vascular smooth muscle KATP channels, composed of the pore-forming Kir6.1 and regulatory SUR2B subunits, control vascular tone, dysfunction of which causes systemic disease. Vascular KATP is regulated by Mg-nucleotides, but the underlying structural mechanism has remained elusive. Here, we determined cryoEM structures of these channels in the presence of MgATP and MgADP. Two key structures captured, one showing the SUR2B-nucleotide binding domains (NBDs) separated and one showing the SUR2B-NBDs dimerized, reveal conformation-specific organization of intrinsically disordered regions (IDRs) found in both Kir6.1 and SUR2B. In the NBD-separated conformation, the Kir6.1-N terminal IDR (KNt) sits within the central cleft of the ABC-core of SUR2B. In the NBD-dimerized conformation, KNt is excluded from the central cleft and instead forms contacts with an ED domain comprising 15 consecutive glutamate and aspartate residues within a SUR2B IDR, the N1-T2 linker connecting NBD1 (N1) to transmembrane domain 2 (T2). Moreover, within the N1-T2 linker a regulatory helix seen between the two NBDs in the NBD-separated conformation moves to outside the dimerized NBDs, interacting with the C-terminal residues unique to SUR2B, in the NBD-dimerized conformation. MD simulations further reveal that transient but frequent interactions mediated by the IDRs may facilitate Mg-nucleotide dependent conformational switch in vascular KATP channels.

biochemistry↗

Probing the sequence variability tolerance in a de novo α-helical barrel biocatalyst

De novo-designed enzymes have recently achieved high catalytic activity and stereoselectivity while demonstrating exceptional thermostability in entirely novel protein scaffolds. Among these, -helical barrel protein scaffolds are attractive structures for biocatalysis due to their structural simplicity, high thermostability, and rationalizable sequence patterning. However, enabling major structural reengineering of these scaffolds while maintaining the structure, stability and catalytic activity while also improving soluble protein production remain major challenges and pose the fundamental question how engineerable a de novo backbone-sequence pair is. Here, we combine deep learning based and classic computational protein design to modify and optimize de novo -helical barrel biocatalysts. Using the previously reported six-helical barrel 6H5L as a model scaffold, AlphaFold2-guided RosettaRemodel enabled the design of a truncated variant, whose crystal structure closely matches the design model. Additional sequence-redesign using ProteinMPNN generated a variant with a tenfold increase of soluble protein yield in Escherichia coli. Biochemical, biophysical, and structural analyses showed that both variants retained the overall barrel architecture, high thermal stability, and catalytic activity for both purified protein and whole-cell systems. Detailed kinetic analysis on the variants showed both variation in kcat and Km, reflecting changes in catalytic turnover and substrate binding. Together, these approaches provide new insights and possibilities for the further engineering of functional de novo -helical barrels, their ability to withstand dramatically large sequence changes and their broader application in biocatalysis and biotechnology.

biochemistry↗

Cytokine-induced nuclear translocation of STAT1 via a non-transferable NLS

The targeting function of nuclear localization signals (NLSs) is generally considered independent of a protein's native sequence or fold and is readily transferable to heterologous cargos. Contrary to this paradigm, rapid nuclear translocation of phosphorylated STAT1 (pSTAT1) following cytokine stimulation requires importin {beta}, Ran-GTP, and the importin 5 isoform, which recognizes a non-transferable NLS. Here, we present cryo-EM structures of pSTAT1 bound to importin 5, revealing an asymmetric 2:1 complex that diverges from canonical NLS-mediated cargo recognition. Importin 5 occupies the DNA-binding groove of the pSTAT1 dimer, with a single STAT1 N-terminal domain positioning the C-terminal Armadillo repeats 9-10 (S1B domain) orthogonal to the DNA-binding interface. This interface is also targeted by the Ebola virus protein VP24, an antagonist of interferon signaling. We further show that Ran-GTP alone is insufficient to trigger nuclear release of pSTAT1, which additionally requires the exportin CAS. A cryo-EM reconstruction of the CAS-Ran-GTP-5 complex, supported by in vitro competition assays, demonstrates that CAS and pSTAT1 are mutually exclusive ligands for importin 5. Together, these findings define the molecular choreography of cytokine-induced STAT1 nuclear translocation and release, establishing a general paradigm for STAT family signaling.

biochemistry↗