bioRxiv Science⌕ Search

Biology subjects

Rigden, D.

Publications and source records attributed to Rigden, D..

8 recordsLinked to original sources

Deep Learning-based structural and functional annotation of Pandoravirus hypothetical proteins

Giant viruses, including Pandoraviruses, contain large amounts of genomic dark matter - genes encoding proteins of unknown function. New generation, deep learning-based protein structure modelling offers new opportunities to apply structure-based function inference to these sequences, often labelled as hypothetical proteins. However, the AlphaFold Protein Structure Database, a convenient resource covering the majority of UniProt, currently lacks models for most viral proteins. Here, we apply a panoply of predictive methods to protein structure predictions representative of large clusters of hypothetical proteins shared among four Pandoraviruses. In several cases, strong functional predictions can be made. Thus, we identify a likely nucleotidyltransferase putatively involved in viral tRNA maturation that has a BTB domain presumably involved in protein-protein interactions. We further identify a cluster of membrane channel sequences presenting three paralogous families which may, as seen in other giant viruses, induce host cell membrane depolarization. And we identify homologues of calcium-activated potassium channel beta subunits and pinpoint their likely Acanthamoeba cellular alpha subunit counterparts. Despite these successes, many other clusters remain cryptic, having folds that are either too functionally promiscuous or too novel to provide strong clues as to their role. These results suggest that significant structural and functional novelty remains to be uncovered in the giant virus proteomes.

bioinformatics↗

CASP15 cryoEM protein and RNA targets: refinement and analysis using experimental maps

CASP assessments primarily rely on comparing predicted coordinates with experimental reference structures. However, errors in the reference structures can potentially reduce the accuracy of the assessment. This issue is particularly prominent in cryoEM-determined structures, and therefore, in the assessment of CASP15 cryoEM targets, we directly utilized density maps to evaluate the predictions. A method for ranking the quality of protein chain predictions based on rigid fitting to experimental density was found to correlate well with the CASP assessment scores. Overall, the evaluation against the density map indicated that the models are of high accuracy although local assessment of predicted side chains in a 1.52 [A] resolution map showed that side-chains are sometimes poorly positioned. The top 136 predictions associated with 9 protein target reference structures were selected for refinement, in addition to the top 40 predictions for 11 RNA targets. To this end, we have developed an automated hierarchical refinement pipeline in cryoEM maps. For both proteins and RNA, the refinement of CASP15 predictions resulted in structures that are close to the reference target structure, including some regions with better fit to the density. This refinement was successful despite large conformational changes and secondary structure element movements often being required, suggesting that predictions from CASP-assessed methods could serve as a good starting point for building atomic models in cryoEM maps for both proteins and RNA. Loop modeling continued to pose a challenge for predictors with even short loops failing to be accurately modeled or refined at times. The lack of consensus amongst models suggests that modeling holds the potential for identifying more flexible regions within the structure.

bioinformatics↗

Microtubule association of TRIM3 revealed by differential extraction proteomics

The microtubule network is formed from polymerised Tubulin subunits and associating proteins, which govern microtubule dynamics and a diverse array of functions. To identify novel microtubule binding proteins, we have developed an unbiased biochemical assay, which relies on the selective extraction of cytosolic proteins from cells, whilst leaving behind the microtubule network. Candidate proteins are linked to microtubules by their sensitivities to the depolymerising drug Nocodazole or the microtubule stabilising drug, Taxol, which is quantitated by mass spectrometry. Our approach is benchmarked by co-segregation of Tubulin and previously established microtubule-binding proteins. We then identify several novel candidate microtubule binding proteins, from which we have selected the ubiquitin E3 ligase TRIM3 (Tripartite motif-containing protein 3) for further characterisation. We map TRIM3 microtubule binding to its C-terminal NHL-repeat region. We show that TRIM3 is required for the accumulation of acetylated Tubulin, following treatment with Taxol. Furthermore, loss of TRIM3, partially recapitulates the reduction in Nocodazole-resistant microtubules characteristic of Alpha-Tubulin Acetyltransferase 1 (ATAT1) depletion. These results can be explained by a decrease in ATAT1 following depletion of TRIM3 that is independent of transcription.

cell biology↗

Assessment of three-dimensional RNA structure prediction in CASP15

The prediction of RNA three-dimensional structures remains an unsolved problem. Here, we report assessments of RNA structure predictions in CASP15, the first CASP exercise that involved RNA structure modeling. Forty two predictor groups submitted models for at least one of twelve RNA-containing targets. These models were evaluated by the RNA-Puzzles organizers and, separately, by a CASP-recruited team using metrics (GDT, lDDT) and approaches (Z-score rankings) initially developed for assessment of proteins and generalized here for RNA assessment. The two assessments independently ranked the same predictor groups as first (AIchemy_RNA2), second (Chen), and third (RNAPolis and GeneSilico, tied); predictions from deep learning approaches were significantly worse than these top ranked groups, which did not use deep learning. Further analyses based on direct comparison of predicted models to cryogenic electron microscopy (cryo-EM) maps and X-ray diffraction data support these rankings. With the exception of two RNA-protein complexes, models submitted by CASP15 groups correctly predicted the global fold of the RNA targets. Comparisons of CASP15 submissions to designed RNA nanostructures as well as molecular replacement trials highlight the potential utility of current RNA modeling approaches for RNA nanotechnology and structural biology, respectively. Nevertheless, challenges remain in modeling fine details such as non- canonical pairs, in ranking among submitted models, and in prediction of multiple structures resolved by cryo-EM or crystallography.

biophysics↗

Structural Insights into Pink-eyed Dilution Protein (Oca2)

Recent innovations in computational structural biology have opened an opportunity to revise our current understanding of the structure and function of clinically important proteins. This study centres on human Oca2 which is located on mature melanosomal membranes. Mutations of Oca2 can result in a form of oculocutanous albinism which is the most prevalent and visually identifiable form of albinism. Sequence analysis predicts Oca2 to be a member of the SLC13 transporter family but it has not been classified into any existing SLC families. The modelling of Oca2 with AlphaFold2 and other advanced methods shows that, like SLC13 members, it consists of a scaffold and transport domain and displays a pseudo inverted repeat topology that includes re-entrant loops. This finding contradicts the prevailing consensus view of its topology. In addition to the scaffold and transport domains the presence of a cryptic GOLD domain is revealed that is likely responsible for its trafficking from the endoplasmic reticulum to the Golgi prior to localisation at the melanosomes and possesses known glycosylation sites. Analysis of the putative ligand binding site of the model shows the presence of highly conserved key asparagine residues that suggest Oca2 may be a Na+/dicarboxylate symporter. Known critical pathogenic mutations map to structural features present in the repeat regions that form the transport domain. Exploiting the AlphaFold2 multimeric modelling protocol in combination with conventional homology modelling allowed the building of a plausible homodimer in both an inward- and outward-facing conformation supporting an elevator-type transport mechanism.

bioinformatics↗

Using deep learning predictions of inter-residue distances for model validation

Determination of protein structures typically entails building a model that satisfies the collected experimental observations and its deposition in the Protein Data Bank (PDB). Experimental limitations can lead to unavoidable uncertainties during the process of model building, which result in the introduction of errors into the deposited model. Many metrics are available for model validation, but most are limited to the consideration of the physico-chemical aspects of the model or its match to the map. The latest advances in the field of deep learning have enabled the increasingly accurate prediction of inter-residue distances, an advance which has played a pivotal role in the recent improvements observed in the field of protein ab initio modelling. Here we present new validation methods based on the use of these precise inter-residue distance predictions, which are compared with the distances observed in the protein model. Sequence register errors are particularly clearly detected, and the register shifts required for their correction can be reliably determined. The method is available in the package ConKit (www.conkit.org).

bioinformatics↗

Slice'N'Dice: Maximising the value of predicted models for structural biologists

With the advent of next generation modelling methods, such as AlphaFold2, structural biologists are increasingly using predicted structures as search models for Molecular Replacement (MR) when experimental structures of homologues are unavailable. Inaccuracy in domain-domain orientations is often a key limitation when using predicted models for MR. SliceNDice is a software package designed to address this issue by first slicing models into distinct structural units and then automatically placing the slices using Phaser. The slicing step can use AlphaFold2s predicted aligned error (PAE), or can operate via a variety of C atom clustering algorithms, extending applicability to structures of any origin. The number of splits can be selected by the user. SliceNDice is available in CCP4 8.0 and is currently being adapted for cryo-EM use cases.

bioinformatics↗

m6A-TSHub: unveiling the context-specific m6A methylation and m6A-affecting mutations in 23 human tissues

As the most pervasive epigenetic marker present on mRNA and lncRNA, N6-methyladenosine (m6A) RNA methylation has been shown to participate in essential biological processes. Recent studies revealed the distinct patterns of m6A methylome across human tissues, and a major challenge remains in elucidating the tissue-specific presence and circuitry of m6A methylation. We present here a comprehensive online platform m6A-TSHub for unveiling the context-specific m6A methylation and genetic mutations that potentially regulate m6A epigenetic mark. m6A-TSHub consists of four core components, including (1) m6A-TSDB: a comprehensive database of 184,554 functionally annotated m6A sites derived from 23 human tissues and 499,369 m6A sites from 25 tumor conditions, respectively; (2) m6A-TSFinder: a web server for high-accuracy prediction of m6A methylation sites within a specific tissue from RNA sequences, which was constructed using multi-instance deep neural networks with gated attention; (3) m6A-TSVar: a web server for assessing the impact of genetic variants on tissue-specific m6A RNA modification; and (4) m6A-CAVar: a database of 587,983 TCGA cancer mutations (derived from 27 cancer types) that were predicted to affect m6A modifications in the primary tissue of cancers. The database should make a useful resource for studying the m6A methylome and genetic factor of epitranscriptome disturbance in a specific tissue (or cancer type). m6A-TSHub is accessible at: www.xjtlu.edu.cn/biologicalsciences/m6ats.

bioinformatics↗