bioRxiv Science⌕ Search

Biology subjects

Litvin, U.

Publications and source records attributed to Litvin, U..

2 recordsLinked to original sources

Viro3D: a comprehensive database of virus protein structure predictions

Viruses are intracellular parasites of organisms from all domains of life. They infect and cause disease in humans, animals and plants but also play crucial roles in the ecology of microbial communities. Tolerance to genetic change, high-mutation rates, adaptations to hosts and immune escape has driven high divergence of viral genes, hampering their functional annotation and phylogenetic inference. The protein structure is more conserved than sequence and can be used for searches of distant homologs and evolutionary analysis of divergent proteins. Structures of viral proteins are traditionally underrepresented in public databases, but recent advances in protein structure prediction allows us to address this issue. Combining two state-of-the-art approaches, AlphaFold2-ColabFold and ESMFold, we predicted models for 85,000 proteins from 4,400 human and animal viruses, expanding the structural coverage for viral proteins by 30 times compared to experimental structures. We also performed structural and network analyses of the models to demonstrate their utility for functional annotation and inference of distant phylogenetic relationships. Taking this approach, we examined the deep evolutionary history of viral class-I fusion glycoproteins, gaining insights on the origins of coronavirus spike protein. To enable further discoveries, we have created Viro3D (https://viro3d.cvr.gla.ac.uk/), a virus species-centred protein structure database. It allows users to search, browse and download protein models from a virus of interest and explore similar structures present in other virus species. This resource will facilitate fundamental molecular virology, investigation of virus evolution, and may enable structure-informed design of therapies and vaccines.

microbiology↗

pyRBDome: A comprehensive computational platform for enhancing and interpreting RNA-binding proteome data

High-throughput proteomics approaches have revolutionised the identification of RNA-binding proteins (RBPome) and RNA-binding sequences (RBDome) across organisms. Yet the extent of noise, including false-positives, associated with these methodologies, is difficult to quantify as experimental approaches for validating the results are generally low throughput. To address this, we introduce pyRBDome, a pipeline for enhancing RNA-binding proteome data in silico. It aligns the experimental results with RNA-binding site (RBS) predictions from distinct machine learning tools and integrates high-resolution structural data when available. Its statistical evaluation of RBDome data enables quick identification of likely genuine RNA-binders in experimental datasets. Furthermore, by leveraging the pyRBDome results, we have enhanced the sensitivity and specificity of RBS detection through training new ensemble machine learning models. pyRBDome analysis of a human RBDome dataset, compared with known structural data, revealed that while UV cross-linked amino acids were more likely to contain predicted RBSs, they infrequently bind RNA in high-resolution structures. This discrepancy underscores the limitations of structural data as benchmarks, positioning pyRBDome as a valuable alternative for increasing confidence in RBDome datasets.

systems biology↗