bioRxiv Science⌕ Search

Biology subjects

Porubska, J.

Publications and source records attributed to Porubska, J..

3 recordsLinked to original sources

AlphaFind v2: Similarity Search in AlphaFold DB and TED Domains across Structural Contexts

The availability of large-scale protein structure collections enables structure-based analysis of their function and evolution beyond what is possible from sequence alone. However, applying three-dimensional structure comparison at scale remains computationally demanding and limits practical exploration of large experimental and predicted collections. This creates a need for fast, structure-based search methods that retain biological relevance while enabling large-scale exploration. In this paper, we present AlphaFind v2, an application for finding structurally similar proteins in the AlphaFold Database (https://alphafold.ebi.ac.uk/) of predicted structures. AlphaFind v2 uses fast pre-filtering via state-of-the-art protein embeddings that preserve structural information, followed by refinement with US-align. The application presents multiple complementary search modes, including (i) search over full protein chains, (ii) search aware of the AlphaFold pLDDT metric, restricting similarity computation to the most stable and structurally relevant regions, (iii) search over protein domains from the TED database (https://ted.cathdb.info/), and (iv) a multidomain search mode, combining multiple chain-level domain matches within a single score and alignment. The application accepts protein identifiers and returns similar proteins with metrics, rich metadata, and interactive superpositions. AlphaFind v2 additionally allows searching within an organism or CATH label and matches the proteins with experimental structures. AlphaFind v2 is accessible at https://alphafind.ics.muni.cz/. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/710735v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@5e7ce7org.highwire.dtl.DTLVardef@15a3458org.highwire.dtl.DTLVardef@12299ccorg.highwire.dtl.DTLVardef@9f21af_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Rare ring conformations in PDB: Facts or wishful thinking?

Protein structural data are highly valuable for research, and many significant results have been published on their basis. A key point for their credibility and applicability is their quality. An important facet of protein structure quality is the validation of ligands. Some aspects of ligand quality have already been validated by established quality metrics. However, validation of ring conformation has yet to be comprehensively performed despite rings strongly influencing the formation of the ligands scaffold and shape. Most rings form several conformations that differ in their stability. The most stable ones occur frequently in nature and should, therefore, be found in Protein Data Bank (PDB) structures. In this article, we examined which conformations of rings occur in PDB structures. Our analysis focused on conformations of all cyclopentane, cyclohexane, and benzene rings in the PDB. Specifically, we examined 123 264 rings of 24 763 distinct ligands, which instances occur in 44 022 protein structures. In general, we found that most of the rings (98.32 %) are in energetically favourable conformations. Surprisingly, the existence of most of the energetically unfavourable ring conformations (2 067 samples, 1.68 %) is not supported by experimental data. Only 291 unfavourable ring conformations (0.24 %) are backed by experimental data that are accurate enough to distinguish the conformation, which shows that the existence of energetically unfavourable ring conformations is rarely supported by structural or experimental evidence. Our results suggest that each occurrence of untypical ring conformation in the PDB may indicate a potential error and should be carefully analysed.

bioinformatics↗

AlphaFind: Discover structure similarity across the entire known proteome

AlphaFind is a web-based search engine that provides fast structure-based retrieval in the entire set of AlphaFold DB structures. Unlike other protein processing tools, AlphaFind is focused entirely on tertiary structure, automatically extracting the main 3D features of each protein chain and using a machine learning model to find the most similar structures. This indexing approach and the 3D feature extraction method used by AlphaFind have both demonstrated remarkable scalability to large datasets as well as to large protein structures. The web application itself has been designed with a focus on clarity and ease of use. The searcher accepts any valid Uniprot ID, PDB ID or gene symbol as input, and returns a set of similar protein chains from AlphaFold DB, including various similarity metrics between the query and each of the retrieved results. In addition to the main search functionality, the application provides 3D visualizations of protein structure superpositions in order to allow researchers to instantly analyze the structural similarity of the retrieved results. The AlphaFind web application is available online for free and without any registration at https://alphafind.fi.muni.cz. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=82 SRC="FIGDIR/small/580465v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@1bf5ae8org.highwire.dtl.DTLVardef@1e975d7org.highwire.dtl.DTLVardef@378598org.highwire.dtl.DTLVardef@123e439_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗