bioRxiv Science⌕ Search

Biology subjects

Cevrim, E.

Publications and source records attributed to Cevrim, E..

2 recordsLinked to original sources

OmniPath: integrated knowledgebase for multi-omics analysis

Analysis and interpretation of omics data largely benefit from the use of prior knowledge. However, this knowledge is fragmented across resources and often is not directly accessible for analytical methods. We developed OmniPath (https://omnipathdb.org/), a database combining diverse molecular knowledge from 168 resources. It covers causal protein-protein, gene regulatory, miRNA, and enzyme-PTM (post-translational modification) interactions, cell-cell communication, protein complexes, and information about the function, localization, structure, and many other aspects of biomolecules. It prioritizes literature curated data, and complements it with predictions and large scale databases. To enable interactive browsing of this large corpus of knowledge, we developed OmniPath Explorer, which also includes a large language model (LLM) agent that has direct access to the database. Python and R/Bioconductor client packages and a Cytoscape plugin create easy access to customized prior knowledge for omics analysis environments, such as scverse. OmniPath can be broadly used for the analysis of bulk, single-cell and spatial multi-omics data, especially for mechanistic and causal modeling. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/675512v1_ufig1.gif" ALT="Figure 1"> View larger version (27K): org.highwire.dtl.DTLVardef@17c2b6borg.highwire.dtl.DTLVardef@1069835org.highwire.dtl.DTLVardef@1f2ce76org.highwire.dtl.DTLVardef@1d0b34f_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

A Benchmarking Platform for Assessing Protein Language Models on Function-related Prediction Tasks

Proteins play a crucial role in almost all biological processes, serving as the building blocks of life and mediating various cellular functions, from enzymatic reactions to immune responses. Accurate annotation of protein functions is essential for advancing our understanding of biological systems and developing innovative biotechnological applications and therapeutic strategies. To predict protein function, researchers primarily rely on classical homology-based methods, which use evolutionary relationships, and increasingly on machine learning (ML) approaches. Lately, protein language models (PLMs) have gained prominence; these models leverage specialised deep learning architectures to effectively capture intricate relationships between sequence, structure, and function. We recently conducted a comprehensive benchmarking study to evaluate diverse protein representations (i.e., classical approaches and PLMs) and discuss their trade-offs. The current work introduces the Protein Representation Benchmark - PROBE tool, a benchmarking framework designed to evaluate protein representations on function-related prediction tasks. Here, we provide a detailed protocol for running the framework via the GitHub repository and accessing our newly developed user-friendly web service. PROBE encompasses four core tasks: semantic similarity inference, ontology-based function prediction, drug target family classification, and protein-protein binding affinity estimation. We demonstrate PROBEs usage through a new use case evaluating ESM2 and three recent multimodal PLMs--ESM3, ProstT5, and SaProt--highlighting their ability to integrate diverse data types, including sequence and structural information. This study underscores the potential of protein language models in advancing protein function prediction and serves as a valuable tool for both PLM developers and users.

bioinformatics↗