bioRxiv Science⌕ Search

Biology subjects

Pock, T.

Publications and source records attributed to Pock, T..

2 recordsLinked to original sources

Inferring and Evaluating Network Medicine-Based Disease Modules with Nextflow

Most human diseases result from complex molecular interactions of genes and proteins. Various network-based computational methods characterize these mechanisms by expanding seed genes into disease-associated subnetworks, or disease modules. Evaluating these diverse methods is tedious due to unique installation and data preparation requirements. Moreover, the underlying algorithmic strategies differ, making it difficult to determine which of the created modules are most useful or biologically plausible. To address this challenge, we developed an all-in-one Nextflow pipeline that enables automated and reproducible analyses. It handles installation, input preparation, execution, and systematic evaluation of six widely used module detection tools, considering module topology, functional coherence, robustness, and the capacity to recover seeds. In addition, it annotates the resulting disease modules with biological context information, prioritizes potential drug candidates, and generates visualizations and a comprehensive summary report. To showcase the value of our pipeline and offer guidance to potential users, we performed a comprehensive evaluation across 50 different disease-network combinations, revealing substantial variability among the derived disease modules. We show that this variability is driven by differences in modeling approach, input network, and seed composition. While most methods are robust to minor perturbations, they struggle to recover omitted seeds, and none consistently outperforms others, underscoring the need for careful method selection. Our work enables the research community to systematically compare approaches for disease module discovery, promoting reproducible network medicine research. Integrated into the nf-core project (https://nf-co.re/diseasemodulediscovery), it is intended as an extendable, long-term resource for tracking progress in the field.

bioinformatics↗

Predicting Loop Quality in Protein Structure Models

1.Typically, sequences designed de novo are assessed in silico using deep learning-based protein structure prediction methods prior to wetlab testing. While these deep learning (DL) models excel at predicting well-ordered regions, accurate prediction of loop regions, which often are flexible and crucial for protein function, remains a significant challenge. To address this, we introduce the Equivariant Loop Evaluation Network (ELEN), a local model quality assessment (MQA) method that is tailored towards evaluating the accuracy of protein loops at the per-residue level. ELEN jointly predicts three quality metrics, local Distance Difference Test (lDDT), Contact Area Difference Score (CAD-score), and Root Mean Squared Deviation (RMSD), by comparing predicted to experimental reference structures. Learning these metrics simultaneously enables ELEN to capture complementary structural insights, providing a richer assessment of model accuracy. The network operates at all-atom resolution and employs 3D equivariant group convolutions to learn the local geometric environment of each atom. By incorporating sequence embeddings from large language models (LLMs), such as SaProt, we enhance the sequence and evolutionary awareness of the model. Furthermore, by informing ELEN with per-residue physicochemical features, the model achieves competitive accuracy relative to state-of-the-art MQA methods on the Continuous Automated Model EvaluatiOn (CAMEO) benchmark. Although ELEN was primarily developed for assessing loop quality, its architecture also demonstrates strong potential for general MQA tasks. We used ELEN to perform detailed analysis, including identification of flexible or disordered regions and assessment of structural effects from single-residue mutations on three sets of redesigned enzymes. We show that for all sets ELEN successfully identifies poor design positions and thus serves as a powerful tool for advancing both the study and modeling of loops in protein structures.

bioinformatics↗