bioRxiv Science⌕ Search

Biology subjects

Villegas Garcia, E. N.

Publications and source records attributed to Villegas Garcia, E. N..

2 recordsLinked to original sources

Decoding the Grammar of Protein-Protein Interaction Interfaces with Multimodal Representations

Protein-protein interactions govern essential cellular processes, making the identification of interacting sites a central challenge in structural biology, with important implications for protein engineering and the development of targeted therapeutics. Existing prediction algorithms include sequence-based methods, which lack structural information, or structure-based approaches, which often struggle to effectively integrate evolutionary context. Here, we present ESM3-PPISites, a supervised model for residue-level classification of interfaces, leveraging the multimodal representations of the ESM3 Protein Language Model. To ensure a bias-free evaluation, a stringent redundancy filtering protocol is adopted, systematically eliminating latent homology between the training data and a curated benchmark set in both sequence and structural space. ESM3-PPISites achieves unprecedented accuracy, vastly outperforming current approaches. Our findings demonstrate that while ESM3 largest proprietary version yields the highest predictive power, targeted fine-tuning of its small open-weight counterpart significantly narrows the performance gap. We also show the practical impact of these predictions by integrating them as spatial restraints within the HADDOCK docking platform. When evaluated on an independent subset of 12 complexes from the Docking Benchmark v5, the prediction-guided pipeline strongly enhances the identification of near-native binding poses over blind docking, while reducing computational runtime by an order of magnitude. This framework establishes a scalable paradigm for high-throughput structural characterization of protein-protein interactions.

bioinformatics↗

Protein family annotation for the Unified Human Gastrointestinal Proteome by DPCfam clustering

Technological advances in massively parallel sequencing have led to an exponential growth in the number of known protein sequences. Much of this growth originates from metagenomic projects producing new sequences from environmental and clinical samples. The Unified Human Gastrointestinal Proteome (UHGP) catalogue is one of the most relevant metagenomic datasets with applications ranging from medicine to biology. However, the lack of sequence annotation impairs its usability. This work aims to produce a family classification of UHGP sequences to facilitate downstream structural and functional annotation. This is achieved through the release of the DPCfam-UHGP50 dataset containing 10,778 putative protein families generated using DPCfam clustering, an unsupervised pipeline grouping sequences into multi-domain architectures. DPCfam-UHGP50 considerably improves family coverage at protein and residue levels compared to the manually curated repository Pfam. It is our hope that DPCfam-UHGP50 will foster future discoveries in the field of metagenomics of the human gut by the release of a FAIR-compliant database easily accessible via a searchable web server and Zenodo repository.

bioinformatics↗