bioRxiv Science⌕ Search

Biology subjects

Saccon, F.

Publications and source records attributed to Saccon, F..

4 recordsLinked to original sources

Chemical Descriptors and Deep Learning Embeddings for Scoring de novo Peptide Designs

Peptides occupy a valuable niche between small molecules and biologics, but the clinical translation of de novo peptide designs requires rigorous scoring to simultaneously optimise target binding affinity alongside multiple developability traits, including stability, membrane permeability, aggregation propensity, and non-fouling behaviour. Here, we evaluate two distinct approaches for scoring these candidates: classical chemical descriptors and modern deep learning representations derived from protein language and folding models. Assembling nine public datasets spanning five developability traits and four binding-affinity endpoints, we find sequence-derived chemical descriptors alone contain sufficient information to predict developability task labels effectively. Given their drastically lower computational cost and higher interpretability, classical machine learning models trained on these simple descriptors frequently match or approach the performance of complex deep learning architectures, emerging as a highly efficient and interpretable alternative for high-throughput scoring. Finally, for scoring binding affinity, we demonstrate that Boltz-2 pair representations capture the most information among the tested representations; however, the model's predictive power is confounded by a significant bias from the molecular weight of the peptides. Together, these results establish a comprehensive assessment of state-of-the-art methods for predicting both peptide developability and binding affinity, highlighting the enduring value of interpretable chemical descriptors alongside deep learning in the scoring and selection of de novo peptide designs.

bioinformatics↗

Charge reversal at the Lhcb2 N-terminus impairs phosphorylation and PSI-LHCII complex formation

State transitions balance excitation-energy distribution between Photosystem I and Photosystem II in higher plants. Stn7-mediated phosphorylation of the N-terminus of the light-harvesting complex II protein Lhcb2 plays a central role in photosynthetic state transitions. However, it remains unclear how the intrinsic charge of this region, independent of its phosphorylation status, influences state transitions and thylakoid membrane organization. Here, we introduced specific charge-altering mutations in the Lhcb2 N-terminus of Arabidopsis thaliana in the lhcb2 knock-out background and analyzed their effects on LHCII phosphorylation, state transition dynamics, PSI-LHCII complex formation, and thylakoid ultrastructure. Substitution of a conserved positively charged arginine with a negatively charged glutamate (R2E) markedly reduced Lhcb1 and Lhcb2 phosphorylation and state transition efficiency, and abolished PSI-LHCII complex formation. In contrast, introducing a negative charge at a downstream position (Q9E) had no detectable effects. Electron microscopy revealed no significant changes in thylakoid organization in either mutant compared to WT Lhcb2 plants. Despite strongly reduced Lhcb1 and Lhcb2 phosphorylation in the R2E mutant, residual state transitions persisted, potentially mediated by Stn7-dependent phosphorylation of other target proteins. Together, these results provide insight into the role of N-terminal LHCII electrostatics in state transitions and thylakoid membrane organization in plants.

plant biology↗

Exploring the Conformational Landscape of Adenylate Kinase and Beyond: A Benchmark of Protein Folding Models

Protein folding models have revolutionized structure prediction but struggle to capture conformational flexibility. Recent studies perturb inputs or parameters to sample alternative conformations, while diffusion-based approaches generate conformational ensembles directly. Although the former have been benchmarked to some extent, the latter have yet to be evaluated, and sub-domain dynamics validation remains limited. Here, we present a systematic benchmark of nine methods across 20 monomeric proteins with active and inactive states. We extend the pairwise aligned error metric to ensembles and reveal that protein identity exerts a non-negligible influence on model performance. Focusing on Adenylate Kinase, a well-studied enzyme with extensive molecular dynamics (MD) data, we find that Chai-1 performs the best in recovering known conformations, identifying mobile regions, and capturing transition trajectories. These results highlight the potential of generative models as efficient alternatives to MD for exploring protein conformational dynamics and provide a rigorous benchmark for dynamic structure prediction.

bioinformatics↗

CoV-UniBind: A Unified Antibody Binding Database for SARS-CoV-2

Since the emergence of SARS-CoV-2, numerous studies have investigated antibody interactions with viral variants in vitro, and several datasets have been curated to compile available protein structures and experimental measurements. However, existing data remain fragmented, limiting their utility for the development and validation of machine learning models for antibody-antigen interaction prediction. Here, we present CoV-UniBind, a unified database comprising over 75,000 entries of SARS-CoV-2 antibody-antigen sequence, binding, and structural data, integrated and standardised from three public sources and multiple peer-reviewed publications. To demonstrate its utility, we benchmarked multiple protein folding and inverse folding models across tasks relevant to antibody design and vaccine development. We expect CoV-UniBind to facilitate future computational efforts in antibody and vaccine development against SARS-CoV-2. Availability and implementationThe curated datasets, structures, model scores and antibody synonyms are free to download at https://huggingface.co/datasets/InstaDeepAI/cov-unibind. Folded structures are available upon request.

immunology↗