bioRxiv Science⌕ Search

Biology subjects

Rota Negroni, M.

Publications and source records attributed to Rota Negroni, M..

2 recordsLinked to original sources

A Pan-Cancer Multi-Omic Analysis of Copy Number Signature Clusters and Genomic Instability

Copy number signatures provide compact representations of the processes that shape cancer genomes, but signatures derived with different feature encodings are often interpreted as if they were interchangeable. We established a matched-sample pan-cancer benchmark of three major copy number signature compendia, comparing their activity structure, cross-framework concordance, patient stratification, outcome associations, and predictability from non-copy-number molecular data. Signature- level concordance was sparse and concentrated in a limited set of biologically related patterns. Clustering of high-activity signatures produced distinct patient partitions with limited overlap between compendia, although one cluster in each framework showed a directionally favorable outcome association after accounting for cancer-type-specific baseline hazards. Prediction from gene expression, DNA methylation, somatic mutations, age, and tumor purity was strongly framework dependent: test-set F1 scores were 0.93 for Drews, 0.80 for Steele, and 0.24 for Tao. Gene expression provided the largest contribution and largely retained the performance of the full models. These results show that compendium choice is an analytical decision rather than an interchangeable preprocessing step. The benchmark provides a reproducible framework for selecting and interpreting copy number signature representations in pan-cancer studies.

cancer biology↗

An AI-driven pipeline for the discovery of hidden peptides in plant proteomes: the CLE family as a case study

Plant proteomes contain evolutionarily conserved peptides with poorly conserved primary sequences, often hindering their identification and classification into families. Homology-based approaches and conventional annotation pipelines frequently fail to detect these family members, particularly in poorly characterized, but agronomically relevant plant species. CLE peptides (CLAVATA3/EMBRYO SURROUNDING REGION-related peptides) constitute a large and evolutionarily conserved family of plant signaling molecules, yet their characterization remains incomplete. Beyond a limited number of well-studied members, a substantial number of CLE peptides remain uncharacterized due to functional redundancy and the intrinsic features of CLE genes, which encode short pre-propeptides with only a small 12-residue conserved motif. Here, we present a novel framework leveraging state-of-the-art Protein Language Models (pLMs) to discover CLE peptides directly from 13 plant proteomes. By coupling sequence embeddings trained on large evolutionary datasets (ESM2 and ProtT5) with supervised machine learning, our dual-model approach captures deep semantic features of the CLE family that are missed by traditional alignment methods. The pipeline demonstrated robust generalization, achieving high classification accuracy (98.9-99.4%) on a held-out set of CLE peptides not used during training. Consequently, we identified a set of high-confidence, previously unannotated CLE candidates prioritized through a stringent consensus-based filtering strategy. This work demonstrates how AI-driven proteome analysis can overcome the limitations of homology-based methods and provides a scalable strategy for uncovering previously unidentified peptide-mediated signaling molecules across plant lineages. HighlightLeveraging Protein Language Models, our AI framework uncovers "hidden" signaling peptides missed by standard tools, revealing the elusive diversity of CLE regulators across plant proteomes.

plant biology↗