bioRxiv Science⌕ Search

Biology subjects

Senoner, T.

Publications and source records attributed to Senoner, T..

6 recordsLinked to original sources

Pseudoperplexity Probes Memorization in Protein Language Models

Protein Language Models (pLMs) have significantly advanced computational biology. Yet their scale and reliance on redundant training data raise a fundamental question: do pLMs generalize the statistical grammar of proteins, or do they simply memorize their training data? To investigate this, we used pseudoperplexity as a probe for sequence-level memorization, comparing ProtT5's pseudoperplexity on a pre-training proxy dataset against a post-training holdout of genuinely novel sequences. To ensure a valid comparison, we matched the datasets by sequence length, cluster size, and taxonomic family. As a statistical baseline, we trained n-gram language models; analysis of higher-order n-gram composition and a statistically significant divergence in perplexity confirmed that the post-training sequences were genuinely novel at the local sequence level. ProtT5 showed a statistically significant difference in pseudoperplexity between seen and unseen sequences, though further analysis revealed this memorization signal to be modest. These findings suggest that ProtT5 exhibits detectable but limited memorization of its training data as measured by a pseudoperplexity-based probe.

bioinformatics↗

ProtSpace: Protein Universe in Your Browser

AO_SCPLOWBSTRACTC_SCPLOWProtein Language Models (pLMs) generate per-protein embeddings that encode functional, structural, and evolutionary information, yet the relationships captured in these representations remain difficult to explore systematically. ProtSpace (https://protspace.app) is a web application for interactive visualization of pLM embedding spaces, enabling hypothesis generation directly in the browser without installation. Unlike traditional network-based tools that exclusively visualize amino acid sequence similarity, ProtSpace explores embedding spaces, revealing relationships often not captured by traditional comparisons. Users provide protein sequences or pre-computed embeddings through a Google Colab notebook or the Python CLI; the pipeline applies dimensionality reduction, retrieves 38 annotation types spanning UniProt, InterPro, NCBI Taxonomy, TED structural domains, and sequence-based predictors served via Biocentral, and produces a portable binary file for the browser-based viewer. WebGL-accelerated rendering supports interactive exploration of over 570,000 proteins. Distinctive features include per-point pie charts for multi-label annotations and integrated 3D structure viewing through AlphaFold2 predictions. All computation happens on the users machine, ensuring data privacy. We demonstrate the utility of ProtSpace through a progressive zoom-in across biological scales: from global proteome organization of Swiss-Prot, through cross-species comparison revealing conserved and lineage-specific families, to functional hypothesis generation within the beta-lactamase superfamily. ProtSpace is freely available at https://protspace.app under the Apache 2.0 license. KO_SCPLOWEYC_SCPLOWO_SCPCAP C_SCPCAPO_SCPLOWPOINTSC_SCPLOWO_LIProtSpace is a free, open-source web application that visualizes protein Language Model (pLM) embeddings as interactive maps, scaling to 570,000 proteins entirely client-side. C_LIO_LIA zero-installation Google Colab notebook and a Python CLI prepare visualization-ready bundles from FASTA files, UniProt queries, or pre-computed HDF5 embeddings, automatically retrieving 38 annotation types from five sources (UniProt, InterPro, NCBI Taxonomy, TED structural domains, and Biocentral sequence predictors) alongside custom CSV metadata. C_LIO_LIApplication examples demonstrate that embedding visualizations generate testable biological hypotheses at multiple scales, from proteome-wide organization through species-level comparison to family-level functional discovery, and that these are complementary to traditional sequence-based analyses. C_LI

bioinformatics↗

UdonPred: Untangling Protein Intrinsic Disorder Prediction

MotivationRegions in intrinsic disordered proteins (IDPs) constitute important continuous aspects of protein function. While their existence on a structural continuum is widely accepted, most computational predictions have, nevertheless, focused on binary classifications. Existing datasets are severely limited in size and experimental evidence for continuous disorder. ResultsBuilding on recently released datasets of continuous protein disorder and flexibility, we introduce UdonPred, a lightweight neural network exclusively inputting embeddings from the protein Language Model (pLM) ProstT5 to predict per-residue protein disorder from sequence alone. Training and evaluating UdonPred on seven datasets with divergent definitions of disorder and flexibility suggests that not model capacity, but agreement and nuance of disorder annotations, remains the main driver of performance. Binary disorder annotations can be reliably predicted from a multitude of different disorder and flexibility datasets, but there is still room for improvement in predicting continuous disorder. AvailabilityAll code and data used for training and evaluation is available under an open-source license at https://github.com/davidwagemann/udonpred. Contactassistant@rostlab.org Supplementary informationSupplementary data are available at Journal Name online.

bioinformatics↗

Which pLM to choose?

AO_SCPLOWBSTRACTC_SCPLOWProtein-language models (pLMs) provide a novel means for mapping the protein space. Which of these new maps best advances specific biological analyses, however, is not obvious. To elucidate the principles of model selection, we benchmarked fourteen pLMs, spanning several orders of magnitude in parameter count, across a hundred million protein pairs, to assess how well they capture sequence, structure, and function similarity. For each model, we distinguish inherent information, i.e. signal recoverable from raw-embedding distances, and extractable information, i.e. signal revealed through additional supervised training. Three key results emerge. First, pLM protein representation space is inherently different from the space of biological protein representations, i.e. sequences or structures. Here, a size-performance paradox is salient - mid-scale foundation models are as good as much larger ones in reflecting all tested biological properties. Second, pLM representations compress and store biological information in proportion to model size. That is, a lightweight feed-forward network can be trained on embedding pairs to predict said biological properties well - a capacity dividend. Finally, we observe that a task-specific learning radically reshapes the embedding space, gaining inherent understanding of the task, but garbling any further extractions. In other words, smaller pLMs can provide efficient and compute-light general insight. Larger models are advantageous only when fine-tuning is planned to accomplish a specific task. Furthermore, representations generated by "specialist" models are not immediately generalizable throughout protein biology. Thus, for pLMs, bigger isnt always better.

bioinformatics↗

FlatProt: 2D visualization eases protein structure comparison

AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBackgroundC_ST_ABSUnderstanding and comparing three-dimensional (3D) structures of proteins can advance bioinformatics, molecular biology, and drug discovery. While 3D models offer detailed insights, comparing multiple structures simultaneously remains challenging, especially on two-dimensional (2D) displays. Existing 2D visualization tools lack standardized approaches for pipelined inspection of large protein sets, limiting their utility in large-scale pre-filtering. ResultsWe introduce FlatProt, a tool designed to complement 3D viewers by enabling standardized 2D visualization of individual protein structures or large sets thereof. By including Foldseek-based family rotation alignment or an inertia-based fallback, FlatProt creates consistent and scalable visual representations for user-defined protein structures. It supports domain-aware decomposition, family-level overlays, and lightweight visual abstraction of secondary structures. FlatProt processes proteins efficiently, as showcased on a subset of the human-proteome. ConclusionFlatProt provides clear, consistent, user-friendly visualizations that support rapid, comparative inspection of protein structures at scale. By bridging the gap between interactive 3D tools and static visual summaries, it enables users to explore conserved features, detect outliers, and prioritize structures for further analysis. AvailabilityGitHub (https://github.com/t03i/FlatProt); Zenodo (https://doi.org/10.5281/zenodo.15697296).

bioinformatics↗

ProtSpace: a tool for visualizing protein space

Protein language models (pLMs) generate high-dimensional representations of proteins, so called embeddings, that capture complex information stored in the set of evolved sequences. Interpreting these embeddings remains an important challenge. ProtSpace provides one solution through an open-source Python package that visualizes protein embeddings interactively in 2D and 3D. The combination of embedding space with protein 3D structure view aids in discovering functional patterns readily missed by traditional sequence analysis. We present two examples to showcase ProtSpace. First, investigations of phage data sets showed distinct clusters of major functional groups and a mixed region, possibly suggesting bias in todays protein sequences used to train pLMs. Second, the analysis of venom proteins revealed unexpected convergent evolution between scorpion and snake toxins; this challenges existing toxin family classifications and added evidence refuting the aculeatoxin family hypothesis. ProtSpace is freely available as a pip-installable Python package (source code & documentation) with examples on GitHub (https://github.com/tsenoner/protspace) and as a web interface (https://protspace.rostlab.org). The platform enables seamless collaboration through portable JSON session files.

bioinformatics↗