bioRxiv Science⌕ Search

Biology subjects

Teufel, F.

Publications and source records attributed to Teufel, F..

6 recordsLinked to original sources

LOL-EVE: Predicting Promoter VariantEffects from Evolutionary Sequences

Disease-associated genetic variants occur extensively in noncoding regions like promoters, but current methods focus primarily on single nucleotide variants (SNVs) that typically have small regulatory effect sizes. Expanding beyond single nucleotide events is essential with insertions and deletions (indels) representing the logical next step as they are readily identifiable in population data and more likely to disrupt regulatory elements. However, existing methods struggle with indel prediction, and clinical interpretation often requires assessing complete promoter haplotypes rather than individual variants. We present LOL-EVE (Language Of Life for Evolutionary Variant Effects), a conditional autoregressive transformer trained on 13.6 million mammalian promoter sequences that enables both zero-shot indel prediction and complete promoter sequence scoring. We introduce three benchmarks for promoter indel prediction: ultra rare variant prioritization, causal eQTL identification, and transcription factor binding site disruption analysis. LOL-EVEs superior performance demonstrates that evolutionary patterns learned from indels enable accurate assessment of broader promoter function. Application to Genomics England clinical data shows that LOL-EVE can prioritize promoter haplotypes in known developmental disorder genes, suggesting potential utility for clinical variant assessment. LOL-EVE bridges individual variant prediction with haplotype-level analysis, demonstrating how evolution-based genomic language models may assist in evaluating regulatory variants in complex genetic cases.

genomics↗

Predicting the subcellular location of prokaryotic proteins with DeepLocPro

Protein subcellular location prediction is a widely explored task in bioinformatics because of its importance in proteomics research. We propose DeepLocPro, an extension to the popular method DeepLoc, tailored specifically to archaeal and bacterial organisms. DeepLocPro is a multiclass subcellular location prediction tool for prokaryotic proteins, trained on experimentally verified data curated from UniProt and PSORTdb. DeepLocPro compares favorably to the PSORTb 3.0 ensemble method, surpassing its performance across multiple metrics on our benchmark experiment. The DeepLocPro prediction tool is available online at https://ku.biolib.com/deeplocpro and https://services.healthtech.dtu.dk/services/DeepLocPro-1.0/.

bioinformatics↗

GraphPart: Homology partitioning for biological sequence analysis

When splitting biological sequence data for the development and testing of predictive models, it is necessary to avoid too closely related pairs of sequences ending up in different partitions. If this is ignored, performance estimates of prediction methods will tend to be exaggerated. Several algorithms have been proposed for homology reduction, where sequences are removed until no too closely related pairs remain. We present GraphPart, an algorithm for homology partitioning, where as many sequences as possible are kept in the dataset, but partitions are defined such that closely related sequences always end up in the same partition. Evaluation of GraphPart on Protein, DNA and RNA datasets shows that it is capable of retaining a larger number of sequences per dataset, while providing homology separation quality on par with reduction approaches.

bioinformatics↗

MembraneFold: Visualising transmembrane protein structure and topology

BackgroundAlphaFolds accuracy, which is often comparable to that of experimentally determined structures, has revolutionized protein structure research. Being a statistical method, AlphaFold implicitly infers the cellular environment, e.g. the cell membrane, from the protein sequence. Membrane protein topology prediction methods predict the cellular environment for each protein residue but not the structure. Current structure and topology tools thus provide complementary information. ResultsWe introduce the web server MembraneFold. MembraneFold combines protein structure (from an uploaded PDB file/AlphaFold DB/OmegaFold) and topology (DeepTMHMM) prediction in one server. The output is shown both as a structure with topology superimposed and as a sequence annotation. MembraneFold uses structures predicted by OmegaFold if neither a PDB file is uploaded nor the structure is available in AlphaFold DB. ConclusionMembraneFold is a user-friendly web server that provides practitioners with fast and accurate information about membrane proteins. It is available at https://ku.biolib.com/MembraneFold/.

bioinformatics↗

Identifying endogenous peptide receptors bycombining structure and transmembrane topologyprediction

Many secreted endogenous peptides rely on signalling pathways to exert their function in the body. While peptides can be discovered through high throughput technologies, their cognate receptors typically cannot, hindering the understanding of their mode of action. We investigate the use of AlphaFold-Multimer for identifying the cognate receptors of secreted endogenous peptides in human receptor libraries without any prior knowledge about likely candidates. We find that AlphaFolds predicted confidence metrics have strong performance for prioritizing true peptide-receptor interactions. By applying transmembrane topology prediction using DeepTMHMM, we further improve performance by detecting and filtering biologically implausible predicted interactions. In a library of 1112 human receptors, the method ranks true receptors in the top percentile on average for 11 benchmark peptide-receptor pairs.

bioinformatics↗

SignalP 6.0 achieves signal peptide prediction across all types using protein language models

Signal peptides (SPs) are short amino acid sequences that control protein secretion and translocation in all living organisms. As experimental characterization of SPs is costly, prediction algorithms are applied to predict them from sequence data. However, existing methods are unable to detect all known types of SPs. We introduce SignalP 6.0, the first model capable of detecting all five SP types. Additionally, the model accurately identifies the positions of regions within SPs, revealing the defining biochemical properties that underlie the function of SPs in vivo. Results show that SignalP 6.0 has improved prediction performance, and is the first model to be applicable to metagenomic data. SignalP 6.0 is available at https://services.healthtech.dtu.dk/service.php?SignalP-6.0

bioinformatics↗