bioRxiv Science⌕ Search

Biology subjects

Hutter, F.

Publications and source records attributed to Hutter, F..

10 recordsLinked to original sources

Prediction of plant organismal complexity based on transcription factor annotation: an AI approach

How morphological complexity evolves is still enigmatic. While there is evidence in algae and plants as well as animals that diversification of the repertoire of transcription factors (TF) is causative for evolution of organismal complexity, there are many examples from lineages that follow their own way of complexity evolution, for example by expansion of particular families. For land plants, correlation of the size of the TF complement with number of cell types (as a proxy for morphological complexity) has been shown, and several families were identified as candidates to drive complexity evolution. Here, we expand a previously available dataset of cell type numbers from 12 to 82 proteomes and introduce a four class body plan scheme. We find that the total TF complement correlates with the number of cell types of Archaeplastida (primary plastid bearing plants and algae). We used TabPFN (Tabular Prior-data Fitted Network) for binary (uni- vs. multicellularity) as well as for four class Bauplan classification. TabPFN is able to predict the morphological complexity with high accuracy. This approach allows to determine organismal complexity based on the gene space of an organism. Based on our results, we can confirm that plant morphological evolution is driven by gain and expansion of TF families.

evolutionary biology↗

Transcriptome-based cell type assignment for kidney cell culture models

BackgroundKidney cell lines are widely used to model kidney physiology and disease; however, their gene expression profiles may differ from primary cells due to immortalization, culture conditions, or experimental treatments. Determining whether a cell line resembles its native cell type is critical for interpreting in vitro findings. We developed a transcriptome-based approach that matches bulk RNA-seq data from kidney cell lines, primary cells, or tissues to reference cell types derived from single-cell RNA-seq (scRNA-seq) datasets. MethodsReference transcriptomic profiles were generated from two human and two murine kidney scRNA-seq datasets by pseudobulk aggregation. Bulk RNA-seq data from microdissected kidney tissue, non-kidney negative controls, and kidney cell lines were matched to these references using three statistical similarity measures (Spearman correlation, Euclidean distance, Poisson distance) and three machine learning classifiers (Random Forest, XGBoost, TabPFN). Each was assessed with global gene expression, curated kidney marker gene lists, and the most variable genes. Matching accuracy was evaluated through a three-step validation strategy: within-dataset matching, cross-reference comparison, and validation against primary kidney tissue and negative controls. ResultsGene expression rank-based Spearman correlation and TabPFN, a foundation model for tabular data, emerged as the most accurate and specific approaches, particularly with curated kidney marker gene lists. Both methods correctly identified microdissected kidney tubule segments and were robust against non-kidney negative controls. Applied to commonly used kidney cell lines, OK cells retained proximal tubule identity, particularly under shear stress, while other proximal tubule lines (HK-2, HKC-8, HKC-11) showed inconsistent matching. Collecting duct-derived mIMCD-3 maintained stable similarity across passages, culture conditions, and genetic modifications. ConclusionWe provide two complementary implementations: CellMatchR, an accessible web-based tool using Spearman correlation for routine use, and comprehensive scripts for TabPFN-based matching (link will be added after peer reviewed publication). Together, these resources enable researchers to make informed decisions about kidney cell culture model selection, interpretation, and stability. Translational StatementKidney cell lines are fundamental tools in nephrology research, yet their transcriptomic similarity to native cell types is rarely validated systematically. We demonstrate that combining bulk RNA-seq data with single-cell reference datasets enables robust assessment of cell line identity using gene expression-rank-based correlation and machine learning approaches. By providing a comprehensive evaluation of matching methods, curated kidney marker gene lists, and reference datasets, our study serves as both a practical resource and a methodological framework for the kidney research community, facilitating informed selection of cell culture models, quality control of experimental conditions, developing new experimental cell culture models, and more reliable translation of in vitro findings to kidney physiology and disease.

bioinformatics↗

Enhancing Intra-Continental Biogeographical Ancestry Prediction Through a Machine Learning Marker Selection Method

While classifiers such as TabPFN (Hollmann et al., 2025) and SNIPPER (Phillips et al., 2007a) achieve strong intercontinental performance (Heinzel et al., 2025), their accuracy in classifying individuals within Europe remains low. One major factor contributing to this limitation is the set of genetic markers used for classification. Marker panels such as the VISAGE Enhanced Tool (Xavier et al., 2022) are commonly employed in forensic genetics because they contain ancestry-informative markers (AIMs) that distinguish very well between major continental populations. However, these panels are often not optimized for fine-scale differentiation within continents, where genetic variation is more subtle and population structure is rather continuous. We apply machine learning to select informative markers for intra-European classification, using data from Consortium et al. (2015). Compared with the VISAGE Enhanced Tool and allele frequency-based approaches (Phillips et al., 2007b; Kosoy et al., 2009; Nassir et al., 2009; Kidd et al., 2014; Phillips et al., 2014a), our marker sets achieve substantially higher accuracy within Europe: For four European populations, accuracy improves from 68.2% (VISAGE, 104 markers) to 73.7% (100 new markers) and 82.3% (200 new markers). For five populations, accuracy rises from 56.1% (VISAGE) to 64.5% (100 new markers). Our results show that tailored marker selection markedly improves intra-continental classification. While optimized here for Europe, the method can be applied to any region with sufficient training data.

genetics↗

Advancing Biogeographical Ancestry Predictions Through Machine Learning

Tools like Snipper or the Admixture Model count as state-of-the-art methods in forensic science for biogeographical ancestry. However, they have not been systematically compared to classifiers widely used in other disciplines. Noting that genetic data have a tabular form, this study addresses this gap by benchmarking forensic classifiers against TabPFN, a cutting-edge, general-purpose machine learning classifier for tabular data. The comparison evaluates performance using metrics such as accuracy--the proportion of correct classifications--and ROC AUC. We examine classification tasks for individuals at both the intracontinental and continental levels, based on a published dataset for training and testing. Our results reveal significant performance differences between methods, with TabPFN consistently achieving the best results for accuracy, ROC AUC and log loss. E.g., for accuracy, TabPFN improves SNIPPER from 84% to 93% on a continental scale using eight populations, and from 43% to 48% for inter-European classification with ten populations.

genetics↗

RNA-Protein Interaction Classification via Sequence Embeddings

RNA-protein interactions (RPI) are ubiquitous in cellular organisms and essential for gene regulation. In particular, protein interactions with non-coding RNAs (ncRNAs) play a critical role in these processes. Experimental analysis of RPIs is time-consuming and expensive, and existing computational methods rely on small and limited datasets. This work introduces RNAInterAct, a comprehensive RPI dataset, alongside RPIembeddor, a novel transformer-based model designed for classifying ncRNA-protein interactions. By leveraging two foundation models for sequence embedding, we incorporate essential structural and functional insights into our task. We demonstrate RPIembeddors strong performance and generalization capability compared to state-of-the-art methods across different datasets and analyze the impact of the proposed embedding strategy on the performance in an ablation study.

bioinformatics↗

KinPFN: Bayesian Approximation of RNA FoldingKinetics using Prior-Data Fitted Networks

AO_SCPLOWBSTRACTC_SCPLOWRNA is a dynamic biomolecule crucial for cellular regulation, with its function largely determined by its folding into complex structures, while misfolding can lead to multifaceted biological sequelae. During the folding process, RNA traverses through a series of intermediate structural states, with each transition occurring at variable rates that collectively influence the time required to reach the functional form. Understanding these folding kinetics is vital for predicting RNA behavior and optimizing applications in synthetic biology and drug discovery. While in silico kinetic RNA folding simulators are often computationally intensive and time-consuming, accurate approximations of the folding times can already be very informative to assess the efficiency of the folding process. In this work, we present KinPFN, a novel approach that leverages prior-data fitted networks to directly model the posterior predictive distribution of RNA folding times. By training on synthetic data representing arbitrary prior folding times, KinPFN efficiently approximates the cumulative distribution function of RNA folding times in a single forward pass, given only a few initial folding time examples. Our method offers a modular extension to existing RNA kinetics algorithms, promising significant computational speed-ups orders of magnitude faster, while achieving comparable results. We showcase the effectiveness of KinPFN through extensive evaluations and real-world case studies, demonstrating its potential for RNA folding kinetics analysis, its practical relevance, and generalization to other biological data.

bioinformatics↗

Towards Generative RNA Design With Tertiary Interactions

AO_SCPLOWBSTRACTC_SCPLOWThe function of an RNA molecule depends on its structure and a strong structure-to-function relationship is already achieved on the secondary structure level of RNA. Therefore, the secondary structure based design of RNAs is one of the major challenges in computational biology. A common approach to RNA design is inverse RNA folding. However, existing RNA design methods cannot invert all folding algorithms because they cannot represent all types of base interactions. In this work, we propose RNAinformer, a novel generative transformer based approach to the inverse RNA folding problem. Leveraging axial-attention, we directly model the secondary structure input represented as an adjacency matrix in a 2D latent space, which allows us to invert all existing secondary structure prediction algorithms. Consequently, RNAinformer is the first model capable of designing RNAs from secondary structures with all base interactions, including non-canonical base pairs and tertiary interactions like pseudoknots and base multiplets. We demonstrate RNAinformers state-of-the-art performance across different RNA design benchmarks and showcase its novelty by inverting different RNA secondary structure prediction algorithms.

bioinformatics↗

RNAformer: A Simple Yet Effective Deep Learning Model for RNA Secondary Structure Prediction

AO_SCPLOWBSTRACTC_SCPLOWPredicting RNA secondary structure is essential for understanding RNA function and developing RNA-based therapeutics. Despite recent advances in deep learning for structural biology, its application to RNA secondary structure prediction remains contentious. A primary concern is the control of homology between training and test data. Moreover, deep learning approaches often incorporate complex multi-model systems, ensemble strategies, or require external data. Here, we present the RNAformer, a scalable axial-attention-based deep learning model designed to predict secondary structure directly from a single RNA sequence without additional requirements. We demonstrate the benefits of this lean architecture by learning an accurate biophysical RNA folding model using synthetic data. Trained on experimental data, our model overcomes previously reported caveats in deep learning approaches with a novel homology-aware data pipeline. The RNAformer achieves state-of-the-art performance on RNA secondary structure prediction, out-performing both traditional non-learning-based methods and existing deep learning approaches, while carefully considering sequence and structure similarities.

bioinformatics↗

RnaBench: A Comprehensive Library for In Silico RNA Modelling

RNA is a crucial regulator in living organisms and malfunctions can lead to severe diseases. To explore RNA-based therapeutics and applications, computational structure prediction and design approaches play a vital role. Among these approaches, deep learning (DL) algorithms show great promise. However, the adoption of DL methods in the RNA community is limited due to various challenges. DL practitioners often underestimate data homologies, causing skepticism in the field. Additionally, the absence of standardized benchmarks hampers result comparison, while tackling low-level tasks requires significant effort. Moreover, assessing performance and visualizing results prove to be non-trivial and task-dependent. To address these obstacles, we introduce RnaBench (RnB), an open-source RNA library designed specifically for the development of deep learning algorithms that mitigate the challenges during data generation, evaluation, and visualization. It provides meticulously curated homology-aware RNA datasets and standardized RNA benchmarks, including a pioneering RNA design benchmark suite featuring a novel real-world RNA design problem. Furthermore, RnB offers baseline algorithms, both existing and novel performance measures, as well as data utilities and a comprehensive visualization module, all accessible through a user-friendly interface. By leveraging RnB, DL practitioners can rapidly develop innovative algorithms, potentially revolutionizing the field of computational RNA research.

bioinformatics↗

Partial RNA Design

RNA design is a key technique to achieve new functionality in fields like synthetic biology or biotechnology. Computational tools could help to find such RNA sequences but they are often limited in their formulation of the search space. In this work, we propose partial RNA design, a novel RNA design paradigm that addresses the limitations of current RNA design formulations. Partial RNA design describes the problem of designing RNAs from arbitrary RNA sequences and structure motifs with multiple design goals. By separating the design space from the objectives, our formulation enables the design of RNAs with variable lengths and desired properties, while still allowing precise control over sequence and structure constraints at individual positions. Based on this formulation, we introduce a new algorithm, libLEARNA, capable of efficiently solving different constraint RNA design tasks. A comprehensive analysis of various problems, including a realistic riboswitch design task, reveals the outstanding performance of libLEARNA and its robustness.

bioinformatics↗