bioRxiv Science⌕ Search

Biology subjects

Znosko, B. M.

Publications and source records attributed to Znosko, B. M..

4 recordsLinked to original sources

Nearest Neighbor Parameters for Estimating the Folding Stability of RNA Including Pseudouridine

Nearest neighbor parameters are widely used in software for estimating the conformational stability of an RNA sequence folding into a specific structure. Folding stability for RNA with canonical nucleotides A, C, G, and U has been widely studied, but the same is not true for most modified nucleotides. In this work, we present a comprehensive set of nearest neighbor parameters for estimating the folding stability of RNA including pseudouridine in helical or loop contexts. These parameters are derived from 210 optical melting experiments involving helices with pseudouridine-A and pseudouridine-G pairs and with pseudouridine in loop motifs. The experiments include sequences with pseudouridine and U in the same strand, including U-A and U-G pairs, allowing us to consider the folding stability of sequences with both U and pseudouridine. On average, pseudouridine stabilizes RNA folding compared to U in an analogous motif, although this effect is sequence-context dependent. These parameters improve the modeling of folding stability for RNA secondary structures containing pseudouridine. We demonstrate that these parameters successfully model the secondary structure change for Saccharomyces cerevisiae U2 snRNA when two additional inducible pseudouridines are present. These parameters are freely available and incorporated into the RNAstructure software package. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=81 SRC="FIGDIR/small/725682v1_ufig1.gif" ALT="Figure 1"> View larger version (14K): org.highwire.dtl.DTLVardef@e1167aorg.highwire.dtl.DTLVardef@18ac7f0org.highwire.dtl.DTLVardef@4c909eorg.highwire.dtl.DTLVardef@aa8bca_HPS_FORMAT_FIGEXP M_FIG C_FIG

biochemistry↗

RNA Folding Nearest Neighbor Parameters Including the Modification 1-Methyl-Pseudouridine

Nearest neighbor analysis is commonly used to estimate RNA folding stabilities. In this contribution, we report a set of RNA folding nearest neighbor parameters for estimating free energy change for RNA sequences including 1-methyl-pseudouridine. Development of mRNA vaccines has identified 1-methyl-pseudouridine as a key nucleobase modification for suppressing innate immune responses. However, the contributions of these modifications to RNA folding stability were unclear. Our new parameters provide helical terms for 1-methyl-pseudouridine-adenine and 1-methyl-pseudouridine-guanine base pairs. The parameters also estimate loop stabilities for loops with 1-methyl-pseudouridine or a combination of 1-methyl-pseudouridine and uridine. These parameters are derived using 208 optical melting experiments and tested against an additional 16 optical melting experiments. On average, we find that substitution of uridine with 1-methyl-pseudouridine stabilizes RNA folding, with the extent of stabilization depending on adjacent sequence. The estimation of tRNA folding ensembles for tRNA sequences with 1-methyl-pseudouridine was significantly improved using the new nearest neighbor parameters. The new nearest neighbor parameters are provided as part of the RNAstructure software package. With these parameters, the secondary structures of natural sequences with 1-methyl-pseudouridine and mRNA therapeutics fully substituted with 1-methyl-pseudouridine can be modeled.

bioinformatics↗

A Large-Scale Cryo-EM RNA Motif Dataset and Benchmark for Machine Learning-Based Structure Modeling

MotivationRNA molecules play critical roles in gene regulation, viral replication, and cellular control, with their functions tightly coupled to three-dimensional structure. Advances in cryogenic electron microscopy (cryo-EM) now enable RNA structure characterization across a broad resolution range. RNA secondary structural motifs, including hairpins, internal loops, and bulges, act as fundamental building blocks of RNA tertiary architecture and are key targets in RNA-focused therapeutic design. Despite this, most computational approaches for RNA structure prediction from cryo-EM density maps do not explicitly utilize secondary structural motifs as intermediate representations, largely due to the absence of large-scale, high-quality, and motif-resolved datasets suitable for machine learning. ResultsHere, we present a large, open-source dataset containing over 125,000 motif-resolved cryo-EM density maps paired with corresponding atomic structures, spanning 25 classes of RNA secondary structural motifs. The dataset covers resolutions from 1.5 [A] to 34.0 [A], encompassing both near-atomic and low-resolution density maps relevant to RNA modeling. Each motif instance includes a segmented cryo-EM density map represented as a standardized 3D voxel grid, with atomic-level motif annotations propagated to voxel-level labels for RNA backbone, ribose sugar, and nucleobase components. Segmentation quality is validated via cross-correlation analysis, demonstrating strong agreement between motif-level density maps and atomic reference models. To illustrate the datasets utility, high-resolution maps (1.5-2.8 [A]) were used to train a machine learning classifier that distinguished five motif classes with a specificity of 0.948. Availability and ImplementationSource code, implementation of the fully automated pipeline, and the benchmark datasets are publicly available at GitHubhttps://github.com/DrDongSi/3DEM-RNA-Motif-Dataset Zenodohttps://zenodo.org/communities/3dem-rna-motif-dataset Contacthoujie@msu.edu, dongsi@uw.edu

bioinformatics↗

Exploring the Efficiency of Deep Graph Neural Networks for RNA Secondary Structure Prediction

Ribonucleic acid (RNA) plays a vital role in various biological processes and forms intricate secondary and tertiary structures associated with its functions. Predicting RNA secondary structures is essential for understanding the functional and regulatory roles of RNA molecules in biological processes. Traditional free-energy-based methods for predicting these structures often fail to capture complex interactions and long-range dependencies within RNA sequences. Recent advancements in machine learning, particularly with graph neural networks (GNNs), have shown promise in enhancing the ability to model the relationships between molecular sequences and their structures. This work specifically explores the efficacy of various GNN architectures in modeling RNA secondary structure. Through benchmarking the GNN methods against traditional energy-based models on standard datasets, our analysis demonstrates that GNN models improves traditional methods, offering a robust framework for accurate RNA structure prediction.

bioinformatics↗