bioRxiv ScienceSearch

Biology subjects

Mathews, D.

Publications and source records attributed to Mathews, D..

2 recordsLinked to original sources

Chemically Accurate Relative Folding Stability of RNA Hairpins from Molecular Simulations

This study describes a comparison between melts and simulated stabilities of the same RNAs that could be used to benchmark RNA force fields, and potentially to determine future melt-ing experiments. Using umbrella sampling molecular simulations of three 12-nucleotide RNA hairpin stem loops, for which there are experimentally determined free energies of unfold-ing, we projected unfolding onto the reaction coordinate of end to end (5' to 3' hydroxyl oxygen) distance. We estimate the free energy change of the transition from the native con-formation to a fully extended conformation--the stretched state--with no hydrogen bonds between non-neighboring bases. Each simulation was performed four times using the AM-BER FF99+bsc0+{chi}OL3 force field and each window, spaced at 1 [A] intervals, was sampled for 1 s, for a total of 552 s of simulation. We compared differences in the simulated free energy changes to analogous differences in free energies from optical melting experiments using ther-modynamic cycles where the free energy change between stretched and random coil sequences is assumed to be sequence independent. The differences between experimental and simulated {Delta}{Delta}G{degrees} are on average 1.00 {+/-} 0.66 kcal/mol, which is chemically accurate and suggests analo-gous simulations could be used predictively. We also report a novel method to identify where replica free energies diverge along the reaction coordinate, thus indicating where additional sampling would most improve convergence. We conclude by discussing methods to more economically perform such simulations.

biophysics

LinearFold: Linear-Time Prediction of RNA Secondary Structures

Predicting the secondary structure of an RNA sequence with speed and accuracy is useful in many applications such as drug design. The state-of-the-art predictors have a fundamental limitation: they have a run time that scales cubically with the length of the input sequence, which is slow for longer RNAs and limits the use of secondary structure prediction in genome-wide applications. To address this bottleneck, we designed the first linear-time algorithm for this problem. which can be used with both thermodynamic and machine-learned scoring functions. Our algorithm, like previous work, is based on dynamic programming (DP), but with two crucial differences: (a) we incrementally process the sequence in a left-to-right rather than in a bottom-up fashion, and (b) because of this incremental processing, we can further employ beam search pruning to ensure linear run time in practice (with the cost of exact search). Even though our search is approximate, surprisingly, it results in even higher overall accuracy on a diverse database of sequences with known structures. More interestingly, it leads to significantly more accurate predictions on the longest sequence families in that database (16S and 23S Ribosomal RNAs), as well as improved accuracies for long-range base pairs (500+ nucleotides apart).

bioinformatics