bioRxiv ScienceSearch

Biology subjects

Das, R.

Publications and source records attributed to Das, R..

At least 19 recordsLinked to original sources

Multicenter validation of a machine learning algorithm for 48 hour all-cause mortality prediction

PurposeThis study evaluates a machine-learning-based mortality prediction tool.\n\nMaterials and MethodsWe conducted a retrospective study with data drawn from three academic health centers. Inpatients of at least 18 years of age and with at least one observation of each vital sign were included. Predictions were made at 12, 24, and 48 hours before death. Models fit to training data from each institution were evaluated on hold-out test data from the same institution and data from the remaining institutions. Predictions were compared to those of qSOFA and MEWS using area under the receiver operating characteristic curve (AUROC).\n\nResultsFor training and testing on data from a single institution, machine learning predictions averaged AUROCs of 0.97, 0.96, and 0.95 across institutional test sets for 12-, 24-, and 48-hour predictions, respectively. When trained and tested on data from different hospitals, the algorithm achieved AUROC up to 0.95, 0.93, and 0.91, for 12-, 24-, and 48-hour predictions, respectively. MEWS and qSOFA had average 48-hour AUROCs of 0.86 and 0.82, respectively.\n\nConclusionThis algorithm may help identify patients in need of increased levels of clinical care.

bioinformatics

A quantitative and predictive model for RNA binding by human Pumilio proteins

High-throughput methodologies have enabled routine generation of RNA target sets and sequence motifs for RNA-binding proteins (RBPs). Nevertheless, quantitative approaches are needed to capture the landscape of RNA/RBP interactions responsible for cellular regulation. We have used the RNA-MaP platform to directly measure equilibrium binding for thousands of designed RNAs and to construct a predictive model for RNA recognition by the human Pumilio proteins PUM1 and PUM2. Despite prior findings of linear sequence motifs, our measurements revealed widespread residue flipping and instances of positional coupling. Application of our thermodynamic model to published in vivo crosslinking data reveals quantitative agreement between predicted affinities and in vivo occupancies. Our analyses suggest a thermodynamically driven, continuous Pumilio binding landscape that is negligibly affected by RNA structure or kinetic factors, such as displacement by ribosomes. This work provides a quantitative foundation for dissecting the cellular behavior of RBPs and cellular features that impact their occupancies.

biochemistry

RNA tertiary structure energetics predicted by an ensemble model of the RNA double helix

Over 50% of residues within functional structured RNAs are base-paired in Watson-Crick helices, but it is not fully understood how these helices geometric preferences and flexibility might influence RNA tertiary structure. Here, we show experimentally and computationally that the ensemble fluctuations of RNA helices substantially impact RNA tertiary structure stability. We updated a model for the conformational ensemble of the RNA helix using crystallographic structures of Watson-Crick base pair steps. To test this model, we made blind predictions of the thermodynamic stability of >1500 tertiary assemblies with differing helical sequences and compared calculations to independent measurements from a high-throughput experimental platform. The blind predictions accounted for thermodynamic effects from changing helix sequence and length with unexpectedly tight accuracies (RMSD of 0.34 and 0.77 kcal/mol, respectively). These comparisons lead to a detailed picture of how RNA base pair steps fluctuate within complex assemblies and suggest a new route toward predicting RNA tertiary structure formation and energetics.

biophysics

Sampling native-like structures of RNA-protein complexes through Rosetta folding and docking

RNA-protein complexes underlie numerous cellular processes including translation, splicing, and posttranscriptional regulation of gene expression. The structures of these complexes are crucial to their functions but often elude high-resolution structure determination. Computational methods are needed that can integrate low-resolution data for RNA-protein complexes while modeling de novo the large conformational changes of RNA components upon complex formation. To address this challenge, we describe a Rosetta method called RNP-denovo to simultaneously fold and dock RNA to a protein surface. On a benchmark set of structurally diverse RNA-protein complexes that are not solvable with prior strategies, this fold-and-dock method consistently sampled native-like structures with better than nucleotide resolution. We revisited three past blind modeling challenges in which previous methods gave poor results: human telomerase, an RNA methyltransferase with a ribosomal RNA domain, and the spliceosome. When coupled with the same sparse FRET, cross-linking, and functional data used in previous work, RNP-denovo gave models with significantly improved accuracy. These results open a route to computationally modeling global folds of RNA-protein complexes from low-resolution data.

biophysics

De novo computational RNA modeling into cryoEM maps of large ribonucleoprotein complexes

RNA-protein assemblies carry out many critical biological functions including translation, RNA splicing, and telomere extension. Increasingly, cryo-electron microscopy (cryoEM) is used to determine the structures of these complexes, but nearly all maps determined with this method have regions in which the local resolution does not permit manual coordinate tracing. Because RNA coordinates typically cannot be determined by docking crystal structures of separate components and existing structure prediction algorithms cannot yet model RNA-protein complexes, RNA coordinates are frequently omitted from final models despite their biological importance. To address these omissions, we have developed a new framework for De novo Ribonucleoprotein modeling in Real-space through Assembly of Fragments Together with Electron density in Rosetta (DRRAFTER). We show that DRRAFTER recovers near-native models for a diverse benchmark set of small RNA-protein complexes, as well as for large RNA-protein machines, including the spliceosome, mitochondrial ribosome, and CRISPR-Cas9-sgRNA complexes where the availability of both high and low resolution maps enable rigorous tests. Blind tests on yeast U1 snRNP and spliceosomal P complex maps demonstrate that the method can successfully build RNA coordinates in real-world modeling scenarios. Additionally, to aid in final model interpretation, we present a method for reliable in situ estimation of DRRAFTER model accuracy. Finally, we apply this method to recently determined maps of telomerase, the HIV-1 reverse transcriptase initiation complex, and the packaged MS2 genome, demonstrating that DRRAFTER can be used to accelerate accurate model building in challenging cases.

biophysics

Ancient ancestry informative markers for identifying fine-scale ancient population structure in Eurasians

The rapid accumulation of ancient human genomes from various areas and time periods potentially allows the expansion of studies of biodiversity, biogeography, forensics, population history, and epidemiology into past populations. However, most ancient DNA (aDNA) data were generated through microarrays designed for modern-day populations known to misrepresent the population structure. Past studies addressed these problems using ancestry informative markers (AIMs). However, it is unclear whether AIMs derived from contemporary human genomes can capture ancient population structure and whether AIM finding methods are applicable to ancient DNA (aDNA) provided that the high missingness rates in ancient, oftentimes haploid, DNA can also distort the population structure. Here, we define ancient AIMs (aAIMs) and develop a framework to evaluate established and novel AIM-finding methods in identifying the most informative markers. We show that aAIMs identified by a novel principal component analysis (PCA)-based method outperforms all competing methods in classifying ancient individuals into populations and identifying admixed individuals. In some cases, predictions made using the aAIMs were more accurate than those made with a complete marker set. We discuss the features of the ancient Eurasian population structure and strategies to identify aAIMs. This work informs the design of population microarrays and the interpretation of aDNA results.

genetics

EternaBrain: Automated RNA design through move sets from an Internet-scale RNA videogame

Emerging RNA-based approaches to disease detection and gene therapy require RNA sequences that fold into specific base-pairing patterns, but computational algorithms generally remain inadequate for these secondary structure design tasks. The Eterna project has crowdsourced RNA design to human video game players in the form of puzzles that reach extraordinary difficulty. Here, we present an eternamoves-large repository consisting of 1.8 million of player moves on 12 of the most-played Eterna puzzles as well as an eternamoves-select repository of 30,477 moves from the top 72 players on a select set of more advanced puzzles. On eternamoves-select, a multilayer convolutional neural network (CNN) EternaBrain achieves test accuracies of 51% and 34% in base prediction and location prediction, respectively, suggesting that top players moves are partially stereotyped. We then show that while this CNNs move predictions are not enough to solve numerous new puzzles, inclusion of six additional strategies compiled by human players solves 61 out of 100 independent puzzles in the Eterna100 benchmark. This EternaBrain-SAP performance is better than previously published methods and in the middle of the performance range of newer algorithms developed by Eterna participants and other groups. Our study provides useful lessons for efforts to achieve human-competitive performance with automated RNA design algorithms.

bioinformatics

LIkelihood-based Fits of Folding Transitions (LIFFT) for Biomolecule Mapping Data

SummaryBiomolecules shift their structures as a function of temperature and concentrations of protons, ions, small molecules, proteins, and nucleic acids. These transitions impact or underlie biological function and are being monitored at increasingly high throughput. For example, folding transitions for large collections of RNAs can now be monitored at single residue resolution by chemical mapping techniques. LIkelihood-based Fits of Folding Transitions (LIFFT) quantifies these data through well-defined thermodynamic models. LIFFT implements a Bayesian framework that takes into account data at all measured residues and enables visual assessment of modeling uncertainties that can be overlooked in least-squares fits. The framework is appropriate for multimodal techniques ranging from chemical mapping including multi-wavelength spectroscopy.\n\nAvailabilityFreely available MATLAB package at https://ribokit.stanford.edu/LIFFT/.\n\nContactrhiju@stanford.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

biochemistry

Multicenter validation of a sepsis prediction algorithm using only vital sign data in the emergency department, general ward and ICU

ObjectivesWe validate a machine learning-based sepsis prediction algorithm (InSight) for detection and prediction of three sepsis-related gold standards, using only six vital signs. We evaluate robustness to missing data, customization to site-specific data using transfer learning, and generalizability to new settings.\n\nDesignA machine learning algorithm with gradient tree boosting. Features for prediction were created from combinations of only six vital sign measurements and their changes over time.\n\nSettingA mixed-ward retrospective data set from the University of California, San Francisco (UCSF) Medical Center (San Francisco, CA) as the primary source, an intensive care unit data set from the Beth Israel Deaconess Medical Center (Boston, MA) as a transfer learning source, and four additional institutions datasets to evaluate generalizability.\n\nParticipants684,443 total encounters, with 90,353 encounters from June 2011 to March 2016 at UCSF.\n\nInterventionsnone\n\nPrimary and secondary outcome measuresArea under the receiver operating characteristic curve (AUROC) for detection and prediction of sepsis, severe sepsis, and septic shock.\n\nResultsFor detection of sepsis and severe sepsis, InSight achieves an area under the receiver operating characteristic (AUROC) curve of 0.92 (95% CI 0.90 - 0.93) and 0.87 (95% CI 0.86 - 0.88), respectively. Four hours before onset, InSight predicts septic shock with an AUROC of 0.96 (95% CI 0.94 -0.98), and severe sepsis with an AUROC of 0.85 (95% CI 0.79 - 0.91).\n\nConclusionsInSight outperforms existing sepsis scoring systems in identifying and predicting sepsis, severe sepsis, and septic shock. This is the first sepsis screening system to exceed an AUROC of 0.90 using only vital sign inputs. InSight is robust to missing data, can be customized to novel hospital data using a small fraction of site data, and retained strong discrimination across all institutions.\n\nStrengths and limitations of this studyO_LIMachine learning is applied to the detection and prediction of three separate sepsis standards in the emergency department, general ward and intensive care settings.\nC_LIO_LIOnly six commonly measured vital signs are used as input for the algorithm.\nC_LIO_LIThe algorithm is robust to randomly missing data.\nC_LIO_LITransfer learning successfully leverages large dataset information to a target dataset.\nC_LIO_LIRetrospective nature of the study does not predict clinician reaction to information.\nC_LI

bioinformatics

Prospects for recurrent neural network models to learn RNA biophysics from high-throughput data

RNA is a functionally versatile molecule that plays key roles in genetic regulation and in emerging technologies to control biological processes. Computational models of RNA secondary structure are well-developed but often fall short in making quantitative predictions of the behavior of multi-RNA complexes. Recently, large datasets characterizing hundreds of thousands of individual RNA complexes have emerged as rich sources of information about RNA energetics. Meanwhile, advances in machine learning have enabled the training of complex neural networks from large datasets. Here, we assess whether a recurrent neural network model, Ribonet, can learn from high-throughput binding data, using simulation and experimental studies to test model accuracy but also determine if they learned meaningful information about the biophysics of RNA folding. We began by evaluating the model on energetic values predicted by the Turner model to assess whether the neural network could learn a representation that recovered known biophysical principles. First, we trained Ribonet to predict the simulated free energy of an RNA in complex with multiple input RNAs. Our model accurately predicts free energies of new sequences but also shows evidence of having learned base pairing information, as assessed by in silico double mutant analysis. Next, we extended this model to predict the simulated affinity between an arbitrary RNA sequence and a reporter RNA. While these more indirect measurements precluded the learning of basic principles of RNA biophysics, the resulting model achieved sub-kcal/mol accuracy and enabled design of simple RNA input responsive riboswitches with high activation ratios predicted by the Turner model from which the training data were generated. Finally, we compiled and trained on an experimental dataset comprising over 600,000 experimental affinity measurements published on the Eterna open laboratory. Though our tests revealed that the model likely did not learn a physically realistic representation of RNA interactions, it nevertheless achieved good performance of 0.76 kcal/mol on test sets with the application of transfer learning and novel sequence-specific data augmentation strategies. These results suggest that recurrent neural network architectures, despite being naive to the physics of RNA folding, have the potential to capture complex biophysical information. However, more diverse datasets, ideally involving more direct free energy measurements, may be necessary to train de novo predictive models that are consistent with the fundamentals of RNA biophysics.\n\nAuthor SummaryThe precise design of RNA interactions is essential to gaining greater control over RNA-based biotechnology tools, including designer riboswitches and CRISPR-Cas9 gene editing. However, the classic model for energetics governing these interactions fails to quantitatively predict the behavior of RNA molecules. We developed a recurrent neural network model, Ribonet, to quantitatively predict these values from sequence alone. Using simulated data, we show that this model is able to learn simple base pairing rules, despite having no a priori knowledge about RNA folding encoded in the network architecture. This model also enables design of new switching RNAs that are predicted to be effective by the \"ground truth\" simulated model. We applied transfer learning to retrain Ribonet using hundreds of thousands of RNA-RNA affinity measurements and demonstrate simple data augmentation techniques that improve model performance. At the same time, data diversity currently available set limits on Ribonets accuracy. Recurrent neural networks are a promising tool for modeling nucleic acid biophysics and may enable design of complex RNAs for novel applications.

biophysics

Formin3 regulates dendritic architecture via microtubule stabilization and is required for somatosensory nociceptive behavior

The acquisition, maintenance and modulation of dendritic architecture are critical to neuronal form, plasticity and function. Morphologically, dendritic shape impacts functional connectivity and is largely mediated by organization and dynamics of cytoskeletal fibers that provide the underlying scaffold and tracks for intracellular trafficking. Identifying molecular factors that regulate dendritic cytoskeletal architecture is therefore important in understanding mechanistic links between cytoskeletal organization and neuronal function. In a neurogenomic-driven genetic screen of cytoskeletal regulatory molecules, we identified Formin3 (Form3) as a critical regulator of cytoskeletal architecture in Drosophila nociceptive sensory neurons. Form3 is a member of the conserved Formin family of multi-functional cytoskeletal regulators and time course analyses reveal Form3 is cell-autonomously required for maintenance of complex dendritic arbors. Cytoskeletal imaging demonstrates form3 mutants exhibit a specific destabilization of the dendritic microtubule (MT) cytoskeleton, together with defective dendritic trafficking of mitochondria, satellite Golgi and the TRPA channel Painless. Biochemical studies reveal Form3 directly interacts with MTs via FH1-FH2 domains and promotes MT stabilization via acetylation. Neurologically, mutations in human Inverted Formin 2 (INF2; ortholog of form3) have been causally linked to Charcot-Marie-Tooth (CMT) disease. CMT sensory neuropathies lead to impaired peripheral sensitivity. Defects in form3 function in nociceptive neurons results in a severe impairment in noxious heat evoked behaviors. Expression of the INF2 FH1-FH2 domains rescues form3 defects in MT stabilization and nocifensive behavior revealing conserved functions in regulating the cytoskeleton and sensory behavior thereby providing novel mechanistic insights into potential etiologies of CMT sensory neuropathies.\n\nSignificance StatementMechanisms governing cytoskeletal architecture are critical in regulating neural function as aberrations are linked to a broad spectrum of neurological and neurocognitive disorders. Formins are important cytoskeletal regulators however their mechanistic roles in neuronal architecture are poorly understood. We demonstrate mutations in Drosophila formin3 lead to progressive destabilization of the dendritic microtubule cytoskeleton resulting in severely reduced arborization coupled to impaired organelle and ion channel trafficking, as well as nociceptive sensitivity. INF2 mutations are implicated in CMT sensory neuropathies, and INF2 expression can rescue microtubule and nociceptive behavioral defects in form3 mutants. While CMT sensory neuropathies have been linked to defects in axonal development and myelination, our studies connect dendritic cytoskeletal defects with peripheral insensitivity suggesting possible alternative etiological bases.

neuroscience

Evaluating a sepsis prediction machine learning algorithm using minimal electronic health record data in the emergency department and intensive care unit

IntroductionSepsis is a major health crisis in US hospitals, and several clinical identification systems have been designed to help care providers with early diagnosis of sepsis. However, many of these systems demonstrate low specificity or sensitivity, which limits their clinical utility. We evaluate the effects of a machine learning algodiagnostic (MLA) sepsis prediction and detection system using a before-and-after clinical study performed at Cabell Huntington Hospital (CHH) in Huntington, West Virginia. Prior to this study, CHH utilized the St. Johns Sepsis Agent (SJSA) as a rules-based sepsis detection system.\n\nMethodsThe Predictive algoRithm for EValuation and Intervention in SEpsis (PREVISE) study was carried out between July 1, 2017 and August 30, 2017. All patients over the age of 18 who were admitted to the emergency department or intensive care units at CHH were monitored during the study. We assessed pre-implementation baseline metrics during the month of July, 2017, when the SJSA was active. During implementation in the month of August, 2017, SJSA and the MLA concurrently monitored patients for sepsis risk. At the conclusion of the study period, the primary outcome of sepsis-related in-hospital mortality and secondary outcome of sepsis-related hospital length of stay were compared between the two groups.\n\nResultsSepsis-related in-hospital mortality decreased from 3.97% to 2.64%, a 33.5% relative decrease (P = 0.038), and sepsis-related length of stay decreased from 2.99 days in the pre-implementation phase to 2.48 days in the post-implementation phase, a 17.1% relative reduction (P < 0.001).\n\nConclusionReductions in patient mortality and length-of-stay were observed with use of a machine learning algorithm for early sepsis detection in the emergency department and intensive care units at Cabell Huntington Hospital, and may present a method for improving patient outcomes.\n\nTrial RegistrationClinicalTrials.gov, NCT03235193, retrospectively registered on July 27th 2017.

clinical trials

Pediatric Severe Sepsis Prediction Using Machine Learning

Early detection of pediatric severe sepsis is necessary in order to administer effective treatment. In this study, we assessed the efficacy of a machine-learning-based prediction algorithm applied to electronic healthcare record (EHR) data for the prediction of severe sepsis onset. The resulting prediction performance was compared with the Pediatric Logistic Organ Dysfunction score (PELOD-2) and pediatric Systemic Inflammatory Response Syndrome score (SIRS) using cross-validation and pairwise t-tests. EHR data were collected from a retrospective set of de-identified pediatric inpatient and emergency encounters drawn from the University of California San Francisco (UCSF) Medical Center, with encounter dates between June 2011 and March 2016. Patients (n = 11,127) were 2-17 years of age and 103 [0.93%] were labeled severely septic. In four-fold cross-validation evaluations, the machine learning algorithm achieved an AUROC of 0.912 for discrimination between severely septic and control pediatric patients at onset and AUROC of 0.727 four hours before onset. Under the same measure, the prediction algorithm also significantly outperformed PELOD-2 (p < 0.05) and SIRS (p < 0.05) in the prediction of severe sepsis four hours before onset. This machine learning algorithm has the potential to deliver high-performance severe sepsis detection and prediction for pediatric inpatients.

bioinformatics

Blind prediction of noncanonical RNA structure at atomic accuracy

Prediction of RNA structure from nucleotide sequence remains an unsolved grand challenge of biochemistry and requires distinct concepts from protein structure prediction. Despite extensive algorithmic development in recent years, modeling of noncanonical base pairs of new RNA structural motifs has not been achieved in blind challenges. We report herein a stepwise Monte Carlo (SWM) method with a unique add-and-delete move set that enables predictions of noncanonical base pairs of complex RNA structures. A benchmark of 82 diverse motifs establishes the methods general ability to recover noncanonical pairs ab initio, including multistrand motifs that have been refractory to prior approaches. In a blind challenge, SWM models predicted nucleotide-resolution chemical mapping and compensatory mutagenesis experiments for three in vitro selected tetraloop/receptors with previously unsolved structures (C7.2, C7.10, and R1). As a final test, SWM blindly and correctly predicted all noncanonical pairs of a Zika virus double pseudoknot during a recent community-wide RNA-puzzle. Stepwise structure formation, as encoded in the SWM method, enables modeling of noncanonical RNA structure in a variety of previously intractable problems.

biochemistry

Prediction of Acute Kidney Injury with a Machine Learning Algorithm using Electronic Health Record Data

BackgroundA major problem in treating acute kidney injury (AKI) is that clinical criteria for recognition are markers of established kidney damage or impaired function; treatment before such damage manifests is desirable. Clinicians could intervene during what may be a crucial stage for preventing permanent kidney injury if patients with incipient AKI and those at high risk of developing AKI could be identified.\n\nMethodsWe used a machine learning technique, boosted ensembles of decision trees, to train an AKI prediction tool on retrospective data from inpatients at Stanford Medical Center and intensive care unit patients at Beth Israel Deaconess Medical Center. We tested the algorithms ability to detect AKI at onset, and to predict AKI 12, 24, 48, and 72 hours before onset, and compared its 3-fold cross-validation performance to the SOFA score for AKI identification in terms of Area Under the Receiver Operating Characteristic (AUROC).\n\nResultsThe prediction algorithm achieves AUROC of 0.872 (95% CI 0.867, 0.878) for AKI onset detection, superior to the SOFA score AUROC of 0.815 (P < 0.01). At 72 hours before onset, the algorithm achieves AUROC of 0.728 (95% CI 0.719, 0.737), compared to the SOFA score AUROC of 0.720 (P < 0.01).\n\nConclusionsThe results of these experiments suggest that a machine-learning-based AKI prediction tool may offer important prognostic capabilities for determining which patients are likely to suffer AKI, potentially allowing clinicians to intervene before kidney damage manifests.

bioinformatics

Computational Design of Asymmetric Three-dimensional RNA Structures and Machines.

The emerging field of RNA nanotechnology seeks to create nanoscale 3D machines by repurposing natural RNA modules, but successes have been limited to symmetric assemblies of single repeating motifs. We present RNAMake, a suite that automates design of RNA molecules with complex 3D folds. We first challenged RNAMake with the paradigmatic problem of aligning a tetraloop and sequence-distal receptor, previously only solved via symmetry. Single-nucleotide-resolution chemical mapping, native gel electrophoresis, and solution x-ray scattering confirmed that 11 of the 16 miniTTR designs successfully achieved clothespin-like folds. A 2.55 [A] diffraction-resolution crystal structure of one design verified formation of the target asymmetric nanostructure, with large sections achieving near-atomic accuracy (< 2.0 [A]). Finally, RNAMake designed asymmetric segments to tether the 16S and 23S rRNAs together into a synthetic singlestranded ribosome that remains uncleaved by ribonucleases and supports life in Escherichia coli, a challenge previously requiring several rounds of trial-and-error.

bioengineering

RNA structure inference through chemical mapping after accidental or intentional mutations

Despite the critical roles RNA structures play in regulating gene expression, sequencing-based methods for experimentally determining RNA base pairs have remained inaccurate. Here, we describe a multidimensional chemical mapping method called M2-seq (mutate-and-map read out through next-generation sequencing) that takes advantage of sparsely mutated nucleotides to induce structural perturbations at partner nucleotides and then detects these events through dimethyl sulfate (DMS) probing and mutational profiling. In special cases, fortuitous errors introduced during DNA template preparation and RNA transcription are sufficient to give M2-seq helix signatures; these signals were previously overlooked or mistaken for correlated double DMS events. When mutations are enhanced through error-prone PCR, in vitro M2-seq experimentally resolves 33 of 68 helices in diverse structured RNAs including ribozyme domains, riboswitch aptamers, and viral RNA domains with a single false positive. These inferences do not require energy minimization algorithms and can be made by either direct visual inspection or by a new neural-net-inspired algorithm called M2-net. Measurements on the P4-P6 domain of the Tetrahymena group I ribozyme embedded in Xenopus egg extract demonstrate the ability of M2-seq to detect RNA helices in a complex biological environment.\n\nSIGNIFICANCE STATEMENTThe intricate structures of RNA molecules are crucial to their biological functions but have been difficult to accurately characterize. Multidimensional chemical mapping methods improve accuracy but have so far involved painstaking experiments and reliance on secondary structure prediction software. A methodology called M2-seq now lifts these limitations. Mechanistic studies clarify the origin of serendipitous M2-seq-like signals that were recently discovered but not correctly explained and also provide mutational strategies that enable robust M2-seq for new RNA transcripts. The method detects dozens of Watson-Crick helices across diverse RNA folds in vitro and within frog egg extract, with low false positive rate (< 5%). M2-seq opens a route to unbiased discovery of RNA structures in vitro and beyond.

biochemistry

Allosteric Logic of the V. vulnificus Adenine Riboswitch Resolved by Four-dimensional Chemical Mapping

The structural logic that define the functions of gene regulatory RNA molecules may be radically different from classic models of allostery, but the relevant structural correlations have remained elusive in even intensively studied RNA model systems. Here, we present a four-dimensional expansion of chemical mapping called lock-mutate-map-rescue (LM2R), which integrates multiple layers of mutation with nucleotide-resolution chemical mapping. This technique resolves the core mechanism of the adenine-responsive V. vulnificus add riboswitch including its gene expression platform, a paradigmatic system for which both Monod-Wyman-Changeux (MWC) conformational selection models and non-MWC alternatives have been proposed. To discriminate amongst these models, we locked each functionally important helix through designed mutations and assessed formation or depletion of other helices via high-throughput compensatory rescue experiments. These LM2R measurements give strong support to the pre-existing correlations predicted by MWC models, disfavor alternative models, and reveal new structural heterogeneities that may be general across ligand-free riboswitches.

biochemistry