bioRxiv Science⌕ Search

Biology subjects

Murphy, T. M.

Publications and source records attributed to Murphy, T. M..

2 recordsLinked to original sources

Design approaches to expand the toolkit for building cotranscriptionally encoded RNA strand displacement circuits

Cotranscriptionally encoded RNA strand displacement (ctRSD) circuits are an emerging tool for programmable molecular computation with potential applications spanning in vitro diagnostics to continuous computation inside living cells. In ctRSD circuits, RNA strand displacement components are continuously produced together via transcription. These RNA components can be rationally programmed through base pairing interactions to execute logic and signaling cascades. However, the small number of ctRSD components characterized to date limits circuit size and capabilities. Here, we characterize 220 ctRSD gate sequences, exploring different input, output, and toehold sequences and changes to other design parameters, including domain lengths, ribozyme sequences, and the order in which gate strands are transcribed. This characterization provides a library of sequence domains for engineering ctRSD components, i.e., a toolkit, enabling circuits with up to four-fold more inputs than previously possible. We also identify specific failure modes and systematically develop design approaches that reduce the likelihood of failure across different gate sequences. Lastly, we show ctRSD gate design is robust to changes in transcriptional encoding, opening a broad design space for applications in more complex environments. Together, these results deliver an expanded toolkit and design approaches for building ctRSD circuits that will dramatically extend capabilities and potential applications.

synthetic biology↗

A comparison of feature selection methodologies and learning algorithms in the development of a DNA methylation-based telomere length estimator

The field of epigenomics holds great promise in understanding and treating disease with advances in machine learning (ML) and artificial intelligence being vitally important in this pursuit. Increasingly, research now utilises DNA methylation measures at cytosine-guanine dinucleotides (CpG) to detect disease and estimate biological traits such as aging. Given the high dimensionality of DNA methylation data, feature-selection techniques are commonly employed to reduce dimensionality and identify the most important subset of features. In this study, we test and compare a range of feature-selection methods and ML algorithms in the development of a novel DNA methylation-based telomere length (TL) estimator. We found that principal component analysis in advance of elastic net regression led to the overall best performing estimator when evaluated using a nested cross-validation analysis and two independent test cohorts. In contrast, the baseline model of elastic net regression with no prior feature reduction stage performed worst - suggesting a prior feature-selection stage may have important utility. The variance in performance across tested approaches shows that estimators are sensitive to data set heterogeneity and the development of an optimal DNA methylation-based estimator should benefit from the robust methodological approach used in this study. Additionally, we observed that different DNA methylation-based TL estimators, which have few common CpGs, are associated with many of the same biological entities. Moreover, our methodology which utilises a range of feature-selection approaches and ML algorithms could be applied to other biological markers and disease phenotypes, to examine their relationship with DNA methylation and predictive value.

bioinformatics↗