bioRxiv Science⌕ Search

Biology subjects

Dingle, K.

Publications and source records attributed to Dingle, K..

6 recordsLinked to original sources

Maximum Mutational Robustness in Genotype-Phenotype Maps Follows a Self-similar Blancmange-like Curve

Phenotype robustness, defined as the average mutational robustness of all the genotypes that map to a given phenotype, plays a key role in facilitating neutral exploration of novel phenotypic variation by an evolving population. By applying results from coding theory, we prove that the maximum phenotype robustness occurs when genotypes are organised as bricklayers graphs, so called because they resemble the way in which a bricklayer would fill in a Hamming graph. The value of the maximal robustness is given by a fractal continuous everywhere but differentiable nowhere sums-of-digits function from number theory. Interestingly, genotype-phenotype (GP) maps for RNA secondary structure and the HP model for protein folding can exhibit phenotype robustness that exactly attains this upper bound. By exploiting properties of the sums-of-digits function, we prove a lower bound on the deviation of the maximum robustness of phenotypes with multiple neutral components from the bricklayers graph bound, and show that RNA secondary structure phenotypes obey this bound. Finally, we show how robustness changes when phenotypes are coarse-grained and derive a formula and associated bounds for the transition probabilities between such phenotypes.

evolutionary biology↗

Genomic Sequencing from Sputum for Tuberculosis Disease Diagnosis, Lineage Determination and Drug Susceptibility Prediction

BackgroundUniversal access to drug susceptibility testing for newly diagnosed tuberculosis patients is recommended. Access to culture-based diagnostics remains limited and targeted molecular assays are vulnerable to emerging resistance conferring mutations. Improved sample preparation protocols for direct-from-sputum sequencing of Mycobacterium tuberculosis would accelerate access to comprehensive drug susceptibility testing and molecular typing. MethodsWe assessed a thermo-protection buffer-based direct-from-sample M. tuberculosis whole-genome sequencing protocol. We prospectively processed and analyzed 60 acid-fast bacilli smear-positive sputum samples from tuberculosis patients in India and Madagascar. A diversity of semi-quantitative smear positivity level samples were included. Sequencing was performed using Illumina and MinION (monoplex and multiplex) technologies. We measured the impact of bacterial inoculum and sequencing platforms on M. tuberculosis genomic mean read depth, drug susceptibility prediction performance and typing accuracy. ResultsM. tuberculosis was identified from 88% (Illumina), 89% (MinION-monoplex) and 83% (MinION-multiplex) of samples for which sufficient DNA could be extracted. The fraction of M. tuberculosis reads from MinION sequencing was lower than from Illumina, but monoplexing grade 3+ sputum samples on MinION produced higher read depth than Illumina (p<0.05) and MinION multiplex (p<0.01). No significant difference in overall sensitivity and specificity of drug susceptibility predictions was seen across these sequencing modalities or within each sequencing technology when stratified by smear grade. Lineage typing agreement percentages between direct and culture-based sequencing were 85% (MinION-monoplex), 88% (Illumina) and 100% (MinION-multiplex) ConclusionsM. tuberculosis direct-from-sample whole-genome sequencing remains challenging. Improved and affordable sample treatment protocols are needed prior to clinical deployment.

microbiology↗

Predicting phenotype transition probabilities via conditional algorithmic probability approximations

Unravelling the structure of genotype-phenotype (GP) maps is an important problem in biology. Recently, arguments inspired by algorithmic information theory (AIT) and Kolmogorov complexity have been invoked to uncover simplicity bias in GP maps, an exponentially decaying upper bound in phenotype probability with increasing phenotype descriptional complexity. This means that phenotypes with very many genotypes assigned via the GP map must be simple, while complex phenotypes must have few genotypes assigned. Here we use similar arguments to bound the probability P (x [->] y) that phenotype x, upon random genetic mutation, transitions to phenotype y. The bound is [Formula], where [Formula] is the estimated conditional complexity of y given x, quantifying how much extra information is required to make y given access to x. This upper bound is related to the conditional form of algorithmic probability from AIT. We demonstrate the practical applicability of our derived bound by predicting phenotype transition probabilities (and other related quantities) in simulations of RNA and protein secondary structures. Our work contributes to a general mathematical understanding of GP maps, and may facilitate the prediction of transition probabilities directly from examining phenotype themselves, without utilising detailed knowledge of the GP map.

evolutionary biology↗

Random and natural non-coding RNA have similar structural motif patterns but can be distinguished by bulge, loop, and bond counts

An important question in evolutionary biology is whether and in what ways genotype-phenotype (GP) map biases can influence evolutionary trajectories. Untangling the relative roles of natural selection and biases (and other factors) in shaping phenotypes can be difficult. Because RNA secondary structure (SS) can be analysed in detail mathematically and computationally, is biologically relevant, and a wealth of bioinformatic data is available, it offers a good model system for studying the role of bias. For quite short RNA (length L [&le;] 126), it has recently been shown that natural and random RNA are structurally very similar, suggesting that bias strongly constrains evolutionary dynamics. Here we extend these results with emphasis on much larger RNA with length up to 3000 nucleotides. By examining both abstract shapes and structural motif frequencies (ie the numbers of helices, bonds, bulges, junctions, and loops), we find that large natural and random structures are also very similar, especially when contrasted to typical structures sampled from the space of all possible RNA structures. Our motif frequency study yields another result, that the frequencies of different motifs can be used in machine learning algorithms to classify random and natural RNA with quite high accuracy, especially for longer RNA (eg ROC AUC 0.86 for L = 1000). The most important motifs for classification are found to be the number of bulges, loops, and bonds. This finding may be useful in using SS to detect candidates for functional RNA within junk DNA regions.

evolutionary biology↗

Symmetry and simplicity spontaneously emerge from the algorithmic nature of evolution

Engineers routinely design systems to be modular and symmetric in order to increase robustness to perturbations and to facilitate alterations at a later date. Biological structures also frequently exhibit modularity and symmetry, but the origin of such trends is much less well understood. It can be tempting to assume - by analogy to engineering design - that symmetry and modularity arise from natural selection. But evolution, unlike engineers, cannot plan ahead, and so these traits must also afford some immediate selective advantage which is hard to reconcile with the breadth of systems where symmetry is observed. Here we introduce an alternative non-adaptive hypothesis based on an algorithmic picture of evolution. It suggests that symmetric structures preferentially arise not just due to natural selection, but also because they require less specific information to encode, and are therefore much more likely to appear as phenotypic variation through random mutations. Arguments from algorithmic information theory can formalise this intuition, leading to the prediction that many genotype-phenotype maps are exponentially biased towards phenotypes with low descriptional complexity. A preference for symmetry is a special case of this bias towards compressible descriptions. We test these predictions with extensive biological data, showing that that protein complexes, RNA secondary structures, and a model gene-regulatory network all exhibit the expected exponential bias towards simpler (and more symmetric) phenotypes. Lower descriptional complexity also correlates with higher mutational robustness, which may aid the evolution of complex modular assemblies of multiple components.

evolutionary biology↗

Phenotype bias determines how RNA structures occupy the morphospace of all possible shapes

Morphospaces representations of phenotypic characteristics are often populated unevenly, leaving large parts unoccupied. Such patterns are typically ascribed to contingency, or else to natural selection disfavouring certain parts of the morphospace. The extent to which developmental bias, the tendency of certain phenotypes to preferentially appear as potential variation, also explains these patterns is hotly debated. Here we demonstrate quantitatively that developmental bias is the primary explanation for the occupation of the morphospace of RNA secondary structure (SS) shapes. Upon random mutations, some RNA SS shapes (the frequent ones) are much more likely to appear than others. By using the RNAshapes method to define coarse-grained SS classes, we can directly compare the frequencies that non-coding RNA SS shapes appear in the RNAcentral database to frequencies obtained upon random sampling of sequences. We show that: a) Only the most frequent structures appear in nature; the vast majority of possible structures in the morphospace have not yet been explored. b) Remarkably small numbers of random sequences are needed to produce all the RNA SS shapes found in nature so far. c) Perhaps most surprisingly, the natural frequencies are accurately predicted, over several orders of magnitude in variation, by the likelihood that structures appear upon uniform random sampling of sequences. The ultimate cause of these patterns is not natural selection, but rather strong phenotype bias in the RNA genotype-phenotype map, a type of developmental bias or "findability constraint", which limits evolutionary dynamics to a hugely reduced subset of structures that are easy to "find".

evolutionary biology↗