bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.06.23.734074

A substrate recursion principle for biological information, with empirical anchoring through a templating-mode taxonomy

Abstract

Biological inheritance can be treated as a class of catalytic templating reactions in which a daughter molecule, generated by a kinetic kernel acting on a parent template, is itself a substrate for the next round of the same catalysis. We give the physicochemical conditions under which such a reaction can support unbounded heritable molecular distinguishability. Four conditions on the template-operator pair organize the analysis: nonzero per-site information content under the activesite recognition kernel (R1), a count of independently variable recognized positions that grows without bound as the reaction extends (R2), catalytic closure under iteration, possibly through a reversible involution such as Watson-Crick complementation (R3), and stochastic drift of the kernel in its recognition alphabet (R4). A fifth, scope-defining condition (R5) restricts the principle to kernels that are intrinsic physicochemistry rather than externally optimized search. These conditions are necessary for two distinct outcomes, separable as two necessity results. The capacity theorem states that linear scaling of substrate Shannon capacity with reaction extent requires R1, R2, and R3 but not R4: a perfect copier transmits an exponentially large configurational ensemble while producing no novelty. The generation theorem states that diversification of the heritable configuration set beyond the deterministic closure of a finite initial repertoire additionally requires R4, because branching trajectories in the recognized alphabet are what produce innovation. Populationlevel kinetics follow as a corollary that organizes six attested templating reactions into a taxonomy, and as a finite-population, finite-horizon proposition tested over five inheritance kinetic schemes, in which only individual-level stochastic drift reaches the target within the model class. We test the framework on the recently characterized bacterial defense system Drt3b, which makes alternating poly(AC) DNA without using a nucleic acid template. The framework classifies Drt3b as a cyclic two-state catalytic templating channel with a 1-bit capacity ceiling, and predicts that the Glu26-to-Gln active-site mutant incorporates dG at 10% probability at the dA-selecting state; the published biochemistry reports 10.16%. Across 1,232 Drt3b homologs, the framework predicts and recovers a 15.7-fold elevation of dG misincorporation in six clades carrying the natural Glu26-to-Asp substitution at this gate. Substitutions at two universal gate residues, Arg253 (architectural) and Gly248 (selectivity), provide single-experiment site-directed mutagenesis tests of the frameworks predictions. POPULAR SUMMARYA bacterial defense protein called Drt3b, recently characterized in E. coli, synthesizes DNA with a strict alternating ACAC pattern without copying any template. Two conserved active-site residues, Glu26 and Arg253, are modeled as enforcing an alternating two-state catalytic cycle that selects which nucleotide enters at each step. This is sequence without a sequence template, and it does not fit the textbook picture of inheritance. We treat inheritance as a class of catalytic templating reactions and ask which chemistries can support openended evolution. Two distinct requirements emerge. Capacity, the ability to transmit exponentially many distinct heritable configurations, requires three conditions on the template-catalyst pair: more than one monomer state recognized at each position, a position count that grows unboundedly with the reaction extent, and applicability of the catalysis to its own product. Generation of novel heritable configurations beyond what is already present requires a fourth condition: stochastic drift of the catalysis in its recognition alphabet. A perfect copier has capacity but cannot innovate; a drifting copier has both. Drt3b fails the second capacity condition, because its cycle has two states regardless of product length. The framework classifies six attested biological templating reactions as instances or partial instances of the same chemical specification, and it identifies two single-residue substitutions at Drt3bs active site whose measured effects would test its predictions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Boggavarapu, K.. 2026-06-29. A substrate recursion principle for biological information, with empirical anchoring through a templating-mode taxonomy. https://doi.org/10.64898/2026.06.23.734074

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Mechanism of molecular recognition revealed through dynamic drug binding pathways to SARS-CoV-2 main protease

Characterization of drug-binding pathways remains experimentally limited by transient intermediates and computationally challenging due to long timescales intractable for conventional molecular dynamics. To address these challenges, we combined solution NMR titrations with weighted ensemble (WE) enhanced sampling simulations to resolve atomistic pathways of nirmatrelvir binding to the SARS-CoV-2 main protease. NMR titration revealed residue-dependent heterogeneity spanning fast, intermediate, and slow exchange regimes. WE simulations complement the NMR by providing insights into unassigned residues and adding time-resolved and three-dimensional structural context. We map key interactions along two distinct binding pathways, provide dynamic explanations for residues involved in resistance, and capture unique backbone conformations compared to those sampled in unbound or bound states. Our comprehensive binding model is consistent with a combined conformational selection and induced fit mechanism in which early transient contacts are made with residues E47 and L50 and allosteric motions are centered around residue V204 of the distal domain. This synergistic application of WE and titration NMR enables a more comprehensive characterization of drug binding than either method alone, providing an integrated framework that may have broader applicability to defining structure-kinetic relationships and guiding design of next-generation inhibitors.

biophysics↗

Discriminating betacoronavirus receptor usage across subgenera using protein structure prediction and molecular dynamics

A critical step in the emergence of a virus is the ability of the viral protein to bind a host receptor and mediate cell entry. For many coronaviruses, this interaction occurs between the Spike S1 subunit and the human ACE2 receptor. Whether this binding interface can be computationally distinguished across unstudied viruses without experimentally resolved protein structures remains an open question. We predicted how 28 emerging coronaviruses may bind to human ACE2 using structural predictions, static interaction prediction programs, and molecular dynamics simulations. To screen the emerging coronaviruses, we predicted a library of S1 structures using AlphaFold. These predicted structures were then used to model the S1-ACE2 interaction with AlphaFold, ClusPro, and HADDOCK. We used known ACE2-binding sarbecoviruses as positive controls and coronaviruses that bind other receptors as negative controls to threshold predicted binding. Contact analysis quantified the predicted binding and revealed that these static interaction prediction methods varied in discriminative power. Less restrained static predictions separated binders from non-binders, whereas heavily restrained docking did not, potentially forcing an interaction where none should exist. This analysis highlighted an emerging coronavirus, Zhejiang2013, as a potential ACE2 binder. We used molecular dynamics simulations to further assess the static predictions and model the interaction over time. Overall, our results indicate that Zhejiang2013 exhibits dynamic interaction patterns consistent with ACE2 binding. Given that two ACE2-binding coronaviruses have caused global pandemics within the past two decades, identifying potential ACE2 binders is critical for early warning and pandemic preparedness.

biophysics↗

De novo design of flexible protein interactions with GuideFlip

De novo design of protein binders requires a target structure. However, for flexible targets, such as intrinsically disordered proteins, this structure does not exist until the binder has stabilized the interaction. Such targets are therefore difficult for methods that separate structure generation from sequence design. We introduce GuideFlip, which co-designs structure and sequence through guided discrete flow matching: binder residues are assigned progressively while the complex is re-predicted at each step, allowing the evolving interface to affect the design process. GuideFlip reduces the hydrophobic bias of direct AlphaFold optimization and improves in silico success rates over existing approaches. We release a database of binder candidates for 177 human disordered proteins. Experimentally, we obtain de novo binders to the C-terminus of -synuclein and the disordered amino terminus of RBX1 with hit rates of 13.5% and 41.7%, respectively, and we confirm the epitopes of selected binders by NMR and mutagenesis. Applying GuideFlip to flexibility on the binder side, we design a nanobody that binds the agonist-bound {beta}1-adrenergic receptor in the active state, but not the receptor in its inactive state, with a 75% hit rate and cryo-EM structure confirming the design. GuideFlip enables protein design where bound structures emerge only upon binding.

biophysics↗