bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.03.08.531607

Infinite Physical Monkey: Do Deep Learning Methods Really Perform Better in Conformation Generation?

Abstract

Conformation Generation is a fundamental problem in drug discovery and cheminformatics. Generally, it can be categorized into three different classes according to physical scales, i.e., micro molecule (organic), meso molecule (nano particle-like), and macro molecule (protein and nucleic acid). Organic molecule conformation generation, particularly in vacuum and protein pocket environments, is most relevant to drug design. Recently, with the development of geometric neural networks, the data-driven schemes have been successfully applied in this field, both for molecular conformation generation (in vacuum) and binding pose generation (in protein pocket). The former beats the traditional ETKDG method, while the latter achieves similar accuracy compared with the widely used molecular docking software. Although these methods have shown promising results for real-world drug design campaigns, some researchers have recently questioned whether deep learning (DL) methods perform better in molecular conformation generation via a "parameter-free" method. To our surprise, what they have designed is some kind analogous to the famous infinite monkey theorem, the monkeys that are even equipped with physics education. To discuss the feasibility of their proving, we constructed a real infinite stochastic monkey for molecular conformation generation, showing that even with a more stochastic sampler for geometry generation, the coverage of the benchmark QM-computed conformations are higher than those of most DL-based methods. By extending their physical monkey algorithm for binding pose prediction (with 2000 random samples), we also discover that the successful docking rate also achieves near-best performance among existing DL-based docking models. Thus, though their conclusions are right, their proof process needs more concern. In addition to evaluating the rationality of their algorithms and conclusions, we dig into the inspirations of infinite physical monkeys. We find that, for docking pose generation, DL-based models truly learn the interaction rules between residues and ligands, and discover an inductive bias hidden in the training of the pocket-given docking problem. The code of the proposed algorithm could be found at: https://github.com/HaotianZhangAI4Science/infinite-physical-monkey. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC="FIGDIR/small/531607v2_ufig1.gif" ALT="Figure 1"> View larger version (126K): org.highwire.dtl.DTLVardef@188ac2eorg.highwire.dtl.DTLVardef@1e006f3org.highwire.dtl.DTLVardef@e83f8forg.highwire.dtl.DTLVardef@1a4c7cd_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOTOC:C_FLOATNO Infinite Physical Monkey. This image was created with the assistance of DALL{middle dot}E 2 C_FIG

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhang, H., Zhang, J., Zhao, H., Jiang, D., Deng, Y.. 2023-03-10. Infinite Physical Monkey: Do Deep Learning Methods Really Perform Better in Conformation Generation?. https://doi.org/10.1101/2023.03.08.531607

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Autonomous Homeostatic Synthetic Cells via Self-Gating DNA Nanopores

Homeostasis is a fundamental hallmark of living organisms, arising from the complex interplay between biochemical reactions and regulatory feedback systems. Reconstituting such self-regulating behaviour in minimal synthetic cells enables continuous, persistent operation of biochemical reactions for extended amount of time. In this work, we demonstrate a minimal homeostatic synthetic cell capable of autonomous flux regulation using DNA nanotechnology and bottom-up synthetic biology. Our homeostatic architecture consists of Giant Unilamellar Vesicles (GUVs) equipped with gated DNA nanopores, encapsulated in vitro transcription (IVT) machinery, and an RNA degradation system. We achieve homeostasis under varying external chemical stimuli specifically varying concentrations of rNTPs by implementing a negative feedback loop between rNTP influx and RNA production. In our system, DNA nanopores facilitate the influx of rNTPs from the external environment, driving internal transcription. Crucially, the transcription process generates RNA "blockers" designed to bind and gate the DNA nanopores, thereby attenuating further rNTP influx. Our system is dynamic as encapsulated RNases slowly degrade the RNA blockers, allowing the pores to reopen as blocker concentration goes down. We first characterise the functionality and gating efficiency of the DNA nanopores using both pre-synthesised and in situ produced DNA and RNA blockers. We then demonstrate that rNTP flux through these pores is sufficient to drive IVT within the GUVs. Finally, by integrating these modules, we demonstrate robust homeostasis: the system maintains a steady-state level of RNA production for up to 16 hours. By harnessing the controllability of negative feedback loop, we demonstrate thresholding of the homeostasis level using single-stranded regulator DNA. This work establishes a versatile framework for engineering adaptive and self-sustaining responsive nanomaterials and synthetic cell chassis.

biophysics↗

A Generic Numbering Scheme for TMEM16 Scramblases

The TMEM16 family of calcium-activated phospholipid scramblases (CaPLSs) and chloride channels (CaCCs) performs diverse physiological functions that include regulation of blood coagulation and apoptotic signaling, through a shared ten-transmembrane-helix (TM) architecture organized around a hydrophilic lipid-translocating groove. Mechanistic studies of TMEM16 family members have been hampered by the absence of a unified positional reference framework that would permit direct comparison of structurally equivalent residues across paralogs with different sequence numbering systems. Here we introduce a generic numbering scheme for TMEM16 scramblases (GNS-TMEM16), modeled on the Ballesteros & Weinstein system established for class A G protein-coupled receptors. A reference alignment (TMEM16-RA) was constructed from twelve human and mouse TMEM16 scramblases (TMEM16C/D/E/F/G/J) using structure-based ClustalW alignment of the ten TM helices. From this alignment, a TM-specific reference residue (TsRR) was identified for each helix by hierarchical application of three criteria: (1) 100% conservation in the core TMEM16-RA; (2) conservation in an augmented reference alignment (TMEM16-ARA) incorporating a group of phylogenetically more distant homologs composed of nhTMEM16, afTMEM16, TMEM16K, TMEM16A, and TMEM16B; and (3) structural and functional considerations, including helix-perturbing character, groove localization, conserved motif membership, and central TM position. The resulting ten TsRRs are Y1.50, W2.50, R3.50, E4.50, F5.50, P6.50, E7.50, D8.50, W9.50, and E10.50, and are illustrated in mTMEM16F. Each residue is assigned the identifier N.m(k), where N is the TM number, m is the position relative to the TsRR (for which m = 50), and k is the absolute sequence number. Loop residues receive dual identifiers referenced to the TsRRs of both flanking helices. Application of the GNS-TMEM16 is illustrated with the comparisons of the groove-opening measurements using pairwise distances between residues identified by their N.m indices to be corresponding across mTMEM16F, afTMEM16, and nhTMEM16. The results bring to light the advantages of corresponding residues identification in different TMEM16 proteins and show that the mammalian scramblase undergoes substantially larger separation at the extracellular groove entrance than either fungal homolog. Comparison of mutagenesis data guided by N.m correspondence shows at the conserved (E3.55,R6.26) salt-bridge locus, Ala substitution reduces activity more than 100-fold in nhTMEM16 but less than 2-fold in afTMEM16, illustrating that the GNS identifies structural equivalence of position without implying functional equivalence of the residue, which is a distinct advantage of GNS in providing mechanistic interpretation across paralogs. Also described is a protocol for extending the GNS-TMEM16 to uncharacterized protein sequences, including AlphaFold-predicted models, using structural superposition to mTMEM16F. Thus, the presented GNS-TMEM16 provides a stable positional reference for the integration and comparative analysis of structural, computational, and functional data across the TMEM16 family, utilizing a construction strategy applicable to yet other polytopic membrane protein families sharing a common transmembrane fold.

biophysics↗

An agent-based 3D model of non-genetic adaptation in cancer tissues under electrical, mechanical, and hypoxic stress

Non-genetic adaptation enables cancer cells to alter their phenotype under stress without requiring new mutations. However, the mechanisms by which electrical, mechanical, and hypoxic cues combine to shape this process in 3D tissues remain poorly understood. This work presents an agent-based tumor model that integrates vascular oxygen supply, a globally imposed electric field, mechanically mediated crowding and compression cues, phenotype transitions, cell growth, mitosis, death, and inheritance of adaptive memory across division. The simulated tumors exhibit a three-stage trajectory consisting of necrosis onset, transient collapse of live mass, and partial regrowth accompanied by progressive accumulation of adapted cells. Continuous electrical stimulation produces a dose-dependent reduction in live mass while markedly increasing the adapted fraction, with comparatively limited changes in final necrotic burden. This response is strongly conditioned by mechanics and reshapes (and is reshaped by) adaptive capacity. Pulsed stimulation further shows that, in the model, electric field amplitude and temporal schedule jointly determine memory phenomena, phenotypic diversification, and growth recovery. These results show that coupling local oxygen availability, mechanical constraints, electrical forcing, and history-dependent phenotype transitions can generate distinct tissue-level patterns of phenotypic heterogeneity. Both stimulus magnitude and temporal protocol influenced the resulting population structure, suggesting that the history of physical stress may be an important determinant of adaptive dynamics in spatially organized tumor models.

biophysics↗