bioRxiv Science⌕ Search

Biology subjects

Sinitskiy, A.

Publications and source records attributed to Sinitskiy, A..

10 recordsLinked to original sources

StructureSAFE: A structure-aware chemical language model for unified hit identification and lead optimization

Structure-based generative models (SBGMs) hold great promises for accelerating drug discovery by enabling target-aware molecular design. However, existing approaches face fundamental challenges: three-dimensional graph-based models can explicitly incorporate protein structural information but often generate chemically implausible molecules due to limited training data, while chemical language models (CLMs) produce chemically plausible molecules but struggle to effectively leverage three-dimensional structural information for structure-conditioned generation and hard to incorporate lead optimization functionality due to the nature of SMILES string. Here, we present StructureSAFE, a structure-aware chemical language model that resolves this trade-off by integrating protein structural and evolutionary encoders with the SAFE molecular representation via pretraining and finetuning training scheme, enabling both de novo hit identification and a comprehensive suite of lead optimization subtasks within a unified framework. Comprehensive benchmarking on the MolGenBench dataset demonstrates that StructureSAFE achieves state-of-the-art (SOTA) performance across multiple metrics, with particularly pronounced improvements in chemical plausibility relative to graph-based models lacking pretraining. Evaluation on a rigorously constructed held-out test set further confirms its ability to generate drug-like, synthetically accessible molecules with competitive predicted binding affinities for previously unseen targets on both hit identification and lead optimization setting. In silico case studies across four therapeutically relevant targets validate its capacity to generate chemically plausible molecules that recapitulate key binding interactions of known high-affinity ligands while proposing novel interactions for potential better affinity and exploring previously unknown regions of chemical space. Taking together, StructureSAFE represents a versatile and practical tool to provide high-quality candidate molecules for augmenting medicinal chemistry workflows in both hit identification and lead optimization campaigns.

bioinformatics↗

Practical Use of Advanced AI Frameworks on Real-Life Scientific Problems: Three Case Studies

Agentic artificial intelligence (AI) systems increasingly claim to automate scientific research, yet independent evaluations report persistent gaps between those claims and demonstrated capability. We tested frontier agentic AI systems on three practical problems: prediction of treatment non-response in immune-mediated inflammatory diseases, optical chemical structure recognition for literature mining, and prediction of drug-design-related properties from small datasets. Each problem was first assigned to autonomous frameworks and then reattempted as human-led, AI-assisted work. Autonomous runs failed in most cases, while human-led work produced reusable resources and modest but defensible performance, including new evidence for possible mechanisms of treatment resistance and a more practical benchmark for mining chemical structures from scientific papers. Property prediction was the single task on which one autonomous AI framework matched the human expert. We conclude that current frameworks can carry out engineering and analysis once a human expert leads the project, but cannot yet engineer a novel solution without oversight. The use of AI on real-life scientific problems remains an art rather than a routine technology.

bioinformatics↗

Real Science Is Harder Than Benchmarks: Evaluating Advanced AI Frameworks on Published Studies. I. Uncertainty Quantification, ML on Therapeutic Data Commons, and Agent-Based Modeling

Artificial Intelligence (AI) frameworks for automating scientific research have shown strong performance on benchmarks, but their capacity to routinely reproduce results from multiple real-life published studies remains largely untested. We evaluated five advanced AI research frameworks (Kosmos, K-Dense, ToolUniverse, BioAgents from bio.xyz, and the AI Scientist-v2 from Sakana AI) on three real-life tasks (including two recently published papers) spanning uncertainty quantification for molecular property predictions, machine learning on Therapeutic Data Commons benchmarks, and agent-based modeling. AI frameworks demonstrated genuine strengths: generating original hypotheses, competently executing routine data acquisition and coding tasks, providing statistical measures of confidence often absent from the original papers, and producing well-formatted final reports. At the same time, our experiments revealed that real-world scientific tasks remain considerably harder than current benchmarks suggest. No AI framework matched the scope or depth of the original studies, results varied across multiple runs of the same framework with the same prompt, and we documented cases of severe hallucinations in final reports, gaps in literature coverage, and overconfident conclusions. Verification of AI outputs required substantial domain expertise. While these three tasks are only partially representative of the broader scientific landscape, they offer a starting point for developing a more rigorous methodology for evaluation of AI performance than what is currently practiced. We conclude that AI frameworks are already valuable for prototyping research directions and stress-testing completed studies, and some of the limitations documented here appear largely tractable through infrastructure improvements and continued development.

bioinformatics↗

Can AI Conduct Autonomous Scientific Research? Case Studies on Two Real-World Tasks

Recent advances in artificial intelligence (AI) have prompted claims about autonomous "AI scientists," yet systematic evaluations of these capabilities remain scarce. This exploratory study investigates whether current AI frameworks can execute scientific research tasks beyond isolated demonstrations. We tested eight open-source AI frameworks (Agent Laboratory, AutoGen, BabyAGI, GPT Researcher, MOOSE-Chem2, SciAgents, SciMON, and Virtual Lab) on two tasks that aimed to reproduce research on algorithm development from recent papers in uncertainty quantification and protein interaction discovery. In our evaluation, no framework completed a full research cycle from literature understanding through computational execution to validated results and scientific paper writing. While all systems showed competence in conceptual tasks such as planning and summarization, they consistently failed at robust implementation. Every framework produced sophisticated hallucinations. Deployment proved demanding, requiring substantial debugging and technical expertise, which undermines common claims about the democratization of science with AI. Despite these limitations, the frameworks showed promise as research assistants for methodological planning and ideation under careful human supervision. Our findings suggest that the explored AI systems cannot yet autonomously conduct scientific research, but may provide real value for specific subtasks within the research workflow. We offer preliminary observations to help researchers and developers better understand the gap between advertised and actual capabilities of AI in science.

bioinformatics↗

SLOGEN: A Structure-based Lead Optimization Model Unifying Fragment Generation and Screening

Lead optimization plays an important role in preclinical drug discovery. While deep learning has accelerated this process, structure-based approaches that leverage 3D protein-ligand information remain underexplored. Existing models could improve predicted affinity but often yield synthetically inaccessible compounds, whereas screening-based methods limit chemical novelty by relying on fixed fragment libraries. To bridge the gap, we introduce Slogen--a Structure-based Lead Optimization algorithm unifying fragment Generation and screENing. To achieve this, Slogen integrates a transformer-based variational autoencoder, pretrained on the BindingNet v2 dataset, with an E(3)-equivariant graph neural network that models 3D protein-fragment interactions. This unified framework enables both fragment generation and similarity-based screening, simultaneously addressing synthetic tractability and structural diversity. Benchmarking study shows that Slogen matches or surpasses state-of-the-art methods while exploring broader chemical space. Case studies on the Smoothened and D1 dopamine receptors demonstrate its capacity to design high-affinity, drug-like molecules, providing a practical method for structure-guided lead optimization.

bioinformatics↗

Evolutionary Tree in Chemical Space of Natural Products

Natural products (NPs) are key to biological function and adaptation, with their distribution shaped by complex evolutionary and ecological forces. While it may seem reasonable to assume that closely related species produce chemically similar NPs, this assumption has not been systematically tested at a broad taxonomic scale. Here, we evaluate whether evolutionary (taxonomic) proximity correlates with chemical similarity in large-scale data from the Lotus database of NPs. We use five deep learning-based encoders, including Chemformer and SMILES Transformer, to embed NPs into a high-dimensional "chemical space." Our results demonstrate that, for flowering plants (Magnoliopsida) and conifers (Pinopsida), species separated by shorter taxonomic distances tend to produce significantly more similar NPs. Similar trends are observed for Fungi and Metazoa, albeit with some complications, possibly due to horizontal gene transfer, convergent evolution, and/or incomplete coverage in the dataset used for NPs. Our findings suggest that the evolutionary tree can be statistically recovered in a chemical space of NPs, provided that this space is constructed with appropriate deep learning techniques, and provide a new computational framework to investigate the evolutionary dynamics of secondary metabolism. These results can inform drug design strategies, for example by enabling the reconstruction of NPs from poorly studied or extinct species.

bioinformatics↗

Case Study of Using AI as Co-Pilot in Biotech Research: Functional Network Analysis of Invasive Cancer

This study presents a case analysis of using AI systems as co-pilots in biological research, focusing on functional protein networks in invasive colorectal cancer. We used public proteomic data alongside ChatGPT, GitHub Copilot, and PaperQA to automate parts of the workflow, including literature review, code generation, and network analysis. While AI tools improved efficiency, they required expert guidance for tasks involving complex metadata, domain-specific parsing, and reproducibility. Our analysis identified cytoskeleton- and signaling-related networks in invasive cancer, aligning with known biology, but attempts to distinguish invasive from non-invasive cases produced inconclusive results. An attempt to conduct fully automated research using Agent Laboratory failed due to hallucinated data, misinterpretation of research goals, and instability as the complexity of the underlying LLM increased. These findings show that current AI can assist but not replace human researchers in complex biotech studies.

bioinformatics↗

Computational model of primitive nervous system controlling chemotaxis in early multicellular heterotrophs

AO_SCPLOWBSTRACTC_SCPLOWThis paper presents a model to study a hypothetical role of a simple nervous systems in chemotaxis in early multicellular heterotrophs. The model views the organism as a network of motor units connected by flexible fibers and driven by realistic neuron excitation functions. Through numerical simulations, we identified the parameters that maximize the survival time of the modeled organism, focusing on its ability to efficiently locate and consume food. This synchronization enhances the ability of the modeled organism to navigate toward food and avoid harmful conditions. The model is described using basic mechanical principles and highlights the relationship between motor activity and energy balance. Our results suggest that even early prototypes of neural networks might provide significant survival advantages by optimizing movement and energy use. This study offers insights into how the first primitive nervous systems might have functioned. By publishing the code used in the simulations, we hope to contribute to the toolkit of computational methods and models used for exploration of neural origin and evolution.

biophysics↗

Simplest Model of Nervous System. IV. General Solution

In this paper, we extend our previous work on a simplified model of the nervous system by solving the general optimization problem for the evolutionary cost of the nervous system. This optimization takes into account constraints on the scales of membrane potential kinetics and sensory response function to ensure finite, biologically plausible solutions. Our analysis reduces the variational problem to a system of two integro-differential equations, which we solve asymptotically using series expansions. This study confirms the emergence of sharp finite neuronal spikes and robust sensory and motor response functions as evolutionarily optimal solutions. We note that, in principle, alternative evolutionary solutions with different biophysical interpretations might exist for this optimization problem. This work provides a rigorous mathematical framework bridging evolutionary optimization with the fundamental properties of nervous systems.

biophysics↗

Simplest Model of Nervous System. III. Partial Optimization

This paper extends our previous work on the simplest model of a nervous system by providing an asymptotic analysis of the evolutionarily optimal solution in this model. Building on the formalism and principles established earlier, we derive an asymptotic solution to the Fokker-Planck-Kolmogorov equation for a given dynamical equation for the state of the neuron. This solution provides the stationary probability distribution for the position of an organism in its environment and the state of its nervous system. Next, exact and asymptotic solutions for the optimal motor and sensory responses to the approach of a predator are derived. These results align with biological expectations and provide a robust mathematical framework for predicting the behavior of simple nervous systems.

biophysics↗