bioRxiv Science⌕ Search

Biology subjects

Waththe Liyanage, W. W.

Publications and source records attributed to Waththe Liyanage, W. W..

2 recordsLinked to original sources

Five hundred million years of methylation: tracing the mutational origins of vertebrate genome composition

Methylation-associated deamination removes CpG from vertebrate genomes, but how it affects the dinucleotide profile remains unresolved. We analysed 753 vertebrate and 481 invertebrate genomes to test whether CpG loss defines a compositional axis and to identify its strongest signature. CpG depletion was the dominant axis of vertebrate dinucleotide variation. Unexpectedly, its strongest between-genome correlate was AG/CT rather than the direct mutational product TpG/CpA, which showed the expected dataset-wide mass balance but varied little among genomes. A forward-evolution model based on measured seven-nucleotide human germline substitution rates produced neither CpG depletion nor AG/CT enrichment when methylated-CpG mutability was excluded. Adding one CpG-specific mutability term, calibrated only to the mammalian CpG ratio, reproduced both features, identifying AG/CT as a second-order consequence of the context-dependent mutation network. Within genomes, CpG depletion was strongest in transposable elements and weakened with distance from them. Across vertebrates, the axis followed Amniota more closely than endothermy and was associated with an expanded GC-rich isochore compartment. A Machine Learning analysis shows that CpG depletion and AG/CT were the principal features separating vertebrates from invertebrates, in which both were markedly attenuated. Thus, a methylation-associated axis organises vertebrate dinucleotide composition, and its strongest marker is not the immediate product of CpG deamination.

genomics↗

LLMsFold: Integrating Large Language Models and Biophysical Simulations for De Novo Drug Design

The discovery of novel small molecules is challenging because of the vastness of chemical space and the complexity of protein-ligand interactions, leading to low success rates and time-consuming workflows. Here, we present LLMsFold, a computational framework that combines Large Language Models (LLMs) and biophysical foundation tools to design and validate new small molecules targeting pathogenic proteins. The pipeline starts by identifying viable binding pockets on a target protein through geometry-based pocket detection. A 70-billion-parameter transformer model from the LlaMA family then generates candidate molecules as SMILES strings under prompt constraints that enforce drug-likeness. Each molecule is evaluated by Boltz-2, a diffusion-based model for protein-ligand co-folding that predicts bound 3D structure and binding affinity. Promising candidates are iteratively optimized through a reinforcement learning loop that prioritizes high predicted affinity and synthetic accessibility. We demonstrate the approach on two challenging targets: ACVR1 (Activin A Receptor Type 1), implicated in fibrodysplasia ossificans progressiva (FOP), and CD19, a surface antigen expressed on most B-cell lymphoma and leukemia cells. Top candidates show strong in silico binding predictions and favorable drug-like profiles. All code and models are made available to support reproducibility and further development.

bioinformatics↗