bioRxiv Science⌕ Search

Biology subjects

No, K. T.

Publications and source records attributed to No, K. T..

4 recordsLinked to original sources

Exploring the conformational space of protein-protein complex with transformer-based generative model

Protein-protein interactions are the basis of many protein functions, and understanding the contact and conformational changes of protein-protein interactions is crucial for linking protein structure to biological function. Although difficult to detect experimentally, molecular dynamics (MD) simulations are widely used to study the conformational ensembles and dynamics of protein-protein complexes, but there are significant limitations in sampling efficiency and computational costs. In this study, a generative neural network was trained on protein-protein complex conformations obtained from molecular simulations to directly generate novel conformations with physical realism. We demonstrated the use of a deep learning model based on the transformer architecture to explore the conformational ensembles of protein-protein complexes through MD simulations. The results showed that the learned latent space can be used to generate unsampled conformations of protein-protein complexes for obtaining new conformations complementing pre-existing ones, which can be used as an exploratory tool for the analysis and enhancement of molecular simulations of protein-protein complexes.

bioinformatics↗

Multimodal generation of astrocyte by integrating single-cell multi-omics data via deep learning

Obtaining positive and negative samples to examining several multifaceted brain diseases in clinical trials face significant challenges. We propose an innovative approach known as Adaptive Conditional Graph Diffusion Convolution (ACGDC) model. This model is tailored for the fusion of single cell multi-omics data and the creation of novel samples. ACGDC customizes a new array of edge relationship categories to merge single cell sequencing data and pertinent meta-information gleaned from annotations. Afterward, it employs network node properties and neighborhood topological connections to reconstruct the relationship between edges and their properties among nodes. Ultimately, it generates novel single-cell samples via inverse sampling within the framework of conditional diffusion model. To evaluate the credibility of the single cell samples generated through the new sampling approach, we conducted a comprehensive assessment. This assessment included comparisons between the generated samples and real samples across several criteria, including sample distribution space, enrichment analyses (GO term, KEGG term), clustering, and cell subtype classification, thereby allowing us to rigorously validate the quality and reliability of the single-cell samples produced by our novel sample method. The outcomes of our study demonstrated the effectiveness of the proposed method in seamlessly integrating single-cell multi-omics data and generating innovative samples that closely mirrored both the spatial distribution and bioinformatic significance observed in real samples. Thus, we suggest that the generation of these reliable control samples by ACGDC holds substantial promise in advancing precision research on brain diseases. Additionally, it offers a valuable tool for classifying and identifying astrocyte subtypes. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=92 SRC="FIGDIR/small/569500v1_ufig1.gif" ALT="Figure 1"> View larger version (18K): org.highwire.dtl.DTLVardef@1749dd8org.highwire.dtl.DTLVardef@1270912org.highwire.dtl.DTLVardef@1c4a65dorg.highwire.dtl.DTLVardef@18651b5_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

An interface-based molecular generative framework for protein-protein interaction inhibitors

Protein-protein interactions (PPIs) play a crucial role in numerous biochemical and biological processes. Although several structure-based molecular generative models have been developed, PPI interfaces and compounds targeting PPIs exhibit distinct physicochemical properties compared to traditional binding pockets and small-molecule drugs. As a result, generating compounds that effectively target PPIs, particularly by considering PPI complexes or interface hotspot residues, remains a significant challenge. In this work, we constructed a comprehensive dataset of PPI interfaces with active and inactive compound pairs. Based on this, we propose a novel molecular generative framework tailored to PPI interfaces, named GENiPPI. Our evaluation demonstrates that GENiPPI captures the implicit relationships between the PPI interfaces and the active molecules, and can generate novel compounds that target these interfaces. Moreover, GENiPPI can generate structurally diverse novel compounds with limited PPI interface modulators. To the best of our knowledge, this is the first exploration of a structure-based molecular generative model focused on PPI interfaces, which could facilitate the design of PPI modulators. The PPI interface-based molecular generative model enriches the existing landscape of structure-based (pocket/interface) molecular generative model.

bioinformatics↗

PhyloSophos: a high-throughput scientific name mapping algorithm augmented with explicit consideration of taxonomic science

SummaryThe nature of taxonomic science and the scientific nomenclature system makes it difficult to use scientific names as identifiers without running into complications. To facilitate high-throughput analysis of biological data involving scientific names, we designed PhyloSophos, a Python package that takes into account the properties of scientific names and taxonomic systems to map name inputs to the entries within the reference database of choice. We would like to present three case-studies which demonstrates how our implementations, including rule-based pre-processing and recursive mapping could improve mapping performance and information availability. We expect PhyloSophos to help with the systematic processing of poorly digitized and curated biological data, such as biodiversity information and ethnopharmacological resources, thus enabling full-scale bioinformatics analysis using these data. Availability and implementationPhyloSophos is available at GitHub https://github.com/mhcho4096/phylosophos. Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗