bioRxiv Science⌕ Search

Biology subjects

Tagirdzhanov, A.

Publications and source records attributed to Tagirdzhanov, A..

5 recordsLinked to original sources

Nerpa 2: linking biosynthetic gene clusters tononribosomal peptide structures

MotivationNonribosomal peptides (NRPs) are bioactive microbial metabolites with high pharmaceutical potential. Although genome mining enables large-scale detection of biosynthetic gene clusters (BGCs) predicted to encode NRPs, reliably linking these clusters to their chemical products remains challenging due to the flexible and heterogeneous organization of NRP assembly pathways. ResultsWe present Nerpa 2, a probabilistic framework for accurate and scalable linking of NRP BGCs to candidate chemical structures. The method represents assembly lines as hidden Markov models (HMMs) that capture uncertainty and alternative biosynthetic routes. On curated datasets of experimentally validated BGC-product pairs, our tool outperforms existing methods in linking accuracy and pathway reconstruction. When applied to large genome mining datasets, Nerpa 2 efficiently identifies BGCs likely associated with known compounds and highlights potential producers of novel chemistry. Availability and implementationNerpa 2 is freely available at https://github.com/gurevichlab/nerpa.

bioinformatics↗

GRIP: physics-informed neural network for gradient retention time prediction in liquid chromatography

Gradient liquid chromatography has numerous applications in life sciences. Retention time prediction remains challenging due to complex underlying physical processes and highly variable chromatographic conditions. Here, we present GRIP, a physics-informed neural network for gradient retention time prediction that explicitly uses experimental setup parameters. GRIP demonstrates zero-shot generalization to unseen chromatographic systems while being on par or out-performing the transfer learning-based baseline. This approach can computationally guide the experimental setup configuration tailored to specific compounds of interest.

bioinformatics↗

BGC Atlas: A Web Resource for Exploring the Global Chemical Diversity Encoded in Bacterial Genomes

Secondary metabolites are compounds not essential for an organisms development, but provide significant ecological and physiological benefits. These compounds have applications in medicine, biotechnology, and agriculture. Their production is encoded in biosynthetic gene clusters (BGCs), groups of genes collectively directing their biosynthesis. The advent of metagenomics has allowed researchers to study BGCs directly from environmental samples, identifying numerous previously unknown BGCs encoding unprecedented chemistry. Here, we present the BGC Atlas (https://bgc-atlas.cs.uni-tuebingen.de), a web resource that facilitates the exploration and analysis of BGC diversity in metagenomes. The BGC Atlas identifies and clusters BGCs from publicly available datasets, offering a centralized database and a web interface for metadata-aware exploration of BGCs and gene cluster families (GCFs). We analyzed over 35,000 datasets from MGnify, identifying nearly 1.8 million BGCs, which were clustered into GCFs. The analysis showed that ribosomally synthesized and post-translationally modified peptides (RiPPs) are the most abundant compound class, with most GCFs exhibiting high environmental specificity. We believe that our tool will enable researchers to easily explore and analyze the BGC diversity in environmental samples, significantly enhancing our understanding of bacterial secondary metabolites, and promote the identification of ecological and evolutionary factors shaping the biosynthetic potential of microbial communities.

bioinformatics↗

ABC-HuMi: the Atlas of Biosynthetic Gene Clusters in the Human Microbiome

The human microbiome has emerged as a rich source of diverse and bioactive natural products, harboring immense potential for therapeutic applications. To facilitate systematic exploration and analysis of its biosynthetic landscape, we present ABC-HuMi: the Atlas of Biosynthetic Gene Clusters (BGCs) in the Human Microbiome. ABC-HuMi integrates data from major human microbiome sequence databases and provides an expansive repository of BGCs compared to the limited coverage offered by existing resources. Employing state-of-the-art BGC prediction and analysis tools, our database ensures accurate annotation and enhanced prediction capabilities. ABC-HuMi empowers researchers with advanced browsing, filtering, and search functionality, enabling efficient exploration of the resource. At present, ABC-HuMi boasts a catalog of 19,218 representative BGCs derived from the human gut, oral, skin, respiratory and urogenital systems. By capturing the intricate biosynthetic potential across diverse human body sites, our database fosters profound insights into the molecular repertoire encoded within the human microbiome and offers a comprehensive resource for the discovery and characterization of novel bioactive compounds. The database is freely accessible at https://www.ccb.uni-saarland.de/abc_humi/. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=149 SRC="FIGDIR/small/558305v2_ufig1.gif" ALT="Figure 1"> View larger version (52K): org.highwire.dtl.DTLVardef@15af4f6org.highwire.dtl.DTLVardef@886a27org.highwire.dtl.DTLVardef@1f12de5org.highwire.dtl.DTLVardef@fc24be_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

MolDiscovery: Learning Mass Spectrometry Fragmentation of Small Molecules

Identification of small molecules is a critical task in various areas of life science. Recent advances in mass spectrometry have enabled the collection of tandem mass spectra of small molecules from hundreds of thousands of environments. To identify which molecules are present in a sample, one can search mass spectra collected from the sample against millions of molecular structures in small molecule databases. This is a challenging task as currently it is not clear how small molecules are fragmented in mass spectrometry. The existing approaches use the domain knowledge from chemistry to predict fragmentation of molecules. However, these rule-based methods fail to explain many of the peaks in mass spectra of small molecules. Recently, spectral libraries with tens of thousands of labelled mass spectra of small molecules have emerged, paving the path for learning more accurate fragmentation models for mass spectral database search. We present molDiscovery, a mass spectral database search method that improves both efficiency and accuracy of small molecule identification by (i) utilizing an efficient algorithm to generate mass spectrometry fragmentations, and (ii) learning a probabilistic model to match small molecules with their mass spectra. We show our database search is an order of magnitude more efficient than the state-of-the-art methods, which enables searching against databases with millions of molecules. A search of over 8 million spectra from the Global Natural Product Social molecular networking infrastructure shows that our probabilistic model can correctly identify nearly six times more unique small molecules than previous methods. Moreover, by applying molDiscovery on microbial datasets with both mass spectral and genomics data we successfully discovered the novel biosynthetic gene clusters of three families of small molecules. AvailabilityThe command-line version of molDiscovery and its online web service through the GNPS infrastructure are available at https://github.com/mohimanilab/molDiscovery.

bioinformatics↗