bioRxiv Science⌕ Search

Biology subjects

Leao, T. F.

Publications and source records attributed to Leao, T. F..

2 recordsLinked to original sources

A supervised fingerprint-based strategy to connect natural product mass spectrometry fragmentation data to their biosynthetic gene clusters

Microbial specialized metabolites are an important source of and inspiration for many pharmaceutical, biotechnological products and play key roles in ecological processes. However, most bioactivity-guided isolation and identification methods widely employed in metabolite discovery programs do not explore the full biosynthetic potential of an organism. Untargeted metabolomics using liquid chromatography coupled with tandem mass spectrometry is an efficient technique to access metabolites from fractions and even environmental crude extracts. Nevertheless, metabolomics is limited in predicting structures or bioactivities for cryptic metabolites. Linking the biosynthetic potential inferred from (meta)genomics to the specialized metabolome would accelerate drug discovery programs. Here, we present a k-nearest neighbor classifier to systematically connect mass spectrometry fragmentation spectra to their corresponding biosynthetic gene clusters (independent of their chemical compound class). Our pipeline offers an efficient method to link biosynthetic genes to known, analogous, or cryptic metabolites that they encode for, as detected via mass spectrometry from bacterial cultures or environmental microbiomes. Using paired data sets that include validated genes-mass spectral links from the Paired Omics Data Platform, we demonstrate this approach by automatically linking 18 previously known mass spectra to their corresponding previously experimentally validated biosynthetic genes (i.e., via NMR or genetic engineering). Finally, we demonstrated that this new approach is a substantial step towards making in silico (and even de novo) structure predictions for peptidic metabolites and a glycosylated terpene. Altogether, we conclude that NPOmix minimizes the need for culturing and facilitates specialized metabolite isolation and structure elucidation based on integrative omics mining. SignificanceThe pace of natural product discovery has remained relatively constant over the last two decades. At the same time, there is an urgent need to find new therapeutics to fight antibiotic-resistant bacteria, cancer, tropical parasites, pathogenic viruses, and other severe diseases. Here, we introduce a new machine learning algorithm that can efficiently connect metabolites to their biosynthetic genes. Our Natural Products Mixed Omics (NPOmix) tool provides access to genomic information for bioactivity, class, (partial) structure, and stereochemistry predictions to prioritize relevant metabolite products and facilitate their structural elucidation. Our approach can be applied to biosynthetic genes from bacteria (used in this study), fungi, algae, and plants where (meta)genomes are paired with corresponding mass fragmentation data.

bioinformatics↗

Heterologous expression of cryptomaldamide in a cyanobacterial host

Filamentous marine cyanobacteria make a variety of bioactive molecules that are produced by polyketide synthases, non-ribosomal peptide synthetases, and hybrid pathways that are encoded by large biosynthetic gene clusters. These cyanobacterial natural products represent potential drugs leads; however, thorough pharmacological investigations have been impeded by the limited quantity of compound that is typically available from the native organisms. Additionally, investigations of the biosynthetic gene clusters and enzymatic pathways have been difficult due to the inability to conduct genetic manipulations in the native producers. Here we report a set of genetic tools for the heterologous expression of biosynthetic gene clusters in the cyanobacteria Synechococcus elongatus PCC 7942 and Anabaena (Nostoc) PCC 7120. To facilitate the transfer of gene clusters in both strains, we engineered a strain of Anabaena that contains S. elongatus homologous sequences for chromosomal recombination at a neutral site and devised a CRISPR-based strategy to efficiently obtain segregated double recombinant clones of Anabaena. These genetic tools were used to express the large 28.7 kb cryptomaldamide biosynthetic gene cluster from the marine cyanobacterium Moorena (Moorea) producens JHB in both model strains. S. elongatus did not produce cryptomaldamide, however high-titer production of cryptomaldamide was obtained in Anabaena. The methods developed in this study will facilitate the heterologous expression of biosynthetic gene clusters isolated from marine cyanobacteria and complex metagenomic samples. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=110 SRC="FIGDIR/small/267179v1_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@660caforg.highwire.dtl.DTLVardef@1cab871org.highwire.dtl.DTLVardef@130de4org.highwire.dtl.DTLVardef@f50c64_HPS_FORMAT_FIGEXP M_FIG C_FIG

synthetic biology↗