bioRxiv Science⌕ Search

Biology subjects

Cordova, R. A.

Publications and source records attributed to Cordova, R. A..

2 recordsLinked to original sources

Mapping the mammalian dark metabolome by in vivo isotope tracing

Despite decades of biochemical study, a comprehensive map of the mammalian metabolome remains elusive. Mass spectrometry-based metabolomics detects thousands of small molecule-associated signals in mammalian tissues, but it is currently unclear how many of these reflect products of endogenous metabolism. Here, we leverage systematic in vivo isotope tracing to infer the biosynthetic origins of unidentified metabolites. We administered 26 different isotopically labelled nutrients to mice, measured circulating and tissue metabolite labelling by mass spectrometry, and developed a statistical framework to infer the number of carbon atoms incorporated from each of these precursors into more than 4,000 putative metabolites. We show this information can be harnessed for biosynthesis-aware structure elucidation using a multimodal AI model that co-embeds isotopic labelling patterns with chemical structures. This approach revealed several previously unrecognized families of mammalian metabolites, including cysteine-derived alkylthiazolidines, dithioacetal mercapturic acid derivatives, short-chain N-acyltaurines, acylglycyltaurines, and N-oxidized taurines. It further uncovered a family of mevalonate-derived isoprenoid metabolites that includes 2,3-dihydrofarnesoic acid, which is markedly depleted in both mouse and human aging. Age-related depletion of these isoprenoids is driven by impaired coenzyme A synthesis. Our work establishes the biosynthetic precursors for thousands of unidentified metabolites and reveals multiple previously unrecognized branches of mammalian metabolism.

biochemistry↗

Language model-guided anticipation and discovery of unknown metabolites

Despite decades of study, large parts of the mammalian metabolome remain unexplored. Mass spectrometry-based metabolomics routinely detects thousands of small molecule-associated peaks within human tissues and biofluids, but typically only a small fraction of these can be identified, and structure elucidation of novel metabolites remains a low-throughput endeavor. Biochemical large language models have transformed the interpretation of DNA, RNA, and protein sequences, but have not yet had a comparable impact on understanding small molecule metabolism. Here, we present an approach that leverages chemical language models to discover previously uncharacterized metabolites. We introduce DeepMet, a chemical language model that learns the latent biosynthetic logic embedded within the structures of known metabolites and exploits this understanding to anticipate the existence of as-of-yet undiscovered metabolites. Prospective chemical synthesis of metabolites predicted to exist by DeepMet directs their targeted discovery. Integrating DeepMet with tandem mass spectrometry (MS/MS) data enables automated metabolite discovery within complex tissues. We harness DeepMet to discover several dozen structurally diverse mammalian metabolites. Our work demonstrates the potential for language models to accelerate the mapping of the metabolome.

bioinformatics↗