bioRxiv Science⌕ Search

Biology subjects

Busta, L.

Publications and source records attributed to Busta, L..

10 recordsLinked to original sources

Small language models enable rapid and accurate extraction of structured data from unstructured text: an example with plants and their specialized metabolites

Transformer-based large language models are receiving considerable attention because of their ability to analyze scientific literature. Small language models (SLMs), however, also have potential in this area, have smaller compute footprints, and allow users to keep data in-house. Here, we quantitatively evaluate the ability of SLMs to: (i) score references according to project-specific relevance and (ii) extract and structuring data from unstructured sources (scientific abstracts). By comparing SLMs outputs against those of a human on hundreds of abstracts, we found that (i) SLMs can effectively filter literature and extract structured information relatively accurately (error rates as low as 10%), but not with perfect yield (as low as 50% in some cases), (ii) that there are tradeoffs between accuracy, model size, and computing requirements, and (iii) that clearly written abstracts are needed to support accurate data extraction. We recommend advanced prompt engineering techniques, full-text resources, and model distillation as future directions.

plant biology↗

Rule-Based Deconstruction and Reconstruction of Diterpene Libraries: Categorizing Patterns & Unravelling the Structural Landscape

Terpenoids make up the largest class of specialized metabolites with over 180,000 reported compounds currently across all kingdoms of life. Their synthesis accentuates one of natures most choreographed enzymatic and non-reversible chemistries, leading to an extensive range of structural functionality and diversity. Current terpenoid repositories provide a seemingly endless landscape to systematically survey for information regarding structure, sourcing, and synthesis. Efforts here investigate entries for the 20-carbon diterpenoid variants and deconstruct the complex patterns into simple, categorical groups. This deconstruction approach reduces over 60,000 unique diterpenoid structures to less than 1,000 categorical structures. Furthermore, the majority of diterpene entries (over 75%) can be represented by less than 25 core skeletons. Natural diterpenoid abundance was mapped throughout the tree of life and structural diversity was correlated at an atom-and-bond resolution. Additionally, all identified core structures provide guidelines for predicting how diterpene diversity originates via the mechanisms catalyzed by diterpene synthases. Over 95% of diterpenoid structures rely on cyclization. Here a reconstructive approach is reapplied based on known biochemical rules to model the birth of compound diversity. Reconstruction enabled prediction of highly probable synthesis mechanisms for bioactive taxane-relatives, which were discovered over three decades ago. This computational synthesis validates previously identified reaction products and pathways, as well as enables predicting trajectories for synthesizing real and theoretical compounds. This deconstructive and reconstructive approach applied to the diterpene landscape provides modular, flexible, and an easy-to-use toolset for categorically simplifying otherwise complex or hidden patterns. Significance StatementWe take a deconstructive and reconstructive approach to explore the origins of the diterpene landscape. Introduction of a navigational toolset enables users to survey compound libraries in ways formerly uncharted. Their utility demonstrated here, maps out diterpene cyclization routes, critical intermediate waypoints, and guidance for how to arrive at compounds previously off-the-map. Information acquired from these tools may imply the diterpene landscape is vastly unexplored, with the plateau for discovery potentially still out of sight.

bioinformatics↗

Advancing Plant Metabolic Research By Using Large Language Models To Expand Databases And Extract Labelled Data

Premise: Recently, plant science has seen transformative advances in scalable data collection for sequence and chemical data. These large datasets, combined with machine learning, revealed that conducting plant metabolic research on large scales yields remarkable insights. A key next step in increasing scale has been revealed with the advent of accessible large language models, which, even in their early stages, can distill structured data from literature. This brings us closer to creating specialized databases that consolidate virtually all published knowledge on a topic. Methods: Here, we first test different prompt engineering technique / language model combinations in the identification of validated enzyme-product pairs. Next, we evaluate automated prompt engineering and retrieval augmented generation applied to identifying compound-species associations. Finally, we build and determine the accuracy of a multimodal language model-based pipeline that transcribes images of tables into machine-readable formats. Results: When tuned for each specific task, these methods perform with high accuracies (80-90 percent for enzyme-product pair identification and table image transcription), or with modest accuracies (50 percent) but lower false-negative rates than previous methods (down to 40 percent from 55 percent) for compound-species pair identification. Discussion: We enumerate several suggestions for working with language models as researchers, among which is the importance of the users domain-specific expertise and knowledge. Significance StatementScientific databases have played a major role in advancing metabolic research. However, even todays advanced databases are incomplete and/or are not built to best suit certain research tasks. Here, we explored and evaluated the use of large language models and various prompt engineering techniques to expand and subset existing databases in task-specific ways. Our results illustrate the potential for high-accuracy additions and restructurings of existing databases using language models, assuming the specific methods by which the models are used are tuned and validated for the specific task. These findings are important because they outline a method by which we could greatly expand existing databases and rapidly tailor them to specific research efforts, leading to greater research productivity and effective utilization of past research findings. All authors collected data, analyzed data, prepared the manuscript, and approved its final version. The authors declare that they have no competing interests.

plant biology↗

Wax bloom dynamics on Sorghum bicolor under different environmental stresses reveal signaling modules associated with wax production

Epicuticular wax blooms are associated with improved drought resistance in many species, including Sorghum bicolor. While the role of wax in drought resistance is well known, we report new insights into how light and drought dynamically influence wax production. We investigated how wax quantity and composition are modulated over time and in response to different environmental stressors, as well as the molecular and genetic mechanisms involved in such. Gas chromatography-mass spectrometry and photographic results showed that sorghum leaf sheath wax load and composition were altered in mature plants grown under drought and simulated shade, though this phenomenon appears to vary by sorghum cultivar. We combined an in vitro wax induction protocol with GC-MS and RNA-seq measurements to identify a draft signaling pathway for wax bloom induction in sorghum. We also explored the potential of spectrophotometry to aid in monitoring wax bloom dynamics. Spec-trophotometric analysis showed primary differences in reflectance between bloom-rich and bloomless tissue surfaces in the 230-500nm range of the spectrum, corresponding to the blue color channel of photographic data. Our smartphone-based system detected significant differences in wax production between control and shade treatment groups, demonstrating its potential for candidate screening. Overall, our data suggest that wax extrusion can be rapidly modulated in response to light, occurring within days compared to the months required for the changes observed under greenhouse drought/simulated shade conditions. These results highlight the dynamic nature of wax modulation in response to varying environmental stimuli, especially light and water availability. Significance StatementAgricultural crops require significant freshwater for irrigation, making food security vulnerable to drought. Epicuticular wax blooms are associated with drought tolerance in many plants, including Sorghum bicolor. We investigated how environmental factors like light and drought influence wax production in sorghum. Wax production, composition, and gene expression were compared between sorghum exposed to different environmental stressors, reavealing dynamic modulation of wax production in response to environmental stress as well as signaling genes potentially involved in regulating wax production. These findings broaden our understanding of wax-related drought tolerance mechanisms, providing a foundation for future efforts to enginner crops with improved climate resilience.

plant biology↗

A molecular representation system with a common reference frame for natural products pathway discovery and structural diversity tasks.

Researchers have uncovered hundreds of thousands of natural products, many of which contribute to medicine, materials, and agriculture. However, missing knowledge of the biosynthetic pathways to these products hinders their expanded use. Nucleotide sequencing is key in pathway elucidation efforts, and analyses of natural products molecular structures, though seldom discussed explicitly, also play an important role by suggesting hypothetical pathways for testing. Structural analyses are also important in drug discovery, where many molecular representation systems - methods of representing molecular structures in a computer-friendly format - have been developed. Unfortunately, pathway elucidation investigations seldom use these representation systems. This gap is likely because those systems are primarily built to document molecular connectivity and topology, rather than the absolute positions of bonds and atoms in a common reference frame, the latter of which enables chemical structures to be connected with potential underlying biosynthetic steps. Here, we present a unique molecular representation system built around a common reference frame. We tested this system using triterpenoid structures as a case study and explored the systems applications in biosynthesis and structural diversity tasks. The common reference frame system can identify structural regions of high or low variability on the scale of atoms and bonds and enable hierarchical clustering that is closely connected to underlying biosynthesis. Combined with phylogenetic distribution information, the system illuminates distinct sources of structural variability, such as different enzyme families operating in the same pathway. These characteristics outline the potential of common reference frame molecular representation systems to support large-scale pathway elucidation efforts. Significance StatementStudying natural products and their biosynthetic pathways aids in identifying, characterizing, and developing new therapeutics, materials, and biotechnologies. Analyzing chemical structures is key to understanding biosynthesis and such analyses enhance pathway elucidation efforts, but few molecular representation systems have been designed with biosynthesis in mind. This study developed a new molecular representation system using a common reference frame, identifying corresponding atoms and bonds across many chemical structures. This system revealed hotspots and dimensions of variation in chemical structures, distinct overall structural groups, and parallels between molecules structural features and underlying biosynthesis. More widespread use of common reference frame molecular representation systems could hasten pathway elucidation efforts.

biochemistry↗

A new spin on chemotaxonomy using non-proteogenic amino acids as a test case

PremiseSpecialized metabolites serve various roles for plants and humans. Unlike core metabolites, specialized metabolites are restricted to certain lineages. Thus, in addition to their ecological functions, specialized metabolites can serve as diagnostic markers of plant lineages. MethodsWe investigate the phylogenetic distribution of plant metabolites using non-proteogenic amino acids (NPAA). Species-NPAA associations for eight NPAAs were identified from the existing literature and placed within a phylogenetic context using R packages and interactive tree of life. To confirm and extend the literature-based NPAA distribution we selected azetidine-2-carboxylic acid (Aze) and screened over 70 diverse plants using GC-MS. ResultsLiterature searches identified > 900 NPAA-relevant articles, which were manually inspected to identify 560 species-NPAA associations. NPAAs were mapped at the order and genus level, revealing that some NPAAs are restricted to single orders, whereas others are present across divergent taxa. The distribution of Aze across plants suggests a convergent evolutionary history. DiscussionThe reliance on chemotaxonomy has decreased over the years. Yet, there is still value in placing metabolites within a phylogenetic context to understand the evolutionary processes of plant chemical diversification. This approach can be applied to metabolites present in any organism and compared at a range of taxonomic levels.

plant biology↗

Phylochemical mapping of natural products onto the plant tree of life using text mining and large language models.

Plants produce a staggering array of chemicals that are the basis for organismal function and diversity and also provide essential human nutrients and medicine. However, it is poorly defined how these compounds have evolved and are distributed across the diverse lineages of the plant kingdom, hindering a systematic view and understanding of plant chemical diversity. Recent advances in plant genome/transcriptome sequencing have provided a well-defined molecular phylogeny of plants, on which the presence of diverse natural products can be mapped to systematically determine their phylogenetic distribution. Here, we built a proof-of-concept workflow via which previously reported diverse tyrosine-derived plant natural products were mapped on to the plant tree of life. Plant chemical-species associations were mined from literature, filtered, evaluated through manual inspection of over 2,500 scientific articles, and mapped onto the plant phylogeny. The resulting phylochemical map confirmed several highly lineage-specific compound class distributions, such as betalain pigments and Amaryllidaceae alkaloids. The map also highlighted several lineages enriched in dopamine-derived compounds, including the orders Caryophyllales, Liliales, and Fabales. Additionally, the application of large language models using our manually curated data as a ground truth set showed that post-mining manual processing steps can largely be automated with a low false positive rate. Our study demonstrates that a workflow combining text mining with language model-based processing can generate broader phylochemical maps, which will serve as a critical community resource to uncover key evolutionary events that underlie plant chemical diversity and enable system-level views of natures millions of years of chemical experimentation. Significance StatementThis study introduces a novel workflow to map the distribution of natural products onto a phylogeny. It also demonstrates the effectiveness of combining text mining of the literature with large language model evaluation for constructing a phylochemical database that can organize vast portions of the chemical diversity present across the plant tree of life.

plant biology↗

Project ChemicalBlooms: Collaborating with Citizen Scientists to Survey the Chemical Diversity and Phylogenetic Distribution of Plant Epicuticular Wax Blooms

Plants use chemistry to overcome diverse challenges. A particularly striking chemical trait that some plants possess is the ability to synthesize massive amounts of epicuticular wax that accumulates on the plants surfaces as a white coating visible to the naked eye. The ability to synthesize basic wax molecules appears to be shared among virtually all land plants and our knowledge of ubiquitous wax compound synthesis is reasonably advanced. However, the ability to synthe-size thick layers of visible epicuticular crystals ("wax blooms") is restricted to specific lineages and our knowledge of how wax blooms differ from ubiquitous wax layers is less developed. Here, we recruited the help of citizen scientists and middle school students to survey the wax bloom chemistry of 78 species spanning dicot, monocot, and gymnosperm lineages. Using gas chromatography-mass spectrometry, we found that the major wax classes reported from bulk wax mixtures can be present in wax bloom crystals, with fatty acids, fatty alcohols, and alkanes being present in many species bloom crystals. In contrast, other compounds including aldehydes, ketones, secondary alcohols, and triterpenoids were present in only a few species wax bloom crystals. By mapping the 78 wax bloom chemical profiles onto a phylogeny and using phylogenetic comparative analyses, we found that secondary alcohol and triterpenoid-rich wax blooms were present in lineage-specific patterns that would not be expected to arise by chance. That finding is consistent with reports that secondary alcohol biosynthesis enzymes are found only in certain lineages, but was a surprise for triterpenoids, which are intracellular components in virtually all plant lineages. Thus, our data suggest that a lineage-specific mechanism other than biosynthesis exists that enables select species to generate triterpenoid-rich surface wax crystals. Overall, our study outlines a general mode in which research scientists can collaborate with citizen scientists as well as middle and high school classrooms not only to enhance data collection and generate testable hypotheses, but also directly involve classrooms in the scientific process and inspire future STEM workers. Significance StatementPlants coat themselves in a protective layers of waxes. Some plants produce exceptionally thick layers of wax ("wax blooms"), and these thick layers are associated with enhanced abilities to tolerate stress, including drought and insect attack. In collaboration with citizen scientists and middle school classrooms, we provide an overview of the chemistry and phylogenetic distribution of plant epicuticular wax blooms. These data constitute a foundation upon which future studies of diverse wax blooms and their functions can build.

biochemistry↗

Traditional medicinal use is linked with apparency, not specialized metabolite profiles in the order Caryophyllales

A better understanding of the relationship between plant specialized metabolism and traditional medicinal use has the potential to aid in bioprospecting and the untangling of cross-cultural plant use patterns. However, given the limited information available for metabolites in most plant species, associating medicinal properties with a metabolite can be difficult. The order Caryophyllales has a unique pattern of lineages of tyrosine- or phenylalanine-dominant specialized metabolism, represented by mutually exclusive anthocyanin and betalain pigments, making the group ideal to work around a lack of detailed knowledge of specific metabolites. We compiled a list of medicinal species in selected tyrosine- or phenylalanine-dominant families of Caryophyllales across the globe (Nepenthaceae, Polygonaceae, Simmondsiaceae, Microteaceae, Caryophyllaceae, Amaranthaceae, Limeaceae, Molluginaceae, Portulacaceae, Cactaceae, and Nyctaginaceae) by searching scientific literature until no new uses were recovered, and tested for phylogenetic clustering of medicinal uses using a "hot nodes" approach. To test potential non-metabolite drivers of medicinal use, like how often humans encounter a species (apparency), we repeated the same analysis in North American species across the entire order and performed phylogenetic generalized least squares regression (PGLS) with occurrence data from the Global Biodiversity Information Facility (GBIF). We hypothesized families with Tyr-enriched metabolism would show clustering of different types of medicinal use compared to the ancestral Phe-enriched metabolism. Instead, weedy, wide-ranging clades in Polygonaceae and Amaranthaceae are overrepresented across nearly all types of medicinal use. Therefore, we found that apparency is a better predictor of medicinal use than metabolite profiles, although metabolism type may still be a contributing factor.

plant biology↗

Herbarium specimens as tools for exploring the evolution of biosynthetic pathways to fatty acid-derived natural products in plants

Plants synthesize natural products via lineage-specific offshoots of their core metabolic pathways, including fatty acid synthesis. Recent studies have shed light on new fatty acid-derived natural products and their biosynthetic pathways in disparate plant species. Inspired by this progress, we set out to expand the tools available for exploring the evolution of biosynthetic pathways to fatty-acid derived products. We sampled representative species from all major clades of euphyllophytes, including ferns, gymnosperms, and angiosperms (monocots and eudicots), and we show that quantitative profiles of fatty-acid derived surface waxes from preserved plant specimens are consistent with those obtained from freshly collected tissue. We then sampled herbarium specimens representing >50 monocot species to assess the phylogenetic distribution and infer the evolutionary origins of two fatty acid-derived natural products found in that clade: beta-diketones and alkyl resorcinols. These chemical data, combined with analyses of 26 monocot genomes, suggest whole genome duplication as a likely mechanism by which both diketone and alkylresorcinol synthesis evolved from an ancestral alkylresorcinol synthase-like polyketide synthase. This work reinforces the widespread utility of herbarium specimens for studying leaf surface waxes (and possibly other chemical classes) and reveals the evolutionary origins of fatty acid-derived natural products within monocots. Significance StatementPlant chemicals are key components in our food and medicine, and advances in genomic technologies are accelerating plant chemical research. However, access to tissue from specific plant species can still be rate-limiting, especially for species that are difficult to cultivate, endangered, or inaccessible. Here, we demonstrate that herbarium specimens provide a semiquantitative proxy for the cuticular wax profiles of their fresh counterparts, thus reducing the need to collect fresh tissue for studies of wax chemicals and suggesting the same may also be true of other plant chemical classes. We also demonstrate the utility of combining herbarium-based plant chemical profiling with genomic analyses to understand the evolution of plant natural products.

plant biology↗