bioRxiv · 10.1101/2020.02.28.970434
Modeling cannabinoids from a large-scale sample of Cannabis sativa chemotypes
Abstract
The accelerating legalization of Cannabis has opened the industry to using contemporary analytical techniques. The gene regulation and pharmacokinetics of dozens of cannabinoids remain poorly understood. Because retailers in many medical and recreational jurisdictions are required to report chemical concentrations of cannabinoids, commercial laboratories have growing chemotype datasets of diverse Cannabis cultivars. Using a data set of 17,600 cultivars tested by Steep Hill Inc., we apply machine learning techniques to interpolate missing chemotype observations and cluster cultivars together based on similarity. Our results show that cultivars cluster based on their chemotype, and that some imputation methods work better than others at grouping these cultivars based on chemotypic identity. However, due to the missing data for some of the cannabinoids their behavior could not be accurately predicted. These findings have implications for characterizing complex interactions in cannabinoid biosynthesis and improving phenotypical classification of Cannabis cultivars.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Vergara, D., Gaudino, R., Blank, T., Keegan, B.. 2020-02-28. Modeling cannabinoids from a large-scale sample of Cannabis sativa chemotypes. https://doi.org/10.1101/2020.02.28.970434
Cite the original work for its findings. Save a collection to share your selection of sources.