bioRxiv · 10.1101/2025.04.30.651464
Incorporating LLM-Derived Information into Hypothesis Testing for Genomics Applications
Abstract
We propose strategies for incorporating the information in large language models (LLMs) into statistical hypothesis tests in genomics studies. Using gene embeddings derived from text inputs to OpenAIs GPT-3.5 model, we show that biological signals in a variety of genomics datasets reside near the principal subspace spanned by the embeddings. We then use a frequentist and Bayesian (FAB) framework to propose several hypothesis tests that are either optimal or approximately optimal with respect to prior information based on the gene embedding subspace. In four real-world genomics examples, the FAB tests guided by the LLM-derived information achieve more power than classical counterparts.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bryan, J. G., Niu, H., Li, D.. 2025-05-07. Incorporating LLM-Derived Information into Hypothesis Testing for Genomics Applications. https://doi.org/10.1101/2025.04.30.651464
Cite the original work for its findings. Save a collection to share your selection of sources.