bioRxiv · 10.1101/2025.05.19.654988
RL-Finetuning of OpenAI o1-mini to Enhance Biomedical Reasoning
Abstract
Recent breakthroughs in advanced reasoning large language models (LLMs), such as OpenAIs o1, have achieved impressive results in domains like math and coding. However, its not clear how much this type of reasoning helps in solving biomedical problems that involve more domain specialized knowledge and open-ended reasoning. Across two biomedical domains--gene characterization and small molecule property prediction--we find that the commercially available o1-mini model does not consistently outperform non-reasoning LLMs like GPT-4o. This motivated us to explore how much we can improve o1-minis biomedical reasoning through reinforcement learning (RL) finetuning. We show that RL finetuning of o1-mini results in large improvements in performance on gene classification, where it surprisingly outperformed domain-specific state-of-the-art models on some tasks. The results are mixed for small molecule prediction, suggesting that chemical reasoning could be more challenging for LLMs. We conclude with a discussion of the challenges and takeaways from this initial exploration of RL finetuning reasoning models for biomedical tasks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Swanson, K., Chen, Y. T., Jaech, A., Zou, J.. 2025-05-24. RL-Finetuning of OpenAI o1-mini to Enhance Biomedical Reasoning. https://doi.org/10.1101/2025.05.19.654988
Cite the original work for its findings. Save a collection to share your selection of sources.