bioRxiv · 10.64898/2026.09.01.748513
Turning Domain Expertise into Multi-Dimensional Evaluation of Biomedical AI with Karenina
Abstract
Language models and agents are increasingly used in biomedicine, but current benchmarks reward correct answers even when the underlying reasoning is flawed. Here we introduce Karenina, an open-source framework that turns expert knowledge into multi-dimensional evaluations of questions, conversations and autonomous agents. Illustrated in Question-Answer pairs, multi-turn conversations and autonomous data-analysis, these dimensions together moves evaluation beyond scoring, enabling trustworthy decision-making with AI in biomedicine.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Carli, F., Rusina, P., Ai, L., Kuechenhoff, L., To, P. K. P., McDonagh, E. M., Lobentanzer, S., Petroni, F., Dugourd, A., Ochoa, D., Saez-Rodriguez, J.. 2026-09-04. Turning Domain Expertise into Multi-Dimensional Evaluation of Biomedical AI with Karenina. https://doi.org/10.64898/2026.09.01.748513
Cite the original work for its findings. Save a collection to share your selection of sources.