bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.06.01.657285

Active learnings impact on student course performance in STEM varies by type and intensity

Abstract

We updated a recent meta-analysis of active learnings impact on student achievement in undergraduate STEM courses by following the same protocol to evaluate studies published from 2010-2017. We screened 1659 papers, coded 1294, and found 210 that met five pre-established inclusion criteria and six pre-established criteria for methodological quality. After further dropping 76 studies with no exam scores data, 134 of these studies contained data on student performance on identical or equivalent exams. We found that on average, active learnings effect size on exam scores was 0.519 {+/-} 0.049, meaning that when students are in active learning classes, they perform roughly half a standard deviation higher on an identical exam. Funnel plots and sensitivity analyses indicated that these results were not due to sampling bias. Active learning had a positive impact on student outcomes regardless of class size, course level, or STEM discipline, though there was heterogeneity in the effects. All of these results are very similar when compared to earlier meta-analyses, however increased resolution in the studies analyzed here revealed two novel results. First, student performance was significantly better in courses that employed high-intensity active learning, defined as students being on task at least two-thirds of class time, versus lower-intensities. Additionally, there was significant heterogeneity in efficacy across different types of active learning employed. These results suggest that most, if not all types of active learning are effective, and that when innovating in their classes, instructors should continually work to increase active learning intensity. We urge caution in interpreting the results on active learning types, however, and propose a preliminary framework for making more-sophisticated and reliable analyses of variation in course design. Finally, the evidence presented here for active learnings impact on student outcomes creates a strong foundation for faculty professional development and administration.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xu, S., Velasco, V., Hill, M. J., Tran, E., Agrawal, S., Arroyo, E. N., Behling, S., Chambwe, N., Cintron, D. L., Cooper, J. D., Dunster, G., Grummer, J. A., Hennessey, K., Hsiao, J., Iranon, N., Jones, L., Jordt, H., Keller, M., Lacey, M. E., Littlefield, C. E., Lowe, A., Newman, S., Okolo, V., Olroyd, S., Peecook, B. R., Pickett, S. B., Slager, D. L., Caviedes-Solis, I. W., Stanchak, K. E., Sundaravaradan, V., Valdebenito, C., Williams, C. R., Zinsli, K. A., Freeman, S., Theobald, E. J.. 2025-06-02. Active learnings impact on student course performance in STEM varies by type and intensity. https://doi.org/10.1101/2025.06.01.657285

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

How do selective serotonin reuptake inhibitors effect metabolism? A lesson for teaching non-majors about enzymes and metabolic pathways

Understanding foundational concepts in molecular biology and biochemistry can be challenging for many college students. For non-science majors, an introductory biology course may be their sole exposure to college-level science and only opportunity to learn about enzymes and metabolism. Framing these topics within issues that are personally relevant can make the material more meaningful and impactful for non-majors. There is a need to develop curricula that instructors of large-enrollment introductory biology courses can use to teach non-biology majors about enzymes and metabolism. Here, we designed and implemented a lesson that integrates enzyme function and metabolism within the context of selective serotonin reuptake inhibitors (SSRIs). SSRIs are a widely prescribed class of antidepressants whose efficacy can be influenced by enzymes. Genetic polymorphisms in the genes encoding these enzymes can result in individual differences in metabolic responses and lead to medication-induced weight gain. Framing the content in this way allows non-major students to connect enzyme function to genetics, energy balance, and metabolic processes through a health-related, societally relevant issue. In our lesson, students engaged in collaborative learning activities that emphasized interpreting biological data, graphing, and scientific reasoning. Pre- and post-assessment data from a conceptual inventory showed learning gains that exceeded baseline data and were greater than those observed with the traditional curriculum. Students also reported a more coherent understanding of enzyme function and a greater appreciation of the contents relevance to their own lives. This lesson provides an example of embedding biochemical concepts in real-world contexts for students who are non-science majors.

scientific communication and education↗

Insights into the Datasets, Tools, and Training Needs of the AnVIL Community: 2024

The NHGRI Genomic Data Science Analysis, Visualization, and Informatics Lab-space (AnVIL) provides a secure cloud-based environment where research and education communities can analyze genomic and biomedical data. The platform supports a wide range of data analysis as well as the ability to safely store and access data in compliance with NIH policies. Work on the AnVIL platform can be easily shared to promote reproducible science and collaboration. The purpose of this study is to better understand the current user base of the AnVIL platform. The AnVIL Community Poll aimed to collect baseline information, identify development opportunities, guide the prioritization of user support strategies, and succinctly but comprehensively describe the current AnVIL Community. The AnVIL Team disseminated the inaugural AnVIL Community Poll by sharing it broadly on social media and through AnVIL and related consortia mailing lists. We categorized respondents as either returning or potential users of the AnVIL platform (based on their provided usage description) and examined user experiences: specifically user backgrounds, technological comfort, research interests, computational needs, and preferences for training and support. Our sample of the AnVIL community found opportunities for platform adoption beyond the current user base and identified areas where training should be enhanced, training preferences, and user computational needs. Specifically, while most respondents were involved in human genomics research, there may be potential for growth in adoption of the platform by prioritizing materials to support clinical researchers. All respondents felt availability of specific tools or datasets was a key feature of the platform. The broader community may also benefit from further development or showcasing of resources to facilitate cost management, finding and incorporating analysis tools, and data import. Our sample greatly preferred virtual training opportunities and returning users of the platform foresaw needing large amounts of storage. This poll provided an insightful snapshot of the current state of the AnVIL and demonstrated areas where the AnVIL Team can take specific steps to address barriers related to platform adoption and further support the existing and varied AnVIL Community. This work can be built upon through user interviews, community discussion, and coordinating a recurring poll.

scientific communication and education↗