bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.01.14.524073

Using a Topic Model to Map and Analyze a Large Curriculum

Abstract

A qualitative and quantitative understanding of curriculum content is critical for knowing whether its meeting its learning objectives. Curricula for medical education present challenges due to amount of content, the diversity of topics and the large number of contributing faculty. To create a manageable representation of the content in the pre-clerkship curriculum at Yale School of Medicine, a topic model was generated from all educational documents given to students during the pre-clerkship period. The model was used to quantitatively map content to school-wide competencies. The model measured how much of the curriculum addressed each topic and identified a new content area of interest, gender identity, whose coverage could be tracked over four years. The model also allowed quantitative measurement of integration of content within and between courses in the curriculum. The methods described here should be applicable to curricula in which texts can be extracted from materials.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Takizawa, P.. 2023-01-17. Using a Topic Model to Map and Analyze a Large Curriculum. https://doi.org/10.1101/2023.01.14.524073

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

A Large Language Model-Powered Map of Metabolomics Research

We present a comprehensive map of the metabolomics research landscape, synthesizing insights from over 80,000 publications. Using PubMedBERT, we transformed abstracts into 768-dimensional embeddings that capture the nuanced thematic structure of the field. Dimensionality reduction with t-SNE revealed distinct clusters corresponding to key domains such as analytical chemistry, plant biology, pharmacology, and clinical diagnostics. In addition, a neural topic modeling pipeline refined with GPT-4o mini reclassified the corpus into 20 distinct topics--ranging from "Plant Stress Response Mechanisms" and "NMR Spectroscopy Innovations" to "COVID-19 Metabolomic and Immune Responses." Temporal analyses further highlight trends including the rise of deep learning methods post-2015 and a continued focus on biomarker discovery. Integration of metadata such as publication statistics and sample sizes provide additional context to these evolving research dynamics. An interactive web application (https://metascape.streamlit.app/) enables dynamic exploration of these insights. Overall, this study offers a robust framework for literature synthesis that empowers researchers, clinicians, and policymakers to identify emerging research trajectories and address critical challenges in metabolomics, while also sharing our perspectives on key trends shaping the field.

scientific communication and education↗

Addressing cultural and knowledge barriers to enable preclinical sex inclusive research

For over thirty years, research has highlighted a sex bias in early research, risking the validity of biological knowledge. The first step towards change is effectively challenging misconceptions allowing researchers to perceive sex inclusive research as do-able. Utilising the theory of planned behaviour, we quantified researchers intention as a proxy measure for conducting sex inclusive research and explored attitude (value of the behaviour), subjective norm (perceived social pressure) and behavioural control (ability to conduct the behaviour). Additionally, we quantified the knowledge gap, prevalence of misconceptions, and assessed perceived benefits and barriers. We tested a workshop intervention that directly challenges the cultural embedded barriers. The data shows researchers intentions were high, but they had weak statistical knowledge and misunderstandings leading to a perception that inclusive research is prohibitive due to cost and animal use. We demonstrate that participation in the training intervention improved knowledge, altered the perceived barriers and cultural expectations.

scientific communication and education↗