bioRxiv ScienceSearch

bioRxiv · 10.1101/510081

Post exam analysis: Implication for intervention

Abstract

Post exam item analysis enables teachers to reduce biases on student achievement assessments and improve their way of instruction. Difficulty indices, discrimination power and distracter efficiencies were commonly investigated in item analysis. This research was intended to investigate the difficulty and discrimination indices, distracters efficiency, whole test reliability and construct defects in summative test for freshman common course at Gondar CTE. In this study, 176 exam papers were analyzed in terms of difficulty index, point bi-serial correlation and distracter efficiencies. Internal consistency reliability and construct defects such as meaningless stems, punctuation errors and inconsistencies in option formats were also investigated. Results revealed that the summative test as a whole has moderate difficulty level (0.56 {+/-} 0.20) and good distracter efficiency (85.71% {+/-} 29%). However, the exam was poor in terms of discrimination power (0.16 {+/-} 0.28) and internal consistency reliability (KR-20 = 0.58). Only one item has good discrimination power and one more item excellent in its discrimination. About 41.9% of the items were either too easy or too difficult. Inconsistency in option formats or inappropriate options, punctuation errors and meaningless stems were also observed. Thus, future test development interventions should give due emphasis on item reliability, discrimination coefficient and item construct defects.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tsegaye, K. N.. 2019-01-04. Post exam analysis: Implication for intervention. https://doi.org/10.1101/510081

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

Strategies for building computing skills to support microbiome analysis: a five-year perspective from the EDAMAME workshop

Here, we report our educational approach and learner evaluations of the first five years of the Explorations in Data Analysis for Metagenomic Advances in Microbial Ecology (EDAMAME) workshop, held annually at Michigan State Universitys Kellogg Biological Station from 2014-2018. We hope this information will be useful for others who want to organize computing-intensive workshops and encourage quantitative skill development among microbiologists.\n\nImportanceHigh-throughput sequencing and related statistical and bioinformatic analyses have become routine in microbiology in the past decade, but there are few formal training opportunities to develop these skills. A week-long workshop can offer sufficient time for novices to become introduced to best computing practices and common workflows in sequence analysis. We report our experiences in executing such a workshop targeted to professional learners (graduate students, post-doctoral scientists, faculty, and research staff).

scientific communication and education

The Role of the Public Health Service Equalization Program in the Control of Hypertension in China: Results from a Cross-sectional Health Interview Survey

ObjectivesNon-communicable diseases (NCDs) have become the main cause of mortality in China. In 2009, the Chinese government introduced the Public Health Service Equalization (PHSE) program to restore the primary healthcare system in both essential medical care and public health service provision. This study evaluates the impact of management on hypertension control and evaluate how the program works.\n\nMethodsThe China National Health Development Research Centre (CNHDRC) undertook the Cross-sectional Health Service Interview Survey (CHSIS) of 62,097 people from primary healthcare reform pilot areas, across 17 provinces from eastern, central and western parts of China in 2014. This study is based on CHSIS survey responses from 9,607 participants, who had been diagnosed with hypertension. Regression analysis was used to estimate the impact of management provided under PHSE on hypertension control adjusting for the effects of other known determinants of hypertension control.\n\nFindingsUncontrolled hypertension was markedly lower among respondents, whose hypertension had been managed (22.4% in managed patients versus 31.1% in unmanaged patients, p<0.001). The interaction between PHSE management and the geographical region was highly significant in the model (p<0.001), suggesting that the PHSE program was not equally effective in all regions. Further analysis suggested that approximately 10% of regional variability was attributed to differences in administrative systems, as there was a significant association (P=0.014) between the presence of established regional Information Management Systems (IMS) and increased PHSE effectiveness. Insurance ({chi}2(5)=4.4, p=0.496) and Hukou ({chi}2(1)=2.4, p=0.121), which denote social security and urban rural differences, respectively, were not significant predictor of hypertension control.\n\nConclusionActive management of hypertension through the PHSE program was effective with 7.31 million more patients receiving hypertension control and equalization of service delivery was reflected to some extent. The link between established IMS and regional variability in the impact of PHSE highlights the importance of effective management of patient referrals and follow-up. Further investigation is needed to explore the factors that influence the effectiveness of PHSE.

scientific communication and education