bioRxiv Science⌕ Search

Biology subjects

Koretsky, M. J.

Publications and source records attributed to Koretsky, M. J..

3 recordsLinked to original sources

Random forest model improves annotation and discovery of variants of uncertain significance in Alzheimer's and other neurological disorders

Variants of uncertain significance (VUS) are a bottleneck for genetic discovery and complicate clinical decision-making in Alzheimers disease and related neurological disorders (ADRD). We developed MoVUS: Model for Variants of Unknown Significance, a random-forest approach that integrates functional predictors to classify missense VUS. MoVUS leverages a balanced random forest model trained on dbNSFP v5.1a with high-confidence ClinVar and HGMD labels, using harmonized functional prediction rankscores. MoVUS produced confident, explainable calls, with [~]98% accuracy (AUC [~]0.998), prioritizing potentially pathogenic candidates and down-ranked likely benign variants on independent validation sets of ClinVar-only and HGMD-only variants. In our discovery analyses on ADRD-implicated variants in dbNSFP and from independent collaborator cohorts, we achieved high-confidence classifications on a majority of the unknown variants (average of 55% of discovery variants). We also had access to medical records and family trees for some variants, further validating our findings. Across held-out and external datasets, MoVUS reports high accuracy alongside confidence scores and helps prioritize actionable candidates, and reduces bias by considering multiple scores for each variant. To facilitate use, we developed a web app for users to browse across 100+ ADRD genes. MoVUS provides transparent, reproducible triage for research follow-up by pairing consensus predictors with SHAP-based visualizations and explanations.

genetics↗

CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research

BackgroundsBiomedical research requires sophisticated understanding and reasoning across multiple specializations. While large language models (LLMs) show promise in scientific applications, their capability to safely and accurately support complex biomedical research remains uncertain. MethodsWe present CARDBiomedBench, a novel question-and-answer benchmark for evaluating LLMs in biomedical research. For our pilot implementation, we focus on neurodegenerative diseases (NDDs), a domain requiring integration of genetic, molecular, and clinical knowledge. The benchmark combines expert-annotated question-answer (Q/A) pairs with semi-automated data augmentation, drawing from authoritative public resources including drug development data, genome-wide association studies (GWAS), and Summary-data based Mendelian Randomization (SMR) analyses. We evaluated seven private and open-source LLMs across ten biological categories and nine reasoning skills, using novel metrics to assess both response quality and safety. ResultsOur benchmark comprises over 68,000 Q/A pairs, enabling robust evaluation of LLM performance. Current state-of-the-art models show significant limitations: models like Claude-3.5-Sonnet demonstrates excessive caution (Response Quality Rate: 25% [95% CI: 25% {+/-} 1], Safety Rate: 76% {+/-} 1), while others like ChatGPT-4o exhibits both poor accuracy and unsafe behavior (Response Quality Rate: 37% {+/-} 1, Safety Rate: 31% {+/-} 1). These findings reveal fundamental gaps in LLMs ability to handle complex biomedical information. ConclusionCARDBiomedBench establishes a rigorous standard for assessing LLM capabilities in biomedical research. Our pilot evaluation in the NDD domain reveals critical limitations in current models ability to safely and accurately process complex scientific information. Future iterations will expand to other biomedical domains, supporting the development of more reliable AI systems for accelerating scientific discovery.

bioinformatics↗

Application of Aligned-UMAP to longitudinal biomedical studies

Longitudinal multi-dimensional biological datasets are ubiquitous and highly abundant. These datasets are essential to understanding disease progression, identifying subtypes, and drug discovery. Discovering meaningful patterns or disease pathophysiologies in these datasets is challenging due to their high dimensionality, making it difficult to visualize hidden patterns. Several methods have been developed for dimensionality reduction, but they are limited to cross-sectional datasets. Recently proposed Aligned-UMAP, an extension of the UMAP algorithm, can visualize high-dimensional longitudinal datasets. In this work, we applied Aligned-UMAP on a broad spectrum of clinical, imaging, proteomics, and single-cell datasets. Aligned-UMAP reveals time-dependent hidden patterns when color-coded with the metadata. We found that the algorithm parameters also play a crucial role and must be tuned carefully to utilize the algorithms potential fully. Altogether, based on its ease of use and our evaluation of its performance on different modalities, we anticipate that Aligned-UMAP will be a valuable tool for the biomedical community. We also believe our benchmarking study becomes more important as more and more high-dimensional longitudinal data in biomedical research becomes available. Highlights- explored the utility of Aligned-UMAP in longitudinal biomedical datasets - offer insights on optimal uses for the technique - provide recommendations for best practices In BriefHigh-dimensional longitudinal data is prevalent yet understudied in biological literature. High-dimensional data analysis starts with projecting the data to low dimensions to visualize and understand the underlying data structure. Though few methods are available for visualizing high dimensional longitudinal data, they are not studied extensively in real-world biological datasets. A recently developed nonlinear dimensionality reduction technique, Aligned-UMAP, analyzes sequential data. Here, we give an overview of applications of Aligned-UMAP on various biomedical datasets. We further provide recommendations for best practices and offer insights on optimal uses for the technique.

bioinformatics↗