bioRxiv · 10.1101/2025.10.24.684446
GRASS-NB: Group-structured variable selection for spatial negative binomial data with applications to cancer registry and spatial omics
Abstract
Spatially structured, overdispersed count data with high-dimensional predictors are increasingly observed across studies from population-level epidemiology to cellular-level spatial omics. Feature selection is critical to identify influential predictors, such as key risk factors or biomarkers. Few Bayesian studies have assessed negative binomial regression (NBR) models with standard variable selection priors, like the mixture spike-and-slab (SS) or continuous horseshoe (HS), but mostly under aspatial settings. Features often form groups; for instance, in population surveys, caloric intake and physical activity may fall under "Diet & Exercise", while cigarette use and smoking laws belong to "Smoking". We propose a flexible NBR model that accommodates spatial autocorrelation and introduces a novel group-structured prior by hybridizing SS and HS shrinkage. The models performance with different priors is evaluated in terms of specificity, precision, and computational cost under challenging scenarios, including "large p, small n" cases. We further apply the model to CDC state-level cancer data, comprising demographic, screening, and behavioral covariates, to identify key drivers and population-level risk factors, and to a melanoma spatial omics dataset for predictive modeling expression of gene. An efficient R package is provided on GitHub.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mattila, C., Neelon, B., Sonawane, K., Cao, S., Angel, P., Hill, E., Seal, S.. 2025-10-25. GRASS-NB: Group-structured variable selection for spatial negative binomial data with applications to cancer registry and spatial omics. https://doi.org/10.1101/2025.10.24.684446
Cite the original work for its findings. Save a collection to share your selection of sources.