bioRxiv · 10.1101/678029
Mixture Network Regularized Generalized Linear Model with Feature Selection
Abstract
High dimensional genomics data in biomedical sciences is an invaluable resource for constructing statistical prediction models. With the increasing knowledge of gene networks and pathways, such information can be utilized in the statistical models to improve prediction accuracy and enhance model interpretability. However, in certain scenarios the network structure may only be partially known or subject to inaccuracy. Thus, the performance of statistical models incorporating such network structure may be compromised. In this paper, we propose a weighted sparse network learning method by optimally combining a data driven network with sparsity property to prior known or partially known network to address this issue. We show that our proposed model attains the oracle property and achieves a parsimonious structure in high dimensional setting for different types of outcomes including continuous, binary and survival data. Simulations studies show that our proposed model is robust and outperforms existing methods. Case study on melanoma gene expression further demonstrates that our proposed model achieves good operating characteristics in identifying informative genes and predicting survival risk. An R package glmaag implementing our method is available on the Comprehensive R Archive Network (CRAN).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Li, K., Wang, X., Kuan, P. F.. 2019-06-21. Mixture Network Regularized Generalized Linear Model with Feature Selection. https://doi.org/10.1101/678029
Cite the original work for its findings. Save a collection to share your selection of sources.