bioRxiv Science⌕ Search

Biology subjects

Rabuzin, L.

Publications and source records attributed to Rabuzin, L..

2 recordsLinked to original sources

DeepCAST-GWAS: Improving the Discovery of Genetic Associations Using Deep Learning-Based Regulatory SNP Prioritization

Genome-wide association studies (GWAS) have uncovered numerous variants linked to complex traits, yet power remains limited by the large multiple testing burden and the inclusion of many variants with minimal regulatory impact. We present Deep learning-based Chromatin Accessibility SNP Targeting for GWAS (DeepCAST-GWAS), a framework that integrates functional annotations derived from deep learning models to improve both the yield and the reliability of GWAS findings. DeepCAST-GWAS uses SNP Activity Difference (SAD) scores from in silico mutagenesis with the Enformer model to estimate the predicted effect of each variant on chromatin accessibility across tissues, allowing statistical testing to focus on variants with stronger regulatory evidence. Using conservative family-wise error rate (FWER) control, DeepCAST-FWER produces fewer associations than existing power-boosting approaches, but the associations it reports replicate in larger cohort GWAS at substantially higher rates. For applications where discovery count is more important, DeepCAST-sFDR increases the number of genome-wide significant findings above baseline GWAS by using the Enformer SAD scores for stratified False Discovery Rate (sFDR) control. DeepCAST-sFDR achieves performance comparable to the strongest competing method, while maintaining reliability on par with a standard GWAS. Subsampling analyses across a wide range of traits confirm these improvements in both sensitivity and replicability. DeepCAST-GWAS offers a principled way to incorporate sequence-based regulatory predictions into population-scale association testing, demonstrating that chromatin accessibility activity scores can improve the stability of GWAS discoveries. The framework is made available at https://github.com/BoevaLab/DeepCAST-GWAS.

bioinformatics↗

Exploring Augmentation-Driven Invariances for Graph Self-supervised Learning in Spatial Omics

Spatial omics technologies provide rich insights into biological processes by jointly capturing molecular profiles and the spatial organization of cells. The resulting high-dimensional data can be naturally represented as graphs, where Graph Neural Networks (GNNs) offer an effective framework to model interactions in the tissue. Self-supervised pretraining methods such as Bootstrapped Graph Latents (BGRL) and GRACE leverage graph augmentations to build invariances without costly labels. Yet, the design of augmentation strategies remains underexplored, particularly in the context of spatial omics. In this work, we systematically investigate how different graph augmentations affect embedding quality and downstream performance in spatial omics. We evaluate a suite of existing and novel augmentations, including transformations tailored to biological variation, across two representative tasks: unsupervised domain identification in healthy tissue and supervised phenotype prediction in cancer tissue. Our results show that carefully chosen augmentations substantially improve performance, whereas poorly aligned or overly complex augmentations may fail to help or even degrade performance. These findings highlight the central role of augmentation design in enforcing meaningful invariances for graph contrastive pretraining in spatial omics.

bioinformatics↗