bioRxiv · 10.64898/2026.03.03.709470
bMAE: Masked Autoencoder Latent Representations for Bulk RNA-seq Tissues
Abstract
Bulk tissue RNA-sequencing data from large-scale consortia such as GTEx provide comprehensive gene expression profiles across diverse human tissues. However, the high-dimensional nature of bulk RNA-seq data, combined with technical noise and batch effects, poses challenges for downstream analyses. While dimensionality reduction methods are routinely applied, standard approaches such as PCA often fail to optimally preserve tissue-discriminative information and exhibit poor generalization to unseen tissue types. We developed a masked autoencoder for bulk tissue RNA-seq that learns compressed latent representations through self-supervised learning with variable masking schedules. Evaluated on GTEx data (31 tissue types, 19,788 samples, 19,308 genes), our method substantially outperformed all baselines in leave-one-tissue-out (LOTO) cross-validation. We achieved mean silhouette 0.20, ARI 0.58, and NMI 0.84 versus best baseline UMAP (0.007, 0.25, 0.47), representing 28.6-fold, 2.3-fold, and 1.8-fold improvements. Remarkably, within held-out tissue categories containing subtissues, our latent space revealed hierarchical structure with enhanced subtissue separation (silhouette 0.16, ARI 0.35, NMI 0.36) versus baselines (silhouette 0.15, ARI 0.20, NMI 0.20), despite never observing these distinctions during training. The method compressed 19,308 genes to 128 dimensions while preserving multi-scale structure.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wan, Z., Untalan, M. Z. G., Vasconcellos Vargas, D.. 2026-03-05. bMAE: Masked Autoencoder Latent Representations for Bulk RNA-seq Tissues. https://doi.org/10.64898/2026.03.03.709470
Cite the original work for its findings. Save a collection to share your selection of sources.