bioRxiv Science⌕ Search

Biology subjects

Nikolova, O.

Publications and source records attributed to Nikolova, O..

3 recordsLinked to original sources

Transfer learning framework via Bayesian group factor analysis incorporating feature-wise dependencies

Transfer learning considers distinct but related tasks defined over heterogeneous domains and aims to improve generalization and performance through knowledge transfer between tasks. This approach can be especially advantageous in biomedical contexts with insufficient labeled training data, where joint learning across domains can enable inference in otherwise underpowered datasets. High-dimensional biomedical data is characterized with redundancy, rendering non-linear dependencies among features. Existing models often fail to leverage such feature dependencies during inference, limiting their ability to model complex biological systems. We present a Bayesian group factor analysis transfer learning framework that supports multitask, multi-modal learning. Our approach learns a shared latent space within each domain, simultaneously across multiple domains, and uses a feature-wise prior to model complex relationships. We evaluate our framework using controlled synthetic data experiments and four disjoint patient cancer datasets from acute myeloid leukemia and neuroblastoma. We show that our method improves drug response prediction and more readily recapitulates consensus biomarkers of drug response. Similarly, our approach improves tumor purity prediction and identifies a robust gene signature associated with it. Our framework is scalable, interpretable, and adaptable across target phenotypes, offering a robust solution for a wide range of heterogeneous multi-omics problems.

bioinformatics↗

Unifying Multimodal Single-Cell Data Using a Mixture of Experts β-Variational Autoencoder-Based Framework

Multimodal single-cell assays profile complementary layers of cell state, but integration is complicated by modality mismatch, sparsity, and uneven cohort coverage. We present UniVI (Unified Variational Inference), a scalable mixture-of-experts {beta}-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI couples modality-specific encoders/de-coders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or pre-annotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA-protein (CITE-seq) and RNA-chromatin (10x Multiome, SHARE-seq) data spanning human PBMCs and mouse back skin--a non-hematopoietic tissue with continuous differentiation hierarchies--UniVI produces coherent embeddings, improves label transfer, and enables cross-modal reconstruction and denoising. Extending to tri-modal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins (TEA-seq), and accommodates DNA methylation in a paired scNMT-seq mouse gastrulation proof-of-concept under beta-binomial likelihoods. Performance degrades gracefully under severe cell-type imbalance and in the presence of modality-exclusive populations. In an acute myeloid leukemia mosaic design, a paired RNA-protein bridge anchors independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that sharpen with mutation-aware fine-tuning. UniVI thus provides a flexible, interpretable framework for multimodal integration across paired, tri-modal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies.

bioinformatics↗

Predicting transcription factor activity using prior biological information

Transcription factors are critical regulators of cellular gene expression programs. Disruption of normal transcription factor regulation is associated with a broad range of diseases. In order to understand the mechanisms that underly disease pathogenesis, it is critical to detect aberrant transcription factor activity. We have developed Priori, a computational method to predict transcription factor activity from RNA sequencing data. Priori has several key advantages over existing methods. Priori utilizes literature-supported regulatory relationship information to identify known transcription factor target genes. Using these transcriptional relationships, Priori uses linear models to determine the impact and direction of transcription factor regulation on the expression of its target genes. In our work, we evaluated the ability of Priori and 16 other methods to detect aberrant activity from 124 single-gene perturbation experiments. We show that Priori identifies perturbed transcription factors with greater sensitivity and specificity than other methods. Furthermore, our work demonstrates that Priori can be used to discover significant determinants of survival in breast cancer as well as identify mediators of drug response in leukemia from primary patient samples.

bioinformatics↗