bioRxiv Science⌕ Search

Biology subjects

Arnoldt, L.

Publications and source records attributed to Arnoldt, L..

2 recordsLinked to original sources

Biologically Guided Variational Inference for Interpretable Multimodal Single-Cell Integration and Mechanistic Discovery

Multi-omics technologies allow detailed characterization of cell types and states across omics layers as well as chemical and genetic perturbations. Variational autoencoders have become a cornerstone of single-cell data integration; however, they are often implemented as black-box models, requiring post hoc interpretation using known markers, pathways, or regulators to make sense of their latent representations. NetworkVI fundamentally flips this paradigm by incorporating biological knowledge directly into the model architecture. By embedding co-regulation networks derived from topologically associated domains and structured ontologies such as the Gene Ontology (GO), NetworkVI introduces a biologically-informed inductive bias that promotes the preservation of meaningful variation during integration while enforcing interpretability at both the gene and GO levels. NetworkVI achieves state-of-the-art data integration, modality imputation, and cell label transfer across bimodal and trimodal datasets. Beyond integration, here we show that NetworkVI facilitates ontology-guided hypothesis generation by exploiting established associations between genes, structured cellular programs, and regulatory domains to interpretably model cellular identities. Furthermore, decomposition of GO activation spaces resolves lineage-specific functional states within immune cell types, including quiescent, inflammatory, and transitional monocyte subpopulations, that are invisible to transcriptomic clustering. NetworkVI prioritizes GO-term programs associated with immunosenescence, consistent with age-associated immune dysregulation and reveals candidate immune evasion mechanisms consistent with CD58 loss in a Perturb-CITE-seq melanoma dataset.

bioinformatics↗

Joint Generation of Protein Sequence and Structure with RoseTTAFold Sequence Space Diffusion

Protein denoising diffusion probabilistic models (DDPMs) show great promise in the de novo generation of protein backbones but are limited in their inability to guide generation of proteins with sequence specific attributes and functional properties. To overcome this limitation, we develop ProteinGenerator, a sequence space diffusion model based on RoseTTAfold that simultaneously generates protein sequences and structures. Beginning from random amino acid sequences, our model generates sequence and structure pairs by iterative denoising, guided by any desired sequence and structural protein attributes. To explore the versatility of this approach, we designed proteins enriched for specific amino acids, with internal sequence repeats, with masked bioactive peptides, with state dependent structures, and with key sequence features of specific protein families. ProteinGenerator readily generates sequence-structure pairs satisfying the input conditioning (sequence and/or structural) criteria, and experimental validation showed that the designs were monomeric by size exclusion chromatography (SEC), had the desired secondary structure content by circular dichroism (CD), and were thermostable up to 95{degrees}C. By enabling the simultaneous optimization of both sequence and structure, ProteinGenerator allows for the design of functional proteins with specific sequence and structural attributes, and paves the way for protein function optimization by active learning on sequence-activity datasets.

biochemistry↗