bioRxiv Science⌕ Search

Biology subjects

XU, K.

Publications and source records attributed to XU, K..

7 recordsLinked to original sources

A Discrete Language of Protein Words for Functional Discovery and Design

Proteins can preserve conserved functions despite extensive sequence and structural divergence, suggesting that functional organization is governed by distributed constraints not captured by conventional representations. Here we develop a hierarchical sequence-based representation framework that compresses proteins into context-dependent latent states while preserving multiscale organizational information. Using this framework, we identified previously uncharacterized ciliary proteins lacking detectable sequence and structure homology, including ADMAP1, which is required for normal sperm axonemal organization and motility in mice. Discrete latent protein states captured species-level organizational signatures correlated with major evolutionary groups and revealed expansion of intrinsically disordered regulatory environments in eukaryotes. Autoregressive sampling within this latent space further enabled design of synthetic actin-remodeling proteins that maintained robust F-actin severing activity despite extensive sequence rewiring across key functional interfaces. These findings demonstrate that distributed protein organization can be inferred directly from sequence, linking functional discovery, evolutionary analysis, and protein design within a shared representational framework.

bioinformatics↗

A comprehensive functional landscape of α-tubulin TUBA1A variants illuminates microtubule biology and refines clinical classification

Missense variant interpretation in highly conserved, paralog-rich gene families remains a critical bottleneck for precision medicine. Here, we developed an integrated experimental-computational platform to systematically assess the functional impact of all possible missense mutations in the -tubulin TUBA1A. Combining high-throughput comprehensive mutagenesis, high-content imaging, convolutional neural network- driven phenotyping, and machine learning-guided prediction, we quantified microtubule assembly phenotypes for every coding variant. This approach outperforms conservation-based predictors and enables functional reinterpretation of disease-associated variants. Structural mapping of the TUBA1A mutational landscape reveals distinct domains critical for GTP binding, chaperone-assisted folding or protofilament interaction, illuminating diverse mechanisms of tubulin-related diseases. Integration with ACMG-AMP guidelines demonstrates this mutational landscape improves clinical variant classification in redundant gene families. This framework is broadly applicable to other structurally conserved proteins, linking variant effect prediction to mechanistic insight and clinical translation.

genetics↗

Interpretable spatial multi-omics data integration and dimension reduction with SpaMV

Spatial multi-omics technologies have revolutionized our understanding of biological systems by providing spatially resolved molecular profiles from multiple perspectives. Existing spatial multi-omics integration methods often assume that data from different modalities share a common underlying distribution, aiming to project them into a single unified latent space. This assumption, however, can obscure the unique insights offered by each modality, thereby limiting the full potential of multi-omics analyses. To address this limitation, we present the Spatial Multi-View (SpaMV) representation learning algorithm, which captures both the shared information across modalities and the distinct, modality-specific information, enabling a more comprehensive and interpretable representation of spatial multi-omics data. Through extensive evaluation on both simulated and real-world datasets, SpaMV demonstrates superior spatial domain clustering performance and provides users with more interpretable dimension reduction for downstream analysis. Moreover, SpaMV effectively annotates cell types within clusters of a mouse thymus dataset, highlighting its effectiveness in interpretable dimensionality reduction.

bioinformatics↗

SynSeg: Generating Synthetic Datasets for Accurate Subcellular Segmentation with U-net

Accurate segmentation of subcellular components is crucial for understanding cellular processes, but traditional methods struggle with noise and complex structures. Convolutional neural networks improve accuracy but require large, time-consuming, and biased manually annotated datasets. Here, we developed SynSeg, a pipeline that generates synthetic training data to train a U-net model for subcellular structure segmentation, eliminating the need for manual annotation. SynSeg leverages synthetic datasets with variations in intensity, morphology, and signal distribution to deliver context-aware segmentations, even in challenging imaging conditions. We demonstrate SynSegs superior performance in segmenting vesicles and cytoskeletal filaments from culture cells and live C. elegans, outperforming traditional methods such as Otsus thresholding, ILEE, and FilamentSensor 2.0. Additionally, SynSeg effectively quantified disease-associated microtubule morphology in live cells, uncovering structural defects caused by mutant Tau proteins linked to neurodegenerative diseases. These results highlight the potential of synthetic data-driven approaches to advance biological segmentation and enhance microscopy techniques. Significance StatementThis study introduces a novel approach for accurately segmenting cellular structures, such as microtubules and vesicles, using synthetic datasets and advanced deep learning techniques. By leveraging a U-Net model trained on thousands of artificially generated images, our method eliminates the need for labor-intensive experimental data and simplifies the data creation process. Importantly, it incorporates noise and variability into the training datasets to make the model more robust and biologically relevant. Our findings demonstrate that the model can successfully identify cellular components, paving the way for its application in real-world microscopy images. This innovation has the potential to accelerate discoveries in cell biology by providing an efficient, scalable tool for analyzing complex cellular structures, even in challenging imaging conditions.

bioinformatics↗

stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images

Spatially resolved transcriptomics (SRT) and spatially resolved proteomics (SRP) data enable the study of gene expression and protein abundances within their precise spatial and cellular contexts in tissues. Certain SRT and SRP tech-nologies also capture corresponding morphology images, adding another layer of valuable information. However, few existing methods developed for SRT data effectively leverage these supplementary images to enhance clustering performance. Here, we introduce stDyer-image, an end-to-end deep learning framework designed for clustering for SRT and SRP datasets with images. Unlike existing methods that utilize images to complement gene expression data, stDyer-image directly links image features to cluster labels. This approach draws inspiration from pathologists, who can visually identify specific cell types or tumor regions from morphological images without relying on gene expression or protein abundances. Benchmarks against state-of-the-art tools demonstrate that stDyer-image achieves superior performance in clustering. Moreover, it is capable of handling large-scale datasets across diverse technologies, making it a versatile and powerful tool for spatial omics analysis.

bioinformatics↗

stDyer enables spatial domain clustering with dynamic graph embedding

Spatially resolved transcriptomics (SRT) data provide critical insights into gene expression patterns within tissue contexts, necessitating effective methods for identifying spatial domains. Traditional clustering techniques often over-look spatial information, leading to disjointed domains. Current computational approaches integrate spatial information but still face challenges in recognizing domain boundaries, scalability, and the need of independent clustering steps. We introduce stDyer, an end-to-end deep learning framework designed for spatial domain clustering in SRT data. stDyer combines a Gaussian Mixture Variational AutoEncoder (GMVAE) with graph attention networks (GATs) to simultaneously learn deep representations and perform clustering for units. A unique feature of stDyer is the dynamic graphs it adopts, which adaptively links units based on Gaussian Mixture assignments in the latent space, thereby improving spatial domain clustering and producing smoother domain boundaries. Additionally, stDyers mini-batch neighbor sampling strategy facilitates scalability to large datasets and enables multi-GPU training. Benchmarking against state-of-the-art tools across various SRT technologies, stDyer demonstrates superior performance in spatial domain clustering, multi-slice analysis, and large-scale dataset handling.

bioinformatics↗

Artificial Intelligence-Enabled AlphaFold II Pipeline Guides Functional Fluorescence Labeling of Tubulin Across Species

Dynamic properties are essential for microtubule (MT) physiology. Current techniques for in vivo imaging of MTs present intrinsic limitations in elucidating the isotype-specific nuances of tubulins, which contribute to their versatile functions. Harnessing the power of AlphaFold II pipeline, we engineered a strategy for the minimally invasive fluorescence labeling of endogenous tubulin isotypes or those harboring missense mutations. We demonstrated that a specifically designed 16-amino acid linker, coupled with sfGFP11 from the split-sfGFP system and integration into the H1-S2 loop of tubulin, facilitated tubulin labeling without compromising MT dynamics, embryonic development, or ciliogenesis in C. elegans. Extending this technique to human cells and murine oocytes, we visualized MTs with the minimal background fluorescence and a pathogenic tubulin isoform with fidelity. The utility of our approach across biological contexts and species set an additional paradigm for studying tubulin dynamics and functional specificity, with implications for understanding tubulin-related diseases known as tubulinopathies.

cell biology↗