bioRxiv Science⌕ Search

Biology subjects

Sui, Z.

Publications and source records attributed to Sui, Z..

6 recordsLinked to original sources

PXN Unlocks the Power of Public Gene Expression Data Through Cross-Technology Integration

The immense value of public gene expression repositories is constrained by the lack of compatibility among datasets generated from diverse experimental technologies. Differences in measurement scales, probe chemistries, and signal distributions create systematic discrepancies across platforms and laboratories. These inconsistencies make large-scale integrative analysis nearly impossible, even though such studies could achieve great statistical power and improved reproducibility. We introduce PXN, a probabilistic machine learning framework that captures a unified representation of biological signal across multiple gene expression technologies. Once trained, PXN can seamlessly translate data between multiple platforms, preserving informative biological variation while removing technology-specific biases. In benchmarking studies, PXN consistently outperforms existing normalization methods in cross-platform accuracy and substantially enhances the power of differential expression analysis. Importantly, we show that PXN is powerful enough to bridge even the most challenging technological divide--between microarray and RNA-seq. This capability provides a scalable route for integrating legacy microarray data with modern RNA-seq studies. By enabling direct comparison and integration of heterogeneous datasets, PXN unlocks the full potential of public repositories for future biological discovery and therapeutic innovation.

bioinformatics↗

Chromosome-level genome assembly of macroalgae Gracilariopsis lemaneiformis

As an important cultivated red alga, Gracilariopsis lemaneiformis has great economic and ecological value. However, its existing genome assembly is highly fragmented and inadequately annotated. In this study, we constructed the first high-quality chromosome-level genome of Gp. lemaneiformis using PacBio long reads, Illumina short reads and Hi-C sequencing data. The assembled genome was approximately 86.66 Mb and the assembled sequences were anchored to 28 pseudo-chromosomes with lengths ranging from 1.70 to 7.81 Mb. 99.91% of the PacBio reads could be mapped to our assembly. In total, 8,664 genes were annotated, and the repeat elements identified in Gp. lemaneiformis constituted 65.04% of the whole genome, including 2.24% tandem repeat sequences and 62.81% interspersed repeats. We also established a high-evidence phylogenetic tree from 19 representative algae species, with the main aim to calculate their divergence times. This high-quality genome of Gp. lemaneiformis provides a crucial foundation for understanding genetic characteristics, investigating the genomic evolution, and facilitating molecular breeding.

genomics↗

Translating Histopathology Foundation Model Embeddings into Cellular and Molecular Features for Clinical Studies

AI-powered pathology foundation models provide general-purpose representations of histopathological images by encoding image tiles into numerical embeddings. However, these embeddings are not directly interpretable in biological or clinical terms and must be translated into biologically meaningful features, such as cell-type composition or gene expression, to enable downstream clinical applications. To bridge this gap, we developed STpath, a framework that integrates histopathology image embeddings derived from existing pathology foundation models with matched, spatially resolved transcriptomics data. STpath consists of cancer-specific XGBoost models trained to infer cell-type compositions and gene expression from histopathology image tiles. We evaluated STpath in colorectal and breast cancer datasets and showed that it provides accurate estimates of the composition of major cell types and the expression of a subset of genes, with further performance gains achieved by combining embeddings from multiple foundation models. Finally, we demonstrated that STpath inferred features that can be used in downstream studies to evaluate their associations with clinical outcomes.

bioinformatics↗

Proteomic Analysis in Alzheimer's Disease with Psychosis Reveals Separate Molecular Signatures for Core AD Proteinopathy and Postsynaptic Density Disruption

Background and Hypothesis: Alzheimers disease with psychosis (AD+P) is a subgroup of AD patients with more rapid cognitive deterioration. While our previous study showed that AD+P is associated with loss of prefrontal cortex postsynaptic density (PSD) proteins, identifying proteins in the broader cellular environment that influence PSD loss addresses a critical knowledge gap about synaptic dysfunction mechanisms in early disease stages. Study Design: We conducted a proteomic analysis comparing prefrontal grey matter cortex tissue homogenates from elderly normal controls (n=18), individuals with AD+P (n=61), and individuals with AD-P (n=48), all with Braak stages 3-5. Study Results: AD+P showed the most pronounced alterations relative to controls (178 proteins with q<0.05), although alterations in AD-P and AD+P relative to controls were highly similar (R{superscript 2}=0.965, p<0.001). Weighted-gene correlation network analysis (WGCNA) identified four modules significantly associated with disease status comparing AD subjects to controls, but none differed significantly between AD+P and AD-P. We identified 15 proteins significantly correlated with PSD yield across all samples, including ENPP6, linked to AD+P by GWAS. Additionally, PSD yield-associated proteins showed minimal overlap with altered AD proteins (1 of 137). WGCNA revealed one module significantly correlated with PSD yield across all samples, enriched for inflammatory terms. Conclusions: Our findings suggest a model in which AD+P arises from the combination of quantitative alterations within a shared AD proteome profile and a superimposed set of protein alterations correlated with PSD yield that are largely independent of the shared AD proteome, conferring distinct mechanisms of synaptic vulnerability and psychosis risk.

neuroscience↗

Functional interaction of hybrid extracellular vesicle-liposome nanoparticles with target cells: absence of toxicity

Building on the success of COVID-19 vaccine development, lipid nanoparticles (LNPs) have emerged as leading vehicles for mRNA delivery in a range of therapeutic applications. Naturally-occurring extracellular vesicles (EVs), which share similar physical properties with LNPs, present a promising alternative platform because of their relative stability and lower immunogenicity. A key challenge common to both EVs and LNPs is enabling efficient vesicle - cell interactions and establishing a polarized permeability pathway required for effective cargo transfer. Membrane recognition and intercalation are essential for the function and delivery capacity of both systems, regardless of their complexity. In this study, we leveraged recent advances to create hybrid extracellular vesicles (HEVs) by using LNPs to load mRNA into EVs. We characterized HEV formation using Forster resonance energy transfer (FRET), cryo-electron microscopy (Cryo-EM), and super-resolution microscopy, and demonstrated their ability to deliver mRNA to recipient cells. In both, in vitro and in vivo models, HEVs exhibited superior transfection efficiency compared to conventional LNPs composed of synthetic lipids, while significantly reducing LNPs cytotoxicity - a not-well-recognized limitation of synthetic lipid-based systems. These results highlight HEVs as a safer and more effective alternative for mRNA and small molecule delivery. Future therapeutic strategies could involve isolating EVs from patients, hybridizing them with synthetic lipid carriers loaded with therapeutic cargo, and reintroducing them for personalized treatment.

cell biology↗

Exploit Spatially Resolved Transcriptomic Data to Infer Cellular Features from Pathology Imaging Data}

Digital pathology is a rapidly advancing field where deep learning methods can be employed to extract meaningful imaging features. However, the efficacy of training deep learning models is often hindered by the scarcity of annotated pathology images, particularly images with detailed annotations for small image patches or tiles. To overcome this challenge, we propose an innovative approach that leverages paired spatially resolved transcriptomic data to annotate pathology images. We demonstrate the feasibility of this approach and introduce a novel transfer-learning neural network model, STpath (Spatial Transcriptomics and pathology images), designed to predict cell type proportions or classify tumor microenvironments. Our findings reveal that the features from pre-trained deep learning models are associated with cell type identities in pathology image patches. Evaluating STpath using three distinct breast cancer datasets, we observe its promising performance despite the limited training data. STpath excels in samples with variable cell type proportions and high-resolution pathology images. As the influx of spatially resolved transcriptomic data continues, we anticipate ongoing updates to STpath, evolving it into an invaluable AI tool for assisting pathologists in various diagnostic tasks.

cancer biology↗