bioRxiv Science⌕ Search

Biology subjects

chen, w.

Publications and source records attributed to chen, w..

4 recordsLinked to original sources

CellTosg2Sequence: A Unified Text-Omics-Signaling-Graph Large Language Model for Single-Cell Analysis

In single-cell (sc)-based scientific discovery, text-formatted biomedical prior knowledge and signaling graphs are essential for annotating and interpreting numeric sc-omics data and for generating novel testable hypotheses. A major limitation of existing single-cell large language models (scLLMs) is that they rely on numeric expression data with gene names as the only textual signal, while comprehensive biomedical priors -- cellular localization, gene function, disease associations, and signaling interaction patterns -- remain absent from the model input. We introduce CellTosg2Sequence, a textual-prior- and signaling-graph-augmented cell-omics-sentence language model. A lightweight heterogeneous graph encoder maps a curated 62,507-node biomedical knowledge graph (KG) into compact virtual tokens that are prepended to each cell sentence, allowing the language model to condition on biological structure with minimal sequence-length overhead. We train CellTosg2Sequence with a three-stage objective: Stage I anchors the KG channel under autoregressive language-model pretraining, leveraging Qwen2.5-32Bs own language reasoning for rapid KG alignment; Stage II aligns labels via supervised fine-tuning with KG-anchored InfoNCE; Stage III applies Group Relative Policy Optimization (GRPO) with an ontology-hierarchy reward, enabling free-generation cell-type prediction that generalizes beyond the closed training vocabulary. Across multiple benchmarks and ablation experiments, CellTosg2Sequence outperforms strong baselines. All results are achieved with lightweight LoRA training and a single unified checkpoint. Data ethicsThis work uses publicly available single-cell datasets from the Human Cell Atlas (https://www.humancellatlas.org) and the Tahoe-100M consortium. All HCA constituent studies were collected under appropriate donor consent and institutional oversight as described in their original publications; we perform computational re-analysis only and introduce no new human subjects data. HCA data access follows the HCA Data Portal terms of use. No new patient data are collected in this study; no additional IRB approval is required for this secondary computational analysis.

bioinformatics↗

Dual-view Guided Context-aware Network for Automated Bone Lesion Segmentation and Quantification in Whole-body SPECT

Whole-body SPECT bone scintigraphy reflects skeletal metabolic activity throughout the body and plays an indispensable role in the screening, treatment evaluation, and prognostic assessment of bone metastases in tumors. However, the automatic detection and segmentation of hypermetabolic bone lesions remain challenging due to low contrast, limited spatial resolution, and complex lesion distributions. In this study, we proposed Bone-Segnet, a dual-view guided automatic segmentation network for hypermetabolic bone lesions that integrated multi-scale feature modeling, global context modeling, and view-conditioned modulation. Pixel-level annotated anterior and posterior whole-body bone scintigraphy images were used for model training and prediction. The proposed network enhanced the recognition of low-contrast and small-scale lesions through small-lesion enhancement and multi-scale contextual modeling. A Transformer module was further introduced to strengthen global feature representation, while cross-view collaborative modeling was achieved by incorporating the complementary characteristics of anterior and posterior imaging. Experimental results demonstrated that the proposed method outperformed existing approaches across multiple evaluation metrics, with the Dice score improving from 0.7440 to 0.8750, indicating a substantial improvement in segmentation performance. Further quantitative analysis based on the segmentation results revealed significant differences among disease types in lesion count, pixel burden, and spatial distribution patterns, reflecting the heterogeneity of disease-related skeletal metabolic activity. Overall, the proposed method improved automatic lesion segmentation performance and enabled quantitative analysis of lesion burden and spatial distribution patterns, providing objective data support for the assessment of related diseases. Index Terms--Whole-body SPECT, bone lesion segmentation, dual-view modeling, quantitative analysis.

bioinformatics↗

A Gm6AG-binding protein from Vibrio cholerae.

Most known modification-dependent restriction endonucleases target 5-methylcytosine, only a few N6-methyladenine (6mA)-dependent restriction endonucleases have been well-characterized, and the majority of them recognize the G6mATC motif (e.g., DpnI, HHPV4I). Here, we report the identification of a novel 6mA-dependent DNA-binding protein from Vibrio cholerae, VchI, which specifically recognizes the G6mAG motif. VchI contains a winged helix (wH) domain that is homologous to the wH domain in DpnI. However, several key residues involved in 6mA recognition differ between VchI and DpnI, which may contribute to the discrepancy in their recognition specificities. These findings advance our understanding of prokaryotic 6mA modification diversity and the 6mA recognition mechanism of the wH domain, while simultaneously providing an innovative tool for epigenetic research.

biochemistry↗

Detection of Multiple Types of Cancer Driver Mutations Us-ing Targeted RNA Sequencing in NSCLC

Currently, DNA and RNA are used separately to capture different types of gene mutations. DNA is commonly used for the detection of SNVs, indels and CNVs; RNA is used for analysis of gene fusion and gene expression. To perform both DNA sequencing (DNA-seq) and RNA-seq, material is divided into two copies, and two different procedures are required for sequencing. Due to overconsumption of samples and experimental process complexity, it is necessary to create an experimental method capable of analyzing SNVs, indels, fusions and expression. We developed an RNA-based hybridization capture panel targeting actionable driver oncogenes in solid tumors and corresponding sample preparation and bioinformatics workflows. Analytical validation with an RNA standard reference containing 16 known fusion mutations and 6 SNV mutations demonstrated a detection specificity of 100.0% [95% CI 88.7%~100.0%] for SNVs and 100.0% [95% CI 95.4%~100.0%] for fusions. The targeted RNA panel achieved a 0.73-2.63 copies/ng RNA lower limit of detection (LOD) for SNVs and 0.21-6.48 copies/ng RNA for fusions. Gene expression analysis revealed a correlation greater than 0.9 across all 15 cancer-related genes between the RNA-seq results and targeted RNA panel. Among 1253 NSCLC FFPE tumor samples, multiple mutation types were called from DNA- and RNA-seq data and compared between the two assays. The DNA panel detected 103 fusions and 21 METex14 skipping events; 124 fusions and 26 METex14 skipping events were detected by the target RNA panel; 21 fusions and 4 METex14 skipping events were only detected by the target RNA panel. Among the 173 NSCLC samples negative for targetable mutations by DNA-seq, 15 (15/173, 8.67%) showed targetable gene fusions that may change clinical decisions with RNA-seq. In total, 226 tier I and tier II missense variants for NSCLC were analyzed at genomic (DNA-seq) and transcriptomic (RNA-seq) levels. The positive percent agreement (PPA) was 97.8%, and the positive predictive value (PPV) was 98.6%. Interestingly, variant allele frequencies were generally higher at the RNA level than at the DNA level, suggesting relatively dominant expression of mutant alleles. PPA was 97.6% and PPV 99.38% for EGFR 19del and 20ins variants. We also explored the relationship of RNA expression with gene copy number and protein expression. The RPKM of EGFR transcripts assessed by the RNA panel showed a linear relationship with copy number quantified by the DNA panel, with an R of 0.8 in 1253 samples. In contrast, MET gene expression is regulated in a more complex manner. In IHC analysis, all 3+ samples exhibited higher RPKM levels; IHC level of 2+ and below showed lower RNA expression. Parallel DNA- and RNA-seq and systematic analysis demonstrated the accuracy and robustness of the RNA sequencing panel in identifying multiple types of variants for cancer therapy. Contact: zhaojia0327@126.com

molecular biology↗