bioRxiv Science⌕ Search

Biology subjects

Kitani, A.

Publications and source records attributed to Kitani, A..

4 recordsLinked to original sources

GlycoMeSH: linking glycan structures to biomedical context for systematic enrichment analysis

Glycan identification has advanced, but glycan structures remain difficult to translate into reproducible biomedical context because reusable glycan-level annotations are sparse. We present GlycoMeSH, a resource that links glycans to Medical Subject Headings (MeSH) through an inference model, a traceable association database and a glycan-set enrichment workflow. GlycoMeSH-BERT recovered ~60% of literature-derived associations at recall@30 and expanded open-vocabulary MeSH coverage beyond closed-label baselines, without higher per-prediction accuracy. At matched candidate counts, its predictions showed motif-level semantic agreement comparable to those baselines, independently of the training labels. GlycoMeSH-DB contains 789,627 associations between 26,954 glycans and 20,302 MeSH terms. GlycoMeSH-EA returned enriched MeSH terms for glycan sets from glycomics and glycoproteomics datasets. Each association represents a biomedical context rather than a validated mechanism, and retains its source PMID or prediction score for audit. GlycoMeSH supplies the missing, evidence-traceable annotation layer that makes glycan sets directly analyzable by enrichment across glycoscience datasets.

bioinformatics↗

GlycanGT: A Foundation Model for Glycan Graphs with Pretrained Representation and Generative Learning

MotivationGlycans are highly diverse biological sequences, but their functional understanding has lagged behind that of proteins and nucleic acids. Many glycans remain incompletely characterized or ambiguously annotated, limiting computational analyses. Existing computational approaches are primarily graph-based, capturing local structural features but struggling to model global patterns and incomplete sequences. ResultsWe present GlycanGT, a foundation model for glycans built on a graph transformer architecture. Glycans were represented as graphs with monosaccharides as nodes and glycosidic bonds as edges, and the model was pretrained using a masked language modeling objective. GlycanGT demonstrated higher performance than existing methods across 8 benchmark classification tasks (e.g., 0.734 Macro-F1 in domain prediction and 0.844 AUPRC for immunogenicity classification), and its embeddings formed biologically meaningful clusters that recovered known N- and O-glycan categories. Moreover, GlycanGT accurately proposed candidates for ambiguous sequences, maintaining >80% top-5 accuracy for both monosaccharide and glycosidic bond predictions under high masking levels. Availability and implementationThe pretrained GlycanGT model weights and usage scripts are available on Hugging Face: https://huggingface.co/Akikitani295/GlycanGT. Additional scripts used for analyses in the paper are publicly available on GitHub: https://github.com/matsui-lab/GlycanGT. Contact: matsui.yusuke.d4@f.mail.nagoya-u.ac.jp

bioinformatics↗

Predicting Alzheimer's Cognitive Resilience Score: A Comparative Study of Machine Learning Models Using RNA-seq Data

BackgroundCognitive resilience (CR) in Alzheimers disease (AD) refers to preserved cognitive function despite substantial AD pathology. Diverse biological processes have been implicated in CR, including synaptic maintenance, neuroimmune regulation, and metabolic homeostasis. However, how these mechanisms are organized into molecularly distinct CR subtypes and relate to clinical and neuroanatomical heterogeneity remains unclear. Here, we applied a machine learning framework to multi-cohort transcriptomic, proteomic, and neuroimaging data to investigate molecular subtypes of CR in AD. MethodsRNA-seq data from the Religious Orders Study and Memory and Aging Project (ROSMAP) cohort were used to train machine learning models classifying individuals with AD pathology as CR or non-CR based on residual-based resilience scores. Model development and performance estimation used nested cross-validation to minimize information leakage. Final ROSMAP-trained models were evaluated in the independent Mount Sinai Brain Bank (MSBB) cohort. Model-derived genes were used for biological interpretation and hierarchical clustering of CR individuals. The subtype structure was further evaluated in the Alzheimers Disease Neuroimaging Initiative (ADNI) cohort using cerebrospinal fluid proteomics, MRI-derived brain measures, and longitudinal MMSE data. ResultsMachine learning models showed modest but consistent predictive performance in ROSMAP, with out-of-fold AUROC values of 0.644-0.688. In the independent MSBB full cohort, AUROC values were 0.586-0.659, with improved discrimination in a top/bottom quartile analysis. Hierarchical clustering identified two major molecular subgroups among CR individuals in ROSMAP/MSBB RNA-seq data. A reduced 22-gene/protein signature showed a partial, cluster-like resemblance to this structure in ADNI cerebrospinal fluid proteomics. In ADNI, both projected CR subtypes showed preserved brain tissue-volume profiles and slower longitudinal MMSE decline compared with non-CR participants, whereas clear differences between CR subtypes were not observed. Differential CSF proteomic analysis suggested partially distinct molecular characteristics. ConclusionsThese findings suggest that CR in AD may encompass molecularly heterogeneous, subtype-like profiles that converge on broadly preserved brain structure and slower cognitive decline. Our results provide a candidate framework for stratifying resilience-associated molecular phenotypes in AD and warrant prospective and experimental validation. We also developed the Resilience Gene Analyzer, a web-based platform for visualizing gene-level contributions to CR prediction (https://igcore.cloud/GerOmics/REsilienceGeneAnalyzer/).

bioinformatics↗

GPNMB+ microglia moderate the amyloid beta-tau interaction in early Alzheimer's disease

BackgroundAlthough interactions between amyloid-beta and tau proteins have been implicated in Alzheimers disease (AD), the precise mechanisms by which these interactions contribute to disease progression are not yet fully understood. Moreover, despite the growing application of deep learning in various biomedical fields, its application in integrating networks to analyze disease mechanisms in AD research remains limited. In this study, we employed BIONIC, a deep learning-based network integration method, to integrate proteomics and protein-protein interaction data, with an aim to uncover factors that moderate the effects of the A{beta}-tau interaction on mild cognitive impairment (MCI) and early-stage AD. MethodsProteomic data from the ROSMAP cohort were integrated with protein-protein interaction (PPI) data using a Deep Learning-based model. Linear regression analysis was applied to histopathological and gene expression data, and mutual information was used to detect moderating factors. Statistical significance was determined using the Benjamini-Hochberg correction (p < 0.05). ResultsOur results suggested that astrocytes and GPNMB+ microglia moderate the A{beta}-tau interaction. Based on linear regression with histopathological and gene expression data, GFAP and IBA1 levels and GPNMB gene expression positively contributed to the interaction of tau with A{beta} in non-dementia cases, replicating the results of the network analysis. ConclusionsThese findings indicate that GPNMB+ microglia moderate the A{beta}-tau interaction in early AD and therefore are a novel therapeutic target. To facilitate further research, we have made the integrated network available as a visualization tool for the scientific community (URL: https://igcore.cloud/GerOmics/AlzPPMap).

systems biology↗