bioRxiv Science⌕ Search

Biology subjects

Himori, K.

Publications and source records attributed to Himori, K..

5 recordsLinked to original sources

Multi-omics definition of the sex-specific glycoproteome of murine tissues

Sex-specific differences in the glycoproteome remain poorly defined despite growing evidence that protein glycosylation is a key determinant of sex biology. Here we present a tissue-resolved glycoproteome atlas of adult male and female C57BL/6J mice, integrating transcriptomics, proteomics and glycoproteomics with sialic acid speciation and lectin microarray profiling across 19 tissues. Quantitative analysis of >26,800 protein- and site-specific N-glycoforms from 1,512 glycoproteins revealed highly distinct tissue glycoproteomes shaped by coordinated regulation of protein abundance and glyco-enzyme expression. Multi-omics integration identified strong glycophenotype-enzyme relationships, including control of tissue sialylation by Cmas and Cmah, suggesting rate-limiting roles in glycosylation. Pronounced sex-linked glycophenotypes were observed in salivary gland, liver and kidney, driven by differences in fucosylation, sialylation and protein abundance, whereas the brain glycome was largely conserved between sexes. An interactive online database (https://igcore.cloud/mta/atlas-viewer/) provides a resource for exploring sex-biased glycosylation across mouse tissues.

cell biology↗

GlycoTraitR: an R package for characterizing structural heterogeneity in N-linked glycoproteomics data

Glycoproteomics data are rapidly accumulating due to advances in mass spectrometry instrumentation and the development of specialized search engines (e.g., pGlyco3, Glyco-Decipher) that enable identification of N-linked glycopeptide spectral matches (GPSMs) together with glycan structures. These advances have greatly expanded the scale and depth of N-linked glycopeptides; however, the intrinsic structural heterogeneity of glycosylation remains challenging to interpret. No existing tool provides a unified trait-based framework for analyzing N-linked GPSM data at both the glycosylation-site and protein levels. We developed glycoTraitR, an R package for trait-based analysis of structural heterogeneity in N-linked glycoproteomics data. GlycoTraitR provides a unified workflow to import GPSMs from search engine outputs, extract biologically interpretable glycan structural traits, and perform comparative analyses of micro- and macro-heterogeneity across experimental conditions using statistical testing. ImplementationThe R package and the source code of glycoTraitR are freely available on github at https://github.com/matsui-lab/glycoTraitR. A more detailed introduction and quick start guide are avaible at https://matsui-lab.github.io/glycoTraitR/.

bioinformatics↗

GlycanGT: A Foundation Model for Glycan Graphs with Pretrained Representation and Generative Learning

MotivationGlycans are highly diverse biological sequences, but their functional understanding has lagged behind that of proteins and nucleic acids. Many glycans remain incompletely characterized or ambiguously annotated, limiting computational analyses. Existing computational approaches are primarily graph-based, capturing local structural features but struggling to model global patterns and incomplete sequences. ResultsWe present GlycanGT, a foundation model for glycans built on a graph transformer architecture. Glycans were represented as graphs with monosaccharides as nodes and glycosidic bonds as edges, and the model was pretrained using a masked language modeling objective. GlycanGT demonstrated higher performance than existing methods across 8 benchmark classification tasks (e.g., 0.734 Macro-F1 in domain prediction and 0.844 AUPRC for immunogenicity classification), and its embeddings formed biologically meaningful clusters that recovered known N- and O-glycan categories. Moreover, GlycanGT accurately proposed candidates for ambiguous sequences, maintaining >80% top-5 accuracy for both monosaccharide and glycosidic bond predictions under high masking levels. Availability and implementationThe pretrained GlycanGT model weights and usage scripts are available on Hugging Face: https://huggingface.co/Akikitani295/GlycanGT. Additional scripts used for analyses in the paper are publicly available on GitHub: https://github.com/matsui-lab/GlycanGT. Contact: matsui.yusuke.d4@f.mail.nagoya-u.ac.jp

bioinformatics↗

HuTAge: a Comprehensive Human Tissue- and Cell-specific Ageing Signature Atlas

SummaryAgeing is a complex process that involves interorgan and intercellular interactions. To obtain a clear understanding of ageing, cross-tissue single-cell data resources are required. However, a complete resource for humans is not available. To bridge this gap, we developed HuTAge, a comprehensive resource that integrates cross-tissue age-related information from The Genotype-Tissue Expression project with cross-tissue single-cell information from Tabula Sapiens to provide human tissue- and cell-specific ageing molecular information. Availability and ImplementationHuTAge is implemented within an R Shiny application and can be freely accessed at https://igcore.cloud/GerOmics/HuTAge/home. The source code is available at https://github.com/matsui-lab/HuTAge. Contacthimori.koichi.b5@f.mail.nagoya-u.ac.jp

bioinformatics↗

A Computational Approach to Interpreting the Embedding Space of Dimension Reduction

Nonlinear dimension reduction methods are widely applied in studies analyzing gene and protein expression, by revealing patterns of discrete groups and continuous orders in high-dimensional data. However, the tools are limited to understanding the obtained embedding structures of biological mechanisms, hindering the full exploitation of data. Here, we propose a novel framework to interpret embedding systematically by identifying and mapping associated biological functions. The method performs statistical tests and visualizes significantly enriched functions essential for the organization of the embedding structure, by applying it to the embedding results of two datasets: the Genotype Tissue Expression dataset and a Caenorhabditis elegans embryogenesis dataset, one capturing distinct cluster structures and the other capturing continuous developmental trajectories. We identified the associated functions for interpreting the two embeddings and confirmed it as a useful explainable AI tool in exploratory data analysis by providing annotations to the embedding space.

bioinformatics↗