bioRxiv Science⌕ Search

Biology subjects

Kiyota, B.

Publications and source records attributed to Kiyota, B..

3 recordsLinked to original sources

Global tree encoding of atlas-scale single-cell genomics

The rapid expansion of single-cell genomic datasets has led to the compilation of biological resources comprising hundreds of millions of cells across tissues, developmental stages, and disease states. This has underscored the need for scalable and interpretable data representations that preserve the complex relationships and multi-scale organization of cellular states, while remaining computationally tractable at atlas scale. Existing approaches based on discrete abstractions have enabled cell annotation, clustering, and trajectory inference, but are often optimized for local inference tasks and may obscure continuous cellular relationships and multi-resolution structure within complex transcriptional and other genomic landscapes. Moreover, increasing dataset sizes often require information-reduction strategies such as random downsampling, limiting the resolution of rare cell populations and heterogeneous cellular states. Here, we present MILK, a scalable computational framework that organizes high-dimensional single-cell populations into unified tree representations. Across large-scale transcriptomic atlases, MILK enables representative subsampling with preserved information, supporting the tractable application of existing algorithms for tasks including deep generative model training and foundation model benchmarking. Additionally, MILK enables holistic, multi-resolution analyses that capture global developmental trajectories, characterize disease-associated cellular perturbations across tissues, and facilitate comparison of transcriptional programs across species within a coherent hierarchical framework. Together, these results establish the hierarchical organization of biological data as a scalable and unifying representation of cellular identity, enabling integrative analysis of single-cell genomic data across diverse contexts.

genomics↗

A cuffed CRISPR guide RNA for microRNA activity-dependent genome editing

Cells in multicellular eukaryotic systems are diverse biological units, with characteristics and functions determined by their molecular profiles. CRISPR-Cas9 genome editing has been widely used across biology to modulate gene expression and study gene function. However, there is currently no versatile and scalable method for editing a cells genome in response to endogenous cellular signals. Here, we report the engineering of a CRISPR guide RNA that efficiently confers genome editing in response to the catalytic activity of a target microRNA (miRNA) within a cell. miRNAs are short non-coding RNAs that are widely conserved across eukaryotes and can cleave their target RNA through almost perfect base pairing. In mammals, miRNAs are largely involved in development and homeostasis as well as disease progression and developmental disorders. To leverage these properties for genome editing, we developed a cuffed guide RNA (cgRNA) which is composed of a permutated order of sequence domains from the commonly used single guide RNA (sgRNA). These permutated domains were then concatenated with a miRNA target sequence, yielding a warped guide RNA that is inactive until cleaved by a complementary miRNA. We demonstrated that cgRNA enabled efficient miRNA activity-dependent genome editing in human and mouse cell lines. Biochemical and structural analyses revealed three stages of inhibition of the CRISPR genome-editing pathway for unprocessed cgRNA. Utilizing a lentiviral library of cgRNAs containing miRNA targets covering mouse genome-wide miRNAs, we identified miRNA cleavage activities and their sequence specificities in mouse embryonic stem cells and during smooth muscle cell differentiation. Furthermore, we showed that endogenous mRNA expression could be irreversibly recorded into a DNA sequence using a cgRNA targeted by a synthetic miRNA repeat. cgRNA is a simple, robust, miRNA activity-gated genome editing system that could facilitate the development of cell state-specific genome editing, the mapping of miRNA activity and gene expression landscapes, and the recording of molecularly determined cell states during the long-term progression of multicellular systems.

bioengineering↗

Detecting and avoiding homology-based data leakage in genome-trained sequence models

Models that predict function from DNA sequence have become critical tools in deciphering the roles of genomic sequences and genetic variation within them. However, traditional approaches for dividing the genomic sequences into training data, used to create the model, and test data, used to determine the models performance on unseen data, fail to account for the widespread homology within genomes. Using simulations, we illustrate how homology-based data leakage can lead to overestimation of model performance. Across a variety of genomics models, we demonstrate that performance on test sequences varies systematically by their similarity with training sequences. Models generally perform well on distant sequences, reflecting the application of learned generalizable principles. At higher and intermediate similarity, models rely on memorized associations, inflating performance when function is conserved between homologs but failing when homologous sequences have functionally diverged. To dissect and mitigate these effects, we introduce hashFrag, a scalable solution for homology detection and data partitioning. Using hashFrag, we demonstrate how to create homology-aware evaluations of model performance, and improve model generalizability by providing improved splits for model training. Altogether, we establish how homology creates a systematic bias in genome-trained models and must be accounted for to ensure reliable evaluation of sequence-to-function predictors.

bioinformatics↗