bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Cell Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Establishment of morphological atlas of Caenorhabditis elegans embryo with cellular resolution using deep-learning-based 4D segmentation

Cell lineage consists of cell division timing, cell migration and cell fate, which are highly reproducible during the development of some nematode species, including C. elegans. Due to the lack of high spatiotemporal resolution of imaging technique and reliable shape-reconstruction algorithm, cell morphology have not been systematically characterized in depth over development for any metazoan. This significantly inhibits the study of space-related problems in developmental biology, including cell segregation, cell-cell contact and cell shape change over development. Here we develop an automated pipeline, CShaper, to help address these issues. By quantifying morphological parameters of densely packed cells in developing C. elegans emrbyo through segmentation of fluorescene-labelled membrance, we generate a time-lapse framework of cellular shape and migration for C. elegans embryos from 4-to 350-cell stage, including a full migration trajectory, morphological dynamics of 226 cells and 877 reproducible cell-cell contacts. In combination with automated cell tracing, cell-fate associated cell shape change becomes within reach. Our work provides a quantitative resource for C. elegans early development, which is expected to facilitate the research such as signaling transduction and cell biology of division.

developmental biology↗

Cell wall forming chitin synthases in a chytrid fungus

Chitin is a critical structural component of fungal cell walls, yet our understanding of its synthesis across the kingdom Fungi remains limited. Here, we investigate chitin synthase diversity, transcription and localisation in the saprotrophic chytrid Rhizoclosmatium globosum (Rg), expanding insights into fungal cell wall biology beyond Dikaryan models. We identified 20 chitin synthase genes in the Rg genome, including canonical Division I and II types, and a distinctive chitin synthase gene containing a glycoside hydrolase domain linked to {beta}-glucan synthesis. Transcriptomic analysis through zoospore, germling and immature thallus developmental stages revealed stage-specific expression patterns, with active gene diversity correlating with increasing morphological complexity. Using electroporation-based transformation and fluorescent fusion constructs, we demonstrated successful expression and localisation of two chitin synthases during cell development. Localisation patterns showed dynamic redistribution from cytoplasmic dispersion in early encysted cells to concentrated signals at the sporangium wall. Expression in the apophysis and at the apophysis-sporangium junction indicates the importance of these structures in cell maintenance. Our findings highlight functional specialisation among chitin synthases and underscore the importance of cell wall integrity in chytrid development. This work establishes Rg as a genetically tractable model for studying chytrid cell biology and contributes to broader understanding of fungal evolution and cell wall dynamics.

microbiology↗

OCellus: A Language-Model Framework for Single-Cell, Spatial, and Perturbation Biology with Natural-Language Reasoning

Computational modeling of cellular behavior--the virtual cell--has emerged as a stated grand challenge at the intersection of artificial intelligence and biology, yet existing foundation models remain specialized: single-cell models process dissociated transcriptomes only, spatial models require dedicated spatial-aware architectures, and perturbation predictors depend on manually curated knowledge bases that cap generalization. Here we introduce OCellus, a single nine-billion-parameter language model (Qwen3.5-9B) fine-tuned on twenty-two biological tasks that simultaneously addresses all three limitations through three coordinated technical contributions on a shared backbone. First, EvenClock encodes two-dimensional spatial coordinates as eighteen clockface sectors of text, enabling spatial reasoning on a vanilla language model without architectural modification; on ten spatial transcriptomics tasks OCellus attains 77 percent spatial-neighborhood accuracy, 96 percent spatial-cellchat accuracy, and 0.70 proportion-cosine similarity on spatial deconvolution, all without any spatial-aware architectural components. Second, per-gene language-model embeddings replace the Gene Ontology annotations that GEARS depends on, achieving Pearson correlation 0.945 on the Replogle 2022 perturbation benchmark versus 0.84 for GEARS across 457 completely unseen knockout genes. Third, OCellus-Agent provides a Planner-Router-Verifier natural-language interface that achieves 75 percent pipeline accuracy on eighty multi-task queries. Removing language-model embeddings collapses perturbation Pearson to 0.06, confirming that learned functional representations--not graph topology--drive the gain. As a cell-type encoder, OCellus ranks first among fourteen foundation models in linear-probe accuracy at 95.1 percent across four benchmark datasets, and reaches 72.6 percent average across twenty-two evaluated biological tasks--a 57-percentage-point absolute gain over the strongest baseline configuration. As a language model, OCellus uniquely generates natural-language explanations of its predictions, a capability absent from all competing methods. Code, pre-trained model weights, the graph-neural-network module, and the agent system will be made available upon publication.

bioinformatics↗

scINSIGHT for interpreting single-cell gene expression from biologically heterogeneous data

The increasing number of scRNA-seq data emphasizes the need for integrative analysis to interpret similarities and differences between single-cell samples. Even though different batch effect removal methods have been developed, none of the existing methods is suitable for het-erogeneous single-cell samples coming from multiple biological conditions. To address this challenge, we propose a method named scINSIGHT to learn coordinated gene expression patterns that are common among or specific to different biological conditions, offering a unique chance to identify cellular identities and key biological processes across single-cell samples. We have evaluated scINSIGHT in comparison with state-of-the-art methods using simulated and real data, which consistently demonstrate its improved performance. In addition, our results show the applicability of scINSIGHT in diverse biomedical and clinical problems.

bioinformatics↗

Mechanistic hierarchical population model identifies latent causes of cell-to-cell variability

All biological systems exhibit cell-to-cell variability, and this variability often has functional implications. To gain a thorough understanding of biological processes, the latent causes and underlying mechanisms of this variability must be elucidated. Cell populations comprising multiple distinct subpopulations are commonplace in biology, yet no current methods allow the sources of variability between and within individual subpopulations to be identified. This limits the analysis of single-cell data, for example provided by flow cytometry and microscopy. In this study, we present a data-driven modeling framework for the analysis of populations comprising heterogeneous subpopulations. Our approach combines mixture modeling with frameworks for distribution approximation, facilitating the integration of multiple single-cell datasets and the detection of causal differences between and within subpopulations. The computational efficiency of our framework allows hundreds of competing hypotheses to be compared, giving unprecedented depth of a study. We demonstrated the ability of our method to capture multiple levels of heterogeneity in the analyzes of simulated data and data from highly heterogeneous sensory neurons involved in pain initiation. Our approach identified the sources of cell-to-cell variability and revealed mechanisms that underlie the modulation of nerve growth factor-induced Erk1/2 signaling by extracellular scaffolds.

systems biology↗

Spatium: A Protein Language Foundation Model for Spatial Proteomics

Spatial proteomics provides single-cell protein measurements under highly constrained and heterogeneous protein panels across datasets, resulting in limited and partially overlapping measurement spaces for cellular characterization. Existing analyses predominantly rely on statistical or task-specific modeling, while learning scalable representations of spatial protein data remain underexplored. This gap motivates the need for models that can learn stable representations of cellular identity from constrained protein measurements. Here we introduce Spatium, a protein language foundation model trained on over 51 million cells across multiple spatial proteomics platforms. Spatium learns intrinsic co-expression hierarchies that capture cell identity in a manner robust to panel composition and measurement scale. Spatium builds a generalizable representation of cell states grounded in biologically interpretable protein expression patterns. We demonstrate that Spatium learns biologically meaningful cell representations across multiple downstream tasks. Spatium recovers accurate cell identities with marker expression patterns consistent with known biology and reveals functionally distinct spatial microenvironments characterized by coherent marker enrichment signatures. It further enables reconstruction of missing protein measurements while preserving biologically meaningful expression patterns. Across these analyses, Spatium demonstrates stable and interpretable performance with lightweight task-specific adaptation, highlighting the robustness of the learned representations across diverse biological and experimental contexts.

bioinformatics↗

Droplet sample preparation for single-cell proteomics applied to the cell cycle

Many biological processes, such as the cell division cycle, are reflected in protein covariation across single cells. This covariation can be quantified and interpreted by single-cell mass-spectrometry (MS) with sufficiently high throughput and accuracy. Towards this goal, we developed nPOP, a method that uses piezo acoustic dispensing to isolate individual cells in 300 picoliter volumes and performs all subsequent sample preparation steps in small droplets on a fluorocarbon-coated slide. This design enabled simultaneous sample preparation of thousands of single cells, including lysing, digesting, and labeling individual cells in volumes of 8-20 nl. Protein covariation analysis identified cell-cycle dynamics that were similar across cell types and dynamics that differed between cell types, even within sub-populations of melanoma cells defined by markers for drug-resistance priming. The melanoma cells expressing these markers accumulated in the G1 phase of the cell cycle, displayed distinct protein covariation across the cell cycle, accumulated glycogen, and had lower abundance of glycolytic enzymes. The non-primed melanoma cells exhibited gradients of protein abundance and covariation, suggesting transition states. These results were validated by different MS methods. Together, they demonstrate that protein covariation across single cells may reveal functionally concerted biological differences between closely related cell states.

bioengineering↗

Uncovering hidden biological processes by probabilistic filtering of single-cell data

Elucidating underlying biological processes in single-cell data is an ongoing challenge and the number of methods that recapitulate dominant signals in such data has increased significantly. However, cellular populations encode multiple biological attributes, related to their spatial configuration, temporal trajectories, cell-cell interactions, and responses to environmental cues, which may be overshadowed by the dominant signal and thus much harder to recover. To approach this task, we developed SiFT (SIgnal FilTering), a method for filtering biological signals in single-cell data, thus uncovering underlying processes of interest. Utilizing existing prior knowledge and reconstruction tools for a specific biological signal, such as spatial structure, SiFT filters the signal and uncovers additional biological attributes. SiFT is applicable to a wide range of tasks, from the removal of unwanted variation in the data as a pre-processing step to revealing hidden biological structures. Applied for pre-processing, SiFT outperforms state-of-the-art methods for the removal of nuisance signals and cell cycle effects. To recover underlying biological structure, we use existing prior knowledge regarding liver zonation to filter the spatial signal from single-cell liver data thereby enhancing the temporal circadian signal the cells are encoding. Lastly, we showcase the applicability of SiFT in the case-control setting for studying COVID-19 disease. Filtering the healthy signal, based on reference samples from healthy donors, exposes disease-related dynamics in COVID-19 data and highlights disease informative cells and their underlying disease response pathways.

bioinformatics↗

SNAP-tag and HaloTag fused proteins for HaSX8-inducible control over synthetic biological functions in engineered mammalian cells

Drug-inducible systems allow biological processes to be regulated through the administration of exogenous chemical inducers. Such methods can be used to study native biological activities, or to control synthetically engineered ones, with temporal and dose-dependent control. However, the number of existing drug-inducible systems is limited, and there remains a need for synthetic biology components that can be combined with the existing toolset and regulated with independent and orthogonal control. Here, we describe new cell engineering components that can be regulated via a heterodimerization of SNAP-tag and HaloTag domains using a selective small molecule crosslinker termed "HaXS8." The construction and validation of multiple HaXS8-sensitive components are described, including systems for regulating transcription, Cre recombinase activity, and caspase-9 activity in mammalian cells. The systems elaborate the ability to control gene expression, DNA recombination, and apoptosis in cell engineered systems.

synthetic biology↗

Scalable probe-based single-cell transcriptional profiling for virtual cell perturbation mapping and synthetic biology phenotyping

Large-scale single-cell transcriptional phenotyping of genetic perturbations (perturb-seq) links genes to phenotypes and should enable virtual cell predictive modeling and cellular engineering. However, current perturb-seq single-cell methods are costly, information sparse and require barcodes for many applications. We developed ProPer-seq, a perturb-seq method that uses multiplexed custom DNA probe panels to measure and phenotype synthetic biology perturbations at single-cell resolution without barcodes, including multidomain proteins and sgRNAs. ProPer-seq faithfully reproduces gold-standard perturb-seq phenotypes while achieving 4-fold cost reduction and 50% increased gene detection per cell. As a scalable fixed-cell profiling method, ProPer-seq enables atlas-scale profiling for virtual-cell initiatives and demonstrates data quality suitable for training and validating predictive models. Lastly, ProPer-seqs targeted detection of modular transgenes enables library-on-library perturbation profiling of combinatorial synthetic protein design spaces. We applied this to 3,550 sgRNA x dCas9 effector combinations as well as 260 CAR x ORF combinations dynamically profiled in primary T cells, revealing principles of transcriptional control and cell state modulation by multidomain synthetic transgenes.

synthetic biology↗

Clinically-Driven Design of Synthetic Gene Regulatory Programs in Human Cells

Synthetic biology seeks to enable the rational design of regulatory molecules and circuits to reprogram cellular behavior. The application of this approach to human cells could lead to powerful gene and cell-based therapies that provide transformative ways to combat complex diseases. To date, however, synthetic genetic circuits are challenging to implement in clinically-relevant cell types and their components often present translational incompatibilities, greatly limiting the feasibility, efficacy and safety of this approach. Here, using a clinically-driven design process, we developed a toolkit of programmable synthetic transcription regulators that feature a compact human protein-based design, enable precise genome-orthogonal regulation, and can be modulated by FDA-approved small molecules. We demonstrate the toolkit by engineering therapeutic human immune cells with genetic programs that enable titratable production of immunotherapeutics, drug-regulated control of tumor killing in vivo and in 3D spheroid models, and the first multi-channel synthetic switch for independent control of immunotherapeutic genes. Our work establishes a powerful platform for engineering custom gene expression programs in mammalian cells with the potential to accelerate clinical translation of synthetic systems.

synthetic biology↗

A Toolkit for Rapid Modular Construction of Biological Circuits in Mammalian Cells

The ability to rapidly assemble and prototype cellular circuits is vital for biological research and its applications in biotechnology and medicine. Current methods that permit the assembly of DNA circuits in mammalian cells are laborious, slow, expensive and mostly not permissive of rapid prototyping of constructs. Here we present the Mammalian ToolKit (MTK), a Golden Gate-based cloning toolkit for fast, reproducible and versatile assembly of large DNA vectors and their implementation in mammalian models. The MTK consists of a curated library of characterized, modular parts that can be easily mixed and matched to combinatorially assemble one transcriptional unit with different characteristics, or a hierarchy of transcriptional units weaved into complex circuits. MTK renders many cell engineering operations facile, as showcased by our ability to use the toolkit to generate single-integration landing pads, to create and deliver libraries of protein variants and sgRNAs, and to iterate through Cas9-based prototype circuits. As a biological proof of concept, we used the MTK to successfully design and rapidly construct in mammalian cells a challenging multicistronic circuit encoding the Ebola virus (EBOV) replication complex. This construct provides a non-infectious biosafety level 2 (BSL2) cellular assay for exploring the transcription and replication steps of the EBOV viral life cycle in its host. Its construction also demonstrates how the MTK can enable important and time sensitive applications such as the rapid testing of pharmacological inhibitors of emerging BSL4 viruses that pose a major threat to human health.

synthetic biology↗

ctQC improves biological inferences from single cell and spatial transcriptomics data

Quality control (QC) is the first critical step in single cell and spatial data analysis pipelines. QC is particularly important when analysing data from primary human samples, since genuine biological signals can be obscured by debris, perforated cells, cell doublets and ambient RNA released into the "soup" by cell lysis. Consequently, several QC methods for single cell data, employ fixed or data-driven quality thresholds. While these approaches efficiently remove empty droplets, they often retain low-quality cells. Here, we propose cell type-specific QC (ctQC), a stringent, data-driven QC approach that adapts to cell type differences and discards soup and debris. Evaluating single cell RNA-seq data from colorectal tumors, human spleen, and peripheral blood mononuclear cells, we demonstrate that ctQC outperforms existing methods by improving cell type separation in downstream clustering, suppressing cell stress signatures, revealing patient-specific cell states, eliminating artefactual clusters and reducing ambient RNA artifacts. When applied to sequencing-based spatial RNA profiling data (Slide-seq), ctQC improved spatial coherence of cell clusters and consistency with anatomical structures. These results demonstrate that strict, data-driven, cell-type-specific QC is applicable to diverse sample types and substantially improves the quality and reliability of biological inferences from single cell and spatial RNA profiles.

bioinformatics↗

The Profiling of Bisecting N-acetylglucosamine (GlcNAc) Modification in Human Amniotic Membrane by Glycomic and Glycoproteomic Analyses

It is acknowledged that the bisecting N-acetylglucosamine (GlcNAc) structure, a GlcNAc linked to the core {beta}-mannose residue via a {beta}1,4 linkage, represents a special type of N-glycosylated modification and has been reported to be involved in various biological processes, such as cell adhesion and fetal development. Clark et al. has found that the majority of N-glycans in human trophoblasts bearing a bisecting GlcNAc. This type of glycan has been reported to help trophoblasts get resistant to natural killer (NK) cell-mediated cytotoxicity, and this would provide a possible explanation for the question how could the mother nourish a fetus within herself without rejection. Herein, we hypothesized that human amniotic membrane which is the last barrier for the fetus may also express bisecting type glycans to protect the fetus. To test this hypothesis, glycomic analysis of human amniotic membrane was performed, and the bisecting N-glycans with high abundance were detected. In addition, we re-analyzed our proteomic data with high fractionation and amino acid sequence coverage from human amniotic membrane, which had been released for the exploration of human missing proteins. The presence of bisecting GlcNAc peptides was revealed and confirmed. A total of 41 glycoproteins with 43 glycopeptides were found to possess a bisecting GlcNAc, 25 of which are for the first time to be reported to have this type of modification. These results provide the profiling of bisecting GlcNAc modification in human amniotic membrane and benefit to the function studies of glycoproteins with bisecting GlcNAc modification and the function studies in immune suppression of human placenta. The mass spectrometry placenta data are available via ProteomeXchange (PXD010630).

cell biology↗

Deep Proteome Profiling of Human Mammary Epithelia at Lineage and Age Resolution

Age is the major risk factor in most carcinomas, yet little is known about how proteomes change with age in any human epithelium. We present comprehensive proteomes comprised of >9,000 total proteins, and >15,000 phosphopeptides, from normal primary human mammary epithelia at lineage resolution from ten women ranging in age from 19 to 68. Data were quality controlled, and results were biologically validated with cell-based assays. Age-dependent protein signatures were identified using differential expression analyses and weighted protein co-expression network analyses. Up-regulation of basal markers in luminal cells, including KRT14 and AXL, were a prominent consequence of aging. PEAK1 was identified as an age-dependent signaling kinase in luminal cells, which revealed a potential age-dependent vulnerability for targeted ablation. Correlation analyses between transcriptome and proteome revealed age-associated loss of proteostasis regulation. Protein expression and phosphorylation changes in the aging breast epithelium identify potential therapeutic targets for reducing breast cancer susceptibility.

cell biology↗

Optogenetic relaxation of actomyosin contractility uncovers mechanistic roles of cortical tension during cytokinesis

Actomyosin contractility generated cooperatively by nonmuscle myosin II and actin filaments plays essential roles in a wide range of biological processes, such as cell motility, cytokinesis, and tissue morphogenesis. However, it is still unknown how actomyosin contractility generates force and maintains cellular morphology. Here, we demonstrate an optogenetic method to induce relaxation of actomyosin contractility. The system, named OptoMYPT, combines a catalytic subunit of the type I phosphatase-binding domain of MYPT1 with an optogenetic dimerizer, so that it allows light-dependent recruitment of endogenous PP1c to the plasma membrane. Blue-light illumination was sufficient to induce dephosphorylation of myosin regulatory light chains and decrease in traction force at the subcellular level. The OptoMYPT system was further employed to understand the mechanics of actomyosin-based cortical tension and contractile ring tension during cytokinesis. We found that the relaxation of cortical tension at both poles by OptoMYPT accelerated the furrow ingression rate, revealing that the cortical tension substantially antagonizes constriction of the cleavage furrow. Based on these results, the OptoMYPT system will provide new opportunities to understand cellular and tissue mechanics.

cell biology↗

Conditional immobilization for live imaging C. elegans using auxin-dependent protein depletion

The visualization of biological processes using fluorescent proteins and dyes in living organisms has enabled numerous scientific discoveries. The nematode Caenorhabditis elegans is a widely used model organism for live imaging studies since the transparent nature of the worm enables imaging of nearly all tissues within a whole, intact animal. While current techniques are optimized to enable the immobilization of hermaphrodite worms for live imaging, many of these approaches fail to successfully restrain the smaller male worms. To enable live imaging of worms of both sexes, we developed a new genetic, conditional immobilization tool that uses the auxin inducible degron (AID) system to immobilize both hermaphrodites and male worms for live imaging. Based on chromosome location, mutant phenotype, and predicted germline consequence, we identified and AID-tagged three candidate genes (unc-18, unc-104, and unc-52). Strains with these AID-tagged genes were placed on auxin and tested for mobility and germline defects. Among the candidate genes, auxin-mediated depletion of UNC-18 caused significant immobilization of both hermaphrodite and male worms that was also partially reversible upon removal from auxin. Notably, we found that male worms require a higher concentration of auxin for a similar amount of immobilization as hermaphrodites, thereby suggesting a potential sex-specific difference in auxin absorption and/or processing. In both males and hermaphrodites, depletion of UNC-18 did not largely alter fertility, germline progression, nor meiotic recombination. Finally, we demonstrate that this new genetic tool can successfully immobilize both sexes enabling live imaging studies of sexually dimorphic features in C. elegans. ARTICLE SUMMARYC. elegans is a powerful model system for visualizing biological processes in live cells. In addition to the challenge of suppressing the worm movement for live imaging, most immobilization techniques only work with hermaphrodites. Here, we describe a new genetic immobilization tool that conditionally immobilizes both worm sexes for live imaging studies. Additionally, we demonstrate that this tool can be used for live imaging the C. elegans germline without causing large defects to germline progression or fertility in either sex.

cell biology↗

Ultrastructural and functional analysis of extra-axonemal structures in trichomonads

Trichomonas vaginalis and Tritrichomonas foetus are extracellular flagellated parasites that inhabit humans and other mammals, respectively. In addition to motility, flagella act in a variety of biological processes in different cell types; and extra-axonemal structures (EASs) has been described as fibrillar structures that provide mechanical support and act as metabolic, homeostatic and sensory platforms in many organisms. Here, we identified the presence of EASs forming prominent flagellar swellings in T. vaginalis and T. foetus and we observed that their formation was associated with the parasites adhesion on the host cells, fibronectin, and precationized surfaces; and parasite:parasite interaction. A high number of rosettes, clusters of intramembrane particles that has been proposed as sensorial structures, and microvesicles protruding from the membrane were observed in the EASs. The protein VPS32, a member of the ESCRT-III complex crucial for diverse membrane remodeling events, the pinching off and release of microvesicles, was found in the surface as well as in microvesicles protruding from EASs. Moreover, we demonstrated that overexpression of VPS32 protein induce EAS formation and increase parasite motility in semi-solid medium. These results provide valuable data about the role of the flagellar EASs in the cell-to-cell communication and pathogenesis of these extracellular parasites.

cell biology↗