bioRxiv Science⌕ Search

Biology subjects

Sims, Z.

Publications and source records attributed to Sims, Z..

5 recordsLinked to original sources

miniMTI: minimal multiplex tissue imaging enhances biomarker expression prediction from histology

Virtual multiplexing from routine histology has advanced rapidly, yet morphology alone provides limited access to molecular state, imposing an intrinsic ceiling on H&E-only inference. Here, we introduce miniMTI, a molecularly anchored virtual staining framework that determines the minimal set of experimentally measured markers required, alongside H&E, to accurately reconstruct large multiplex tissue imaging (MTI) panels while preserving biologically and clinically relevant information. miniMTI learns from paired same-section H&E-MTI data using a unified multimodal generative model that can condition on arbitrary combinations of measured marker channels, coupled with an iterative panel selection strategy to rank informative molecular anchors. Across colorectal and prostate cancer cohorts spanning two MTI platforms and over 40 million cells, miniMTI reduces a 40-marker MTI assay to H&E plus as few as three measured molecular markers, while accurately recovering withheld markers, preserving cellular phenotypes and spatial tissue architecture, and disease-associated molecular programs, including Gleason grade-linked signatures. By integrating histology context with sparse molecular grounding, miniMTI overcomes the limitations of morphology-only virtual staining and provides a scalable, cost-effective approach for expanding MTI-level biomarker coverage with retained biological interpretability and clinical relevance.

bioinformatics↗

Language of Stains: Tokenization Enhances Multiplex Immunofluorescence and Histology Image Synthesis

Multiplex tissue imaging (MTI) is a powerful tool in cancer research, allowing spatially resolved, single-cell phenotype analysis. However, MTI platforms face challenges such as high costs, tissue loss, lengthy acquisition times, and complex analysis of large, multichannel images with batch effects. To address these challenges, we propose a novel computational method to model the interactions between dozens of panel markers and Hematoxylin & Eosin (H&E) staining, enabling in-silico generation of marker stains. This approach reduces the reliance on experimentally measured markers, bridging low-cost H&E data with MTIs high-content information. Our approach uses a two-stage frame-work for channel-wise bioimage synthesis: first, vector quantization learns a visual token vocabulary, then a bidirectional transformer infers missing markers through masked language modeling. Comprehensive bench-marking across different MTI platforms and tissue types demonstrates the effectiveness of our method in improving marker prediction while maintaining biological relevance. This advance makes high-dimensional multiplex tissue imaging more accessible and scalable, supporting deeper insights and potential clinical applications in cancer research.

bioinformatics↗

UniFORM: Towards Universal Immunofluorescence Normalization for Multiplex Tissue Imaging

Multiplexed tissue imaging (MTI) technologies enable high-dimensional spatial analysis of tumor microenvironments but face challenges with technical variability in staining intensities. Existing normalization methods, including Z-score, ComBat, and MxNorm, often fail to account for the heterogeneous, right-skewed expression patterns of MTI data, compromising signal alignment and downstream analyses. We present UniFORM, a non-parametric, Python-based pipeline that uses an automated rigid landmark functional data registration approach for normalizing both feature- and pixel-level MTI data. Designed specifically for the distributional characteristics of MTI datasets, UniFORM operates without prior distributional assumptions and performs robustly regardless of distribution modality, including both unimodal and bimodal patterns. It removes technical variation by aligning the biologically invariant component of the signal, typically the negative (non-expressing) population, while preserving biologically meaningful variation in the positive population, thereby maintaining tissue-specific expression patterns essential for downstream analysis. Benchmarking across three distinct MTI platform datasets demonstrates that UniFORM outperforms existing methods in mitigating batch effects while maintaining biological signal fidelity. This is evidenced by improved marker distribution alignment and positive population preservation, enhanced kBET and Silhouette scores, and improved downstream analyses such as UMAP visualizations and Leiden clustering. UniFORM also introduces a novel guided fine-tuning option for complex and heterogeneous datasets. Although optimized for fluorescence-based platforms, UniFORM provides a scalable and robust solution for MTI data normalization, enabling accurate and biologically meaningful interpretations.

bioinformatics↗

A Masked Image Modeling Approach to Cyclic Immunofluorescence (CyCIF) Panel Reduction and Marker Imputation

CyCIF quantifies multiple biomarkers, but panel capacity is compromised by technical challenges including tissue loss. We propose a computational panel reduction, inferring surrogate CyCIF data from a subset of biomarkers. Our model reconstructs the information content from 25 markers using only 9 markers, learning co-expression and morphological patterns. We demonstrate strong correlations in predictions and generalizability across breast and colorectal cancer tissue microarrays, illustrating broader applicability to diverse tissue types.

bioinformatics↗

SEG: Segmentation Evaluation in absence of Ground truth labels

Identifying individual cells or nuclei is often the first step in the analysis of multiplex tissue imaging (MTI) data. Recent efforts to produce plug-and-play, end-to-end MTI analysis tools such as MCMICRO1- though groundbreaking in their usability and extensibility - are often unable to provide users guidance regarding the most appropriate models for their segmentation task among an endless proliferation of novel segmentation methods. Unfortunately, evaluating segmentation results on a users dataset without ground truth labels is either purely subjective or eventually amounts to the task of performing the original, time-intensive annotation. As a consequence, researchers rely on models pre-trained on other large datasets for their unique tasks. Here, we propose a methodological approach for evaluating MTI nuclei segmentation methods in absence of ground truth labels by scoring relatively to a larger ensemble of segmentations. To avoid potential sensitivity to collective bias from the ensemble approach, we refine the ensemble via weighted average across segmentation methods, which we derive from a systematic model ablation study. First, we demonstrate a proof-of-concept and the feasibility of the proposed approach to evaluate segmentation performance in a small dataset with ground truth annotation. To validate the ensemble and demonstrate the importance of our method-specific weighting, we compare the ensembles detection and pixel-level predictions - derived without supervision - with the datas ground truth labels. Second, we apply the methodology to an unlabeled larger tissue microarray (TMA) dataset, which includes a diverse set of breast cancer phenotypes, and provides decision guidelines for the general user to more easily choose the most suitable segmentation methods for their own dataset by systematically evaluating the performance of individual segmentation approaches in the entire dataset.

bioinformatics↗