bioRxiv Science⌕ Search

Biology subjects

Bayard, Q.

Publications and source records attributed to Bayard, Q..

4 recordsLinked to original sources

Multi-modal benchmarking of the Ultima UG100 and Illumina NovaSeq sequencing platforms using clinically relevant FFPE tissues

Emerging high-throughput sequencing technologies promise lower costs and higher scalability, yet their performance on archival clinical samples remains poorly characterized. Here, we benchmarked Ultima Genomics UG100 against Illumina Novaseq platforms across single-nuclei RNA-seq (snRNA-seq), whole-transcriptome (WTS), whole-exome (WES), and whole-genome sequencing (WGS) using FFPE tissues from oncologic and immune-mediated diseases. Across matched samples, we systematically assessed data quality, coverage profiles, error spectra, variant concordance and transcriptomic reproducibility. UG100 produced highly comparable results to Illumina, capturing key oncogenic and immune-related transcripts, accurately resolving cellular composition in snRNA-seq, and maintaining sensitivity for lowly expressed genes, despite characteristic insertion-biased indels and modest differences in multi-mapping reads. Discrepancies were subtle, largely limited to pseudogene and non-coding transcripts, and did not affect pathway-level conclusions. Ultima UG100 platform prioritized high precision and reduced low-frequency artifacts, offering a cleaner but more conservative variant-calling profile compared to the more sensitive, yet noisier, Illumina/DRAGEN workflow. This multimodal, clinically oriented assessment provides the first comprehensive evaluation of UG100, demonstrating its translational utility in population-scale genomics, and highlighting the potential for emerging sequencing technologies to lower the cost of biomedical research and clinical diagnostics.

genomics↗

Deciphering Cellular Ecosystems Driving Tumor Progression and Immune Escape from Spatial Transcriptomics and Single-Cell with COMPOTES

Cell-cell communication is central to understanding the complex interactions within the tumor microenvironment. However, current methods fail to identify recurrent communication patterns across patient cohorts from spatial transcriptomics, as they are often limited to single samples or lack essential spatial context. Yet this is essential for understanding how local environments influence cell phenotype and states, and shape the entire cellular ecosystem. We introduce a machine-learning approach that models local, spatially aware ligand-receptor interactions and uses matrix factorization to extract global multicellular programs from large cohorts representing the complex biology of cancer. Applied to a multimodal muscle-invasive bladder cancer cohort of 146 patients, it uncovered 45 communication programs defined by distinct ligand-receptor pairs and cellular niches. In particular, we identified a conserved immune program linked to stalled anti-tumor immunity and a program linking KMT2D loss-of-function mutations with early-stage (T2) tumors, intense proliferation and a favorable response to neoadjuvant chemotherapy.

bioinformatics↗

Multiple instance learning with spatial transcriptomics for interpretable patient-level predictions: application in glioblastoma

Accurate prediction of patient outcomes remains a major challenge in oncology. While recent machine learning (ML) approaches often rely on bulk omics lacking spatial resolution or histology-based multiple instance learning (MIL), spatial transcriptomics (SpT) provides a unique opportunity to capture both molecular content and tissue architecture. However, no generalizable ML framework has yet been established to exploit SpT for patient-level outcomes. We present SpaMIL, a flexible and interpretable MIL framework designed for SpT, with a distillation strategy that enables deployment for hematoxylin and eosin (H&E) slides alone. We evaluate the framework by predicting survival from glioblastoma (GBM) patients, a clinically compelling setting given its aggressiveness with a median survival of only 15 months and the lack of prognostic clinical variables. We analyzed 76 GBM cases from the MOSAIC dataset: 43 with matched SpT, H&E, single-nucleus RNA-seq (scRNA-seq), bulk RNA-seq, and clinical variables, and 33 with H&E for external validation. We developed two main architectures: abMIL, tailored to SpTs spatial molecular structure, and MabMIL, which distills SpT-derived representations into H&E. Model interpretability was achieved through a Shapley-based framework linking prognostic predictions to cell-type compositions via SpT deconvolution. In benchmarking across the five GBM MOSAIC modalities, SpT-based abMIL achieved unprecedented prognostic accuracy (median C-index: 0.72, standard deviation: 0.04), outperforming all other modalities, including established clinical predictors. PCA and deconvolution-based SpT representations surpassed recent foundation models, suggesting the need for further research on SpT foundation models. Our interpretability analysis highlighted malignant and non-malignant cell subpopulations associated with favorable or poor prognosis, consistent with recent reports. Finally, MabMIL maintained strong performance while enabling H&E-only deployment, with improved condorance index over H&E-only baselines in both internal (0.59 vs. 0.57) and external (0.62 vs. 0.55) cohorts.

cancer biology↗

Transcriptome Analysis of Archived Tumor Tissues by Visium, GeoMx DSP, and Chromium Methods Reveals Inter- and Intra-Patient Heterogeneity

Recent advancements in probe-based, full-transcriptome, high-resolution technologies for Formalin-Fixed Paraffin-Embedded (FFPE) tissues, such as Visium CytAssist, Chromium Flex (10X Genomics), and GeoMx DSP (Nanostring), have opened new opportunities for studying decades-old archival samples in biobanks, facilitating the generation of data from extensive cohorts. However, the experimental protocols can be labor-intensive and costly; therefore, it is thus essential for researchers to carefully evaluate the strengths and limitations of each technology in relation to their specific research objectives. Here, we report the results of a comparative analysis of the three methods mentioned above on FFPE archival tumor samples from four non-small cell lung cancer, four breast cancer and six diffuse large B-cell lymphoma. We highlight some relative advantages and disadvantages of each method in the context of operational challenges, bioinformatic analysis and biological discovery. Our results show that: 1) all three methods yielded good-quality, highly reproducible transcriptomic data from serial sections of the same FFPE block; 2) GeoMx data contained mixtures of cell types, even when pre-selecting areas with cell type-specific markers; 3) high-throughput spot-level (Visium) or cell-level (Chromium) data enabled the identification of tumor heterogeneity within and between patients, which could be used to identify targeted therapies. Our data support the use of Visium and Chromium for high-throughput and discovery-driven projects, while the GeoMx platform could be suited for addressing specialized questions on targeted regions. All data generated from this study, including GeoMx, Visium, Chromium, H&E, and expert annotations are publicly available.

bioinformatics↗