bioRxiv Science⌕ Search

Biology subjects

Espin-Perez, A.

Publications and source records attributed to Espin-Perez, A..

3 recordsLinked to original sources

Multi-modal benchmarking of the Ultima UG100 and Illumina NovaSeq sequencing platforms using clinically relevant FFPE tissues

Emerging high-throughput sequencing technologies promise lower costs and higher scalability, yet their performance on archival clinical samples remains poorly characterized. Here, we benchmarked Ultima Genomics UG100 against Illumina Novaseq platforms across single-nuclei RNA-seq (snRNA-seq), whole-transcriptome (WTS), whole-exome (WES), and whole-genome sequencing (WGS) using FFPE tissues from oncologic and immune-mediated diseases. Across matched samples, we systematically assessed data quality, coverage profiles, error spectra, variant concordance and transcriptomic reproducibility. UG100 produced highly comparable results to Illumina, capturing key oncogenic and immune-related transcripts, accurately resolving cellular composition in snRNA-seq, and maintaining sensitivity for lowly expressed genes, despite characteristic insertion-biased indels and modest differences in multi-mapping reads. Discrepancies were subtle, largely limited to pseudogene and non-coding transcripts, and did not affect pathway-level conclusions. Ultima UG100 platform prioritized high precision and reduced low-frequency artifacts, offering a cleaner but more conservative variant-calling profile compared to the more sensitive, yet noisier, Illumina/DRAGEN workflow. This multimodal, clinically oriented assessment provides the first comprehensive evaluation of UG100, demonstrating its translational utility in population-scale genomics, and highlighting the potential for emerging sequencing technologies to lower the cost of biomedical research and clinical diagnostics.

genomics↗

Histology and spatial transcriptomic integration revealed infiltration zone with specific cell composition as a prognostic hotspot in glioblastoma

BackgroundGlioblastoma (GBM), the most aggressive primary brain tumor, has a median survival of approximately 15 months. Twenty percent of patients survive beyond three years, but known clinical factors like age, performance status, resection extent, and MGMT promoter methylation status do not fully explain the observed outcomes. ObjectiveOur objective was to identify novel histology derived biomarkers associated with end-of-spectrum overall survival (OS) to provide novel biological insight with a translational potential. MethodsWe analyzed a total of 748 GBM patients from 3 different cohorts, uniquely enriched in long survivors (n=98 with overall survival (OS) > 5y including n=196 with OS[&ge;]3y), with clinical data and H&E slides obtained from the primary tumor at baseline. We propose an interpretable machine learning (ML) methodology for the discovery of histological biomarkers. Our method learned to segment each H&E slide into three distinct regions associated with long-term survival, short-term survival, and non-informative tissue. We characterized these regions by integrating unsupervised learning, nuclei segmentation, blood vessels detection, pathologist annotations, and multimodal data including spatial transcriptomics from n=31 patients of the GBM MOSAIC dataset to discover fully interpretable biomarkers. ResultsOur OS prediction model using histology and clinical data as input achieved an area under the curve (AUC) of 0.85 for the classification of patients between OS<2 and OS[&ge;]3y in external cohort validation, outperforming significantly models trained on clinical data or on histology alone (AUC of 0.76; 0.73, respectively). Two novel biomarkers were predicting poor survival: the presence of regions of lowly infiltrated white matter enriched in malignant cells with a mesenchymal-like phenotype, and lower levels of angiogenesis associated with higher hypoxia response in the main tumor regions. We also found that a subtype of immunosuppressive tumor macrophages - defined by high PLIN2 expression and lipid accumulation- is consistently enriched in histological areas predictive of poor prognosis. ConclusionOur interpretable ML methodology identified a novel prognostic impact of biological processes and cell types according to distinct tumor regions of GBM. These results pave the way for spatially-informed biomarkers to improve risk stratification and for personalized spatially-targeted therapeutic strategies. Key highlightsO_LIOur ML model identified histological biomarkers predicting prognosis independently from known clinical factors C_LIO_LIThe region of lowly infiltrated white matter enriched in malignant cells including a mesenchymal-like phenotype is predictive of poor prognosis C_LIO_LIAngiogenesis is increased in areas predictive of long survival in main non-necrotic tumor regions. C_LIO_LIThe subtype of macrophages expressing PLIN2 and associated with increased lipid metabolism was associated with poor prognosis in all GBM regions. C_LI Highlights O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=116 SRC="FIGDIR/small/681087v4_ufig1.gif" ALT="Figure 1"> View larger version (43K): org.highwire.dtl.DTLVardef@45e619org.highwire.dtl.DTLVardef@105a4b0org.highwire.dtl.DTLVardef@17f257dorg.highwire.dtl.DTLVardef@7669b4_HPS_FORMAT_FIGEXP M_FIG C_FIG

cancer biology↗

Joint probabilistic modeling of pseudobulk and single-cell transcriptomics enables accurate estimation of cell type composition

Bulk RNA sequencing provides an averaged gene expression profile of the numerous cells in a tissue sample, obscuring critical information about cellular heterogeneity. Computational deconvolution methods can estimate cell type proportions in bulk samples, but current approaches can lack precision in key scenarios due to simplistic statistical assumptions, limited modeling of cell-type heterogeneity and poor handling of rare populations. We present MixupVI, a deep generative model that learns representations of single-cell transcriptomic data and introduces a mixup-based regularization to enable reference-free deconvolution of bulk samples. Our method creates a latent representation with an additive property, where the representation of a pseudobulk sample corresponds to the weighted sum of its constituent cell types. We demonstrate how MixupVI enables accurate estimation of cell type proportions through benchmarking on pseudobulks simulated from a large immune single-cell atlas. To support reproducibility and foster progress in the field, we also release PyDeconv, a Python library that implements multiple state-of-the-art deconvolution algorithms and provides a comprehensive benchmark on simulated pseudobulk datasets.

bioinformatics↗