bioRxiv Science⌕ Search

Biology subjects

Egidi, L.

Publications and source records attributed to Egidi, L..

3 recordsLinked to original sources

Tumour evolution as ground truth for cancer whole-genome sequencing

Cancer genomes are shaped by evolutionary processes that couple mutagenesis, clonal selection, chromosomal instability, spatial growth and treatment response into structured genomic patterns, yet current benchmarking strategies largely ignore this evolutionary dependency. Here, we present SCOUT, a large-scale synthetic whole-genome sequencing resource of over 200 samples, designed for systematic benchmarking of tumour genomic analysis and evolutionary inference under controlled evolutionary ground truth. Unlike conventional task-specific simulations, SCOUT models tumour evolution as a latent generative process that simultaneously shapes mutations, copy-number alterations, variant allele frequencies, mutational signatures and clonal architectures. SCOUT recapitulates key features of solid and haematological malignancies, including driver mutations, chromosomal instability, intratumour heterogeneity, spatial sampling and treatment-associated evolutionary dynamics in tumour and matched-normal longitudinal and multi-region sequencing designs. Using SCOUT, we benchmarked widely used methods for somatic variant detection, copy-number analysis, mutational signature inference and tumour evolutionary reconstruction. Across analytical tasks, performance deteriorated in low-purity, highly subclonal and structurally complex tumours, while spatial sampling bias and hypermutation generated spurious evolutionary signals that confounded tumour interpretation across multiple inference layers. Evolutionary simulations further distinguished lineage-restricted genetic bottlenecks from multi-lineage resistance dynamics associated with tumour plasticity. Tumour purity consistently exerted a stronger effect on inference accuracy than sequencing depth. Together, our results establish evolutionary ground truth as a prerequisite for reproducible benchmarking and biologically interpretable analysis of cancer whole-genome sequencing data.

bioinformatics↗

Scalable, fast and accurate differential gene expression testing from millions of cells of multiple patients

Since the development of DNA microarrays and later RNA bulk sequencing, testing with statistically independent samples has been the standard method for detecting genes with different transcription patterns. Single-cell assays challenge these assumptions because individual cells are statistically dependent, and all proposed methodologies present mathematical limitations or computational bottlenecks that prevent a seamless integration of data from many cells and patients simultaneously. In this work, we solve this crucial limitation by introducing a Bayesian framework that retrieves the independence structure at the level of individual patients, separating differences across individuals from actual transcriptional differences. Leveraging multi-GPU and variational inference, our approach excels across different experimental designs and scales to analyse over 10 million cells. This framework enables single-cell differential expression analysis that can finally integrate datasets from large clinical cohorts, atlas projects, or drug-response screens with thousands of samples and millions of cells.

bioinformatics↗

Model-based Bayesian inference of cancer dynamics from heterogenous longitudinal data

The kinetic parameters of cancer population dynamics are critical for developing reliable predictors of tumour growth patterns, extracting metrics for patient stratification and creating algorithms that can forecast clinically significant events. Here, we introduce a model-based Bayesian framework that leverages longitudinal phenotypic (e.g., tumour volume, cell counts) or genotypic (e.g., mutation frequency) data to infer critical parameters of tumour progression within a single patient. Our models uses population genetics to estimate probability distributions for tumour growth rates, initiation and extinction times, pinpointing abrupt shifts in tumour dynamics due to treatment response and revealing associations between drug resistance and pre-existing cancer cell populations. We apply our framework to address pivotal clinical questions across three major cancer types. In colorectal cancer, we use tumour markers data to identify extensive pre-existing RAS-linked resistance to cetuximab. In lung cancer, we use somatic mutation frequencies in citculating tumour DNA to determine prognostic growth rates and develop a test for monitoring minimal residual disease. In chronic lymphocytic leukaemia, we use white blood cell counts to stratify patients by growth patterns and predict time to treatment, advancing adaptive monitoring strategies.

cancer biology↗