bioRxiv ScienceSearch

Biology subjects

Chkhaidze, K.

Publications and source records attributed to Chkhaidze, K..

3 recordsLinked to original sources

Model-based tumor subclonal reconstruction

The vast majority of cancer next-generation sequencing data consist of bulk samples composed of mixtures of cancer and normal cells. To study tumor evolution, subclonal reconstruction approaches based on machine learning are used to separate subpopulation of cancer cells and reconstruct their ancestral relationships. However, current approaches are entirely data-driven and agnostic to evolutionary theory. We demonstrate that systematic errors occur in subclonal reconstruction if tumor evolution is not accounted for, and that those errors increase when multiple samples are taken from the same tumor. To address this issue, we present a novel approach for model-based subclonal reconstruction that combines data-driven machine learning with evolutionary theory. Using public, synthetic and newly generated data, we show the method is more robust and accurate than current techniques in both single-sample and multi-region sequencing data. With careful data curation and interpretation, we show how the method allows minimizing the confounding factors that affect non-evolutionary methods, leading to a more accurate recovery of the evolutionary history of human tumors.

bioinformatics

Measuring single cell divisions in human cancers from multi-region sequencing data

Cancer is driven by complex evolutionary dynamics involving billions of cells. Increasing effort has been dedicated to sequence single tumour cells, but obtaining robust measurements remains challenging. Here we show that multi-region sequencing of bulk tumour samples contains quantitative information on single-cell divisions that is accessible if combined with evolutionary theory. Using high-throughput data from 16 human cancers, we measured the in vivo per-cell point mutation rate (mean: 1.69x10-8 bp per cell division) and per-cell survival rate (mean: 0.57) in individual patient tumours from colon, lung and renal cancers. Per-cell mutation rates varied 50-fold between individuals, and per-cell survival rates were between nearly-homeostatic and almost perfect cell doublings, equating to tumour ages between 1 and 19 years. Furthermore, reanalysing a recent dataset of 89 whole-genome sequenced healthy haematopoietic stem cells, we find 1.14 mutations per genome per cell division and near perfect cell doublings (per-cell survival rate: 0.96) during early haematopoietic development. Our analysis measures in vivo the most fundamental properties of human cancer and healthy somatic evolution at single-cell resolution within single individuals.

cancer biology

Spatially constrained tumour growth affects the patterns of clonal selection and neutral drift in cancer genomic data

Quantification of the effect of spatial tumour sampling on the patterns of mutations detected in next-generation sequencing data is largely lacking. Here we use a spatial stochastic cellular automaton model of tumour growth that accounts for somatic mutations, selection, drift and spatial constrains, to simulate multi-region sequencing data derived from spatial sampling of a neoplasm. We show that the spatial structure of a solid cancer has a major impact on the detection of clonal selection and genetic drift from bulk sequencing data and single-cell sequencing data. Our results indicate that spatial constrains can introduce significant sampling biases when performing multi-region bulk sampling and that such bias becomes a major confounding factor for the measurement of the evolutionary dynamics of human tumours. We present a statistical inference framework that takes into account the spatial effects of a growing tumour and allows inferring the evolutionary dynamics from patient genomic data. Our analysis shows that measuring cancer evolution using next-generation sequencing while accounting for the numerous confounding factors requires a mechanistic model-based approach that captures the sources of noise in the data. SummarySequencing the DNA of cancer cells from human tumours has become one of the main tools to study cancer biology. However, sequencing data are complex and often difficult to interpret. In particular, the way in which the tissue is sampled and the data are collected, impact the interpretation of the results significantly. We argue that understanding cancer genomic data requires mathematical models and computer simulations that tell us what we expect the data to look like, with the aim of understanding the impact of confounding factors and biases in the data generation step. In this study, we develop a spatial simulation of tumour growth that also simulates the data generation process, and demonstrate that biases in the sampling step and current technological limitations severely impact the interpretation of the results. We then provide a statistical framework that can be used to overcome these biases and more robustly measure aspects of the biology of tumours from the data.

bioinformatics