bioRxiv ScienceSearch

Biology subjects

Yuan, K.

Publications and source records attributed to Yuan, K..

7 recordsLinked to original sources

Portraits of genetic intra-tumour heterogeneity and subclonal selection across cancer types

Intra-tumor heterogeneity (ITH) is a mechanism of therapeutic resistance and therefore an important clinical challenge. However, the extent, origin and drivers of ITH across cancer types are poorly understood. To address this question, we extensively characterize ITH across whole-genome sequences of 2,658 cancer samples, spanning 38 cancer types. Nearly all informative samples (95.1%) contain evidence of distinct subclonal expansions, with frequent branching relationships between subclones. We observe positive selection of subclonal driver mutations across most cancer types, and identify cancer type specific subclonal patterns of driver gene mutations, fusions, structural variants and copy-number alterations, as well as dynamic changes in mutational processes between subclonal expansions. Our results underline the importance of ITH and its drivers in tumor evolution, and provide an unprecedented pan-cancer resource of comprehensively annotated subclonal events from whole-genome sequencing data.

cancer biology

SVclone: inferring structural variant cancer cell fraction

We present SVclone, a computational method for inferring the cancer cell fraction of structural variant breakpoints from whole-genome sequencing data. We validate our approach using simulated and real tumour samples, and demonstrate its utility on 2,778 whole-genome sequenced tumours. We find a subset of liver, breast and ovarian cancer cases with decreased overall survival that have subclonally enriched copy-number neutral rearrangements, an observation that could not be discovered with currently available methods.

cancer biology

The evolutionary history of 2,658 cancers

Cancer develops through a process of somatic evolution. Here, we use whole-genome sequencing of 2,778 tumour samples from 2,658 donors to reconstruct the life history, evolution of mutational processes, and driver mutation sequences of 39 cancer types. The early phases of oncogenesis are driven by point mutations in a small set of driver genes, often including biallelic inactivation of tumour suppressors. Early oncogenesis is also characterised by specific copy number gains, such as trisomy 7 in glioblastoma or isochromosome 17q in medulloblastoma. By contrast, increased genomic instability, a nearly four-fold diversification of driver genes, and an acceleration of point mutation processes are features of later stages. Copy-number alterations often occur in mitotic crises leading to simultaneous gains of multiple chromosomal segments. Timing analysis suggests that driver mutations often precede diagnosis by many years, and in some cases decades, providing a window of opportunity for early cancer detection.

cancer biology

Interphase-Arrested Drosophila Embryos Initiate Mid-Blastula Transition At A Low Nuclear-Cytoplasmic Ratio

Externally deposited eggs begin development with an immense cytoplasm and a single overwhelmed nucleus. Rapid mitotic cycles restore normality as the ratio of nuclei to cytoplasm (N/C) increases. At the 14th cell cycle in Drosophila embryos, the cell cycle slows, transcription increases, and morphogenesis begins at the Mid-Blastula Transition (MBT). To explore the role of N/C in MBT timing, we blocked N/C-increase by downregulating cyclin/Cdk1 to arrest early cell cycles. Embryos arrested in cell cycle 12 cellularized, initiated gastrulation movements and activated transcription of genes previously described as N/C dependent. Thus, occurrence of these events is not directly coupled to N/C-increase. However, N/C might act indirectly. Increasing N/C promotes cyclin/Cdk1 downregulation which otherwise inhibits many MBT events. By experimentally inducing downregulation of cyclin/Cdk1, we bypassed this input of N/C-increase. We describe a regulatory cascade wherein the increasing N/C downregulates cyclin/Cdk1 to promote increasing transcription and the MBT.\n\nImpact statementBy showing that cell-cycle arrest allows early Drosophila embryos to progress to later stages, this work eliminates numerous models for embryonic timing and shows the dominating influence of cell-cycle slowing.

developmental biology

Inference of Multiple-wave Admixtures by Length Distribution of Ancestral Tracks

The ancestral tracks in admixed genomes are of valuable information for population history inference. A few methods have been developed to infer admixture history based on ancestral tracks. Nonetheless, these methods suffered the same flaw that only population admixture history under some specific models can be inferred. In addition, the inference of history might be biased or even unreliable if the specific model is deviated from the real situation. To address this problem, we firstly proposed a general discrete admixture model to describe the admixture history with multiple ancestral populations and multiple-wave admixtures. We next deduced the length distribution of ancestral tracks under the general discrete admixture model. We further developed a new method, MultiWaver, to explore the multiple-wave admixture histories. Our method could automatically determine an optimal admixture model based on the length distribution of ancestral tracks, and estimate the corresponding parameters under this optimal model. Specifically, we used a likelihood ratio test (LRT) to determine the number of admixture waves, and implemented an expectation??maximization (EM) algorithm to estimate parameters. We used simulation studies to validate the reliability and effectiveness of our method. Finally, good performance was observed when our method was applied to real datasets of African Americans, Mexicans, Uyghurs, and Hazaras.

genetics

BaalChIP: Bayesian analysis of allele-specific transcription factor binding in cancer genomes

Allele-specific measurements of transcription factor binding from ChIP-seq data are key to dissecting the allelic effects of non-coding variants and their contribution to phenotypic diversity. However, most methods to detect allelic imbalance assume diploid genomes. This assumption severely limits their applicability to cancer samples with frequent DNA copy number changes. Here we present a Bayesian statistical approach called BaalChIP to correct for the effect of background allele frequency on the observed ChIP-seq read counts. BaalChIP allows the joint analysis of multiple ChIP-seq samples across a single variant and outperforms competing approaches in simulations. Using 548 ENCODE ChIP-seq and 6 targeted FAIRE-seq samples we show that BaalChIP effectively corrects allele-specific analysis for copy number variation and increases the power to detect putative cis-acting regulatory variants in cancer genomes.

cancer biology

Inference of multiple-wave population admixture by modeling decay of linkage disequilibrium with polynomial functions

To infer the histories of population admixture, one important challenge with methods based on the admixture linkage disequilibrium (ALD) is to get rid of the effect of source LD (SLD) which is directly inherited from source populations. In previous methods, only the decay curve of weighted LD between pairs of sites whose genetic distance were larger than a certain starting distance was fitted by single or multiple exponential functions, for the inference of recent single- or multiple-wave of admixture. However, the effect of SLD has not been well defined and no tool has been developed to estimate the effect of SLD on weighted LD decay. In this study, we defined the SLD in the formularized weighted LD statistic under the two-way admixture model, and proposed polynomial spectrum (p-spectrum) to study the weighted SLD and weighted LD. We also found reference populations could be used to reduce the SLD in weighted LD statistic. We further developed a method, iMAAPs, to infer Multiple-wave Admixture by fitting ALD using Polynomial spectrum. We evaluated the performance of iMAAPs under various admixture models in simulated data and applied iMAAPs into analysis of genome-wide single nucleotide polymorphism data from the Human Genome Diversity Project (HGDP) and the HapMap Project. We showed that iMAAPs is a considerable improvement over other current methods and further facilitates the inference of the histories of complex population admixtures.

evolutionary biology