bioRxiv ScienceSearch

Biology subjects

Mustonen, V.

Publications and source records attributed to Mustonen, V..

7 recordsLinked to original sources

Uncovering natural longevity alleles from intercrossed pools of aging fission yeast cells

Quantitative traits often show large variation caused by multiple genetic factors. One such trait is the chronological lifespan of non-dividing yeast cells, serving as a model for cellular aging. Screens for genetic factors involved in ageing typically assay mutants of protein-coding genes. To identify natural genetic variants contributing to cellular aging, we exploited two strains of the fission yeast, Schizosaccharomyces pombe, that differ in chronological lifespan. We generated segregant pools from these strains and subjected them to advanced intercrossing over multiple generations to break up linkage groups. We chronologically aged the intercrossed segregant pool, followed by genome sequencing at different times to detect genetic variants that became reproducibly enriched as a function of age. A region on Chromosome II showed strong positive selection during ageing. Based on expected functions, two candidate variants from this region in the long-lived strain were most promising to be causal: small insertions and deletions in the 5-untranslated regions of ppk31 and SPBC409.08. Ppk31 is an orthologue of Rim15, a conserved kinase controlling cell proliferation in response to nutrients, while SPBC409.08 is a predicted spermine transmembrane transporter. Both Rim15 and the spermine-precursor, spermidine, are implicated in ageing as they are involved in autophagy-dependent lifespan extension. Single and double allele replacement suggests that both variants, alone or combined, have subtle effects on cellular longevity. Furthermore, deletion mutants of both ppk31 and SPBC409.08 rescued growth defects caused by spermidine. We propose that Ppk31 and SPBC409.08 may function together to modulate lifespan, thus linking Rim15/Ppk31 with spermidine metabolism.

genetics

Precise prediction of antibiotic resistance in Escherichia coli from full genome sequences

The emergence of microbial antibiotic resistance is a global health threat. In clinical settings, the key to controlling spread of resistant strains is accurate and rapid detection. As traditional culture-based methods are time consuming, genetic approaches have recently been developed for this task. The diagnosis is typically made by measuring a few known determinants previously identified from whole genome sequencing, and thus is restricted to existing information on biological mechanisms. To overcome this limitation, we employed machine learning models to predict resistance to 11 compounds across four classes of antibiotics from existing and novel whole genome sequences of 1936 E. coli strains. We considered a range of methods, and examined population structure, isolation year, gene content, and polymorphism information as predictors. Gradient boosted decision trees consistently outperformed alternative models with an average F1 score of 0.88 on held-out data (range 0.66-0.96). While the best models most frequently employed all inputs, an average F1 score of 0.73 could be obtained using population structure information alone. Single nucleotide variation data were less useful, and failed to improve prediction for ten out of 11 antibiotics. These results demonstrate that antibiotic resistance in E. coli can be accurately predicted from whole genome sequences without a priori knowledge of mechanisms, and that both genomic and epidemiological data are informative. This paves way to integrating machine learning approaches into diagnostic tools in the clinic.\n\nSummaryOne of the major health threats of 21st century is emergence of antibiotic resistance. To manage its economic impact, efforts are made to develop novel diagnostic tools that rapidly detect resistant strains in clinical settings. In our study, we employed a range machine learning tools to predict antibiotic resistance from whole genome sequencing data for E. coli. We used the presence or absence of genes, population structure and isolation year of isolates as predictors, and could attain average precision of 0.93 and recall of 0.83, without prior knowledge about the causal mechanisms. These results demonstrate the potential application of machine learning methods as a diagnostic tool in healthcare settings.

bioinformatics

The Repertoire of Mutational Signatures in Human Cancer

Somatic mutations in cancer genomes are caused by multiple mutational processes each of which generates a characteristic mutational signature. Using 84,729,690 somatic mutations from 4,645 whole cancer genome and 19,184 exome sequences encompassing most cancer types we characterised 49 single base substitution, 11 doublet base substitution, four clustered base substitution, and 17 small insertion and deletion mutational signatures. The substantial dataset size compared to previous analyses enabled discovery of new signatures, separation of overlapping signatures and decomposition of signatures into components that may represent associated, but distinct, DNA damage, repair and/or replication mechanisms. Estimation of the contribution of each signature to the mutational catalogues of individual cancer genomes revealed associations with exogenous and endogenous exposures and defective DNA maintenance processes. However, many signatures are of unknown cause. This analysis provides a systematic perspective on the repertoire of mutational processes contributing to the development of human cancer including a comprehensive reference set of mutational signatures in human cancer.

cancer biology

Portraits of genetic intra-tumour heterogeneity and subclonal selection across cancer types

Intra-tumor heterogeneity (ITH) is a mechanism of therapeutic resistance and therefore an important clinical challenge. However, the extent, origin and drivers of ITH across cancer types are poorly understood. To address this question, we extensively characterize ITH across whole-genome sequences of 2,658 cancer samples, spanning 38 cancer types. Nearly all informative samples (95.1%) contain evidence of distinct subclonal expansions, with frequent branching relationships between subclones. We observe positive selection of subclonal driver mutations across most cancer types, and identify cancer type specific subclonal patterns of driver gene mutations, fusions, structural variants and copy-number alterations, as well as dynamic changes in mutational processes between subclonal expansions. Our results underline the importance of ITH and its drivers in tumor evolution, and provide an unprecedented pan-cancer resource of comprehensively annotated subclonal events from whole-genome sequencing data.

cancer biology

Patterns of selection reveal shared molecular targets over short and long evolutionary timescales

Standing and de novo genetic variants can both drive adaptation to environmental changes, but their relative contributions and interplay remain poorly understood. Here we investigated the dynamics of drug adaptation in yeast populations with different levels of standing variation by experimental evolution coupled with time-resolved sequencing and phenotyping. We found a doubling of standing variation alone boost the adaptation by 64.1% and 51.5% in hydroxyuea and rapamycin respectively. The causative standing and de novo variants were selected on shared targets of RNR4 in hydroxyurea and TOR1, TOR2 in rapamycin. The standing and de novo TOR variants map to different functional domains and act via distinct mechanisms. Interestingly, standing TOR variants from two domesticated strains exhibited opposite resistance effects, reflecting lineage-specific functional divergence. This study provides a dynamic view on how standing and de novo variants interactively drive adaptation and deepens our understanding of clonally evolving diseases.

genetics

The asexual genome of Drosophila

The rate of recombination affects the mode of molecular evolution. In high-recombining sequence, the targets of selection are individual genetic loci; under low recombination, selection collectively acts on large, genetically linked genomic segments. Selection under linkage can induce clonal interference, a specific mode of evolution by competition of genetic clades within a population. This mode is well known in asexually evolving microbes, but has not been traced systematically in an obligate sexual organism. Here we show that the Drosophila genome is partitioned into two modes of evolution: a local interference regime with limited effects of genetic linkage, and an interference condensate with clonal competition. We map these modes by differences in mutation frequency spectra, and we show that the transition between them occurs at a threshold recombination rate that is predictable from genomic summary statistics. We find the interference condensate in segments of low-recombining sequence that are located primarily in chromosomal regions flanking the centromeres and cover about 20% of the Drosophila genome. Condensate regions have characteristics of asexual evolution that impact gene function: the efficacy of selection and the speed of evolution are lower and the genetic load is higher than in regions of local interference. Our results suggest that multicellular eukaryotes can harbor heterogeneous modes and tempi of evolution within one genome. We argue that this variation generates selection on genome architecture.\n\nAuthor SummaryThe Drosophila genome is an ideal system to study how the rate of recombination affects molecular evolution. It harbors a wide range of local recombination rates, and its high-recombining parts show broad signatures of adaptive evolution. The low-recombining parts, however, have remained dark genomic matter that has been omitted from most studies on the inference of selection. Here we show that these genomic regions evolve in a different way, which involves clonal competition and is akin to the evolution of asexual systems. This regime shows a lower efficacy of selection, a lower speed of evolution, and a higher genetic load than high-recombining regions. We argue these evolutionary differences have functional consequences: protein stability and protein expression are gene traits likely to be partially compromised by low recombination rates.

genetics

The evolutionary history of 2,658 cancers

Cancer develops through a process of somatic evolution. Here, we use whole-genome sequencing of 2,778 tumour samples from 2,658 donors to reconstruct the life history, evolution of mutational processes, and driver mutation sequences of 39 cancer types. The early phases of oncogenesis are driven by point mutations in a small set of driver genes, often including biallelic inactivation of tumour suppressors. Early oncogenesis is also characterised by specific copy number gains, such as trisomy 7 in glioblastoma or isochromosome 17q in medulloblastoma. By contrast, increased genomic instability, a nearly four-fold diversification of driver genes, and an acceleration of point mutation processes are features of later stages. Copy-number alterations often occur in mitotic crises leading to simultaneous gains of multiple chromosomal segments. Timing analysis suggests that driver mutations often precede diagnosis by many years, and in some cases decades, providing a window of opportunity for early cancer detection.

cancer biology