bioRxiv ScienceSearch

Biology subjects

Morganella, S.

Publications and source records attributed to Morganella, S..

5 recordsLinked to original sources

The Repertoire of Mutational Signatures in Human Cancer

Somatic mutations in cancer genomes are caused by multiple mutational processes each of which generates a characteristic mutational signature. Using 84,729,690 somatic mutations from 4,645 whole cancer genome and 19,184 exome sequences encompassing most cancer types we characterised 49 single base substitution, 11 doublet base substitution, four clustered base substitution, and 17 small insertion and deletion mutational signatures. The substantial dataset size compared to previous analyses enabled discovery of new signatures, separation of overlapping signatures and decomposition of signatures into components that may represent associated, but distinct, DNA damage, repair and/or replication mechanisms. Estimation of the contribution of each signature to the mutational catalogues of individual cancer genomes revealed associations with exogenous and endogenous exposures and defective DNA maintenance processes. However, many signatures are of unknown cause. This analysis provides a systematic perspective on the repertoire of mutational processes contributing to the development of human cancer including a comprehensive reference set of mutational signatures in human cancer.

cancer biology

Partially methylated domains are hypervariable in breast cancer and fuel widespread CpG island hypermethylation

Global loss of DNA methylation and CpG island (CGI) hypermethylation are regarded as key epigenomic aberrations in cancer. Global loss manifests itself in partially methylated domains (PMDs) which can extend up to megabases. However, the distribution of PMDs within and between tumor types, and their effects on key functional genomic elements including CGIs are poorly defined. Using whole genome bisulfite sequencing (WGBS) of breast cancers, we comprehensively show that loss of methylation in PMDs occurs in a large fraction of the genome and represents the prime source of variation in DNA methylation. PMDs are hypervariable in methylation level, size and distribution, and display elevated mutation rates. They impose intermediate DNA methylation levels incognizant of functional genomic elements including CGIs, underpinning a CGI methylator phenotype (CIMP). However, significant repression effects on cancer-genes are negligible as tumor suppressor genes are generally excluded from PMDs. The genomic distribution of PMDs reports tissue-of-origin of different cancers and may represent tissue-specific silent regions of the genome, which tolerate instability at the epigenetic, transcriptomic and genetic level.

cancer biology

Non-canonical secondary structures arising from non-B DNA motifs are determinants of mutagenesis

Somatic mutations show variation in density across cancer genomes. Previous studies have shown that chromatin organization and replication time domains are correlated with and thus predictive of this variation 1,2,3,4,5. Here, we analyse 1,809 whole-genome sequences from nine cancer types 6,7,8 to show that a subset of repetitive DNA sequences called non-B motifs that predict non-canonical secondary structure formation 9,10,11,12 can independently account for variation in mutation density. However, combined with epigenetic factors and replication timing, the variance explained can be improved to 43-76%. Intriguingly, ~2-fold mutation enrichment is observed directly within non-B motifs, is focused on exposed structural components, and is dependent on physical properties that are optimal for secondary structure formation. Therefore, there is mounting evidence that secondary structures arising from non-B motifs are not simply associated with increased mutation density, they are possibly causally implicated. Our results suggest that they are determinants of mutagenesis and increase the likelihood of recurrent mutations in the genome 13,6. This analysis calls for caution in the interpretation of recurrent mutations and highlights the importance of taking non-B motifs, that can simply be inferred from the reference sequence, into consideration in background models of mutability henceforth.

genomics

ChromoTrace: Reconstruction of 3D Chromosome Configurations by Super-Resolution Microscopy

MotivationThe three-dimensional structure of chromatin plays a key role in genome function, including gene expression, DNA replication, chromosome segregation, and DNA repair. Furthermore the location of genomic loci within the nucleus, especially relative to each other and nuclear structures such as the nuclear envelope and nuclear bodies strongly correlates with aspects of function such as gene expression. Therefore, determining the 3D position of the 6 billion DNA base pairs in each of the 23 chromosomes inside the nucleus of a human cell is a central challenge of biology. Recent advances of super-resolution microscopy in principle enable the mapping of specific molecular features with nanometer precision inside cells. Combined with highly specific, sensitive and multiplexed fluorescence labeling of DNA sequences this opens up the possibility of mapping the 3D path of the genome sequence in situ.\n\nResultsHere we develop computational methodologies to reconstruct the sequence configuration of all human chromosomes in the nucleus from a super-resolution image of a set of fluorescent in situ probes hybridized to the genome in a cell. To test our approach we develop a method for the simulation of chromatin packing in an idealized human nucleus. Our reconstruction method, ChromoTrace, uses suffix trees to assign a known linear ordering of in situ probes on the genome to an unknown set of 3D in situ probe positions in the nucleus from super-resolved images using the known genomic probe spacing as a set of physical distance constraints between probes. We find that ChromoTrace can assign the 3D positions of the majority of loci with high accuracy and reasonable sensitivity to specific genome sequences. By simulating spatial resolution, label multiplexing and noise scenarios we assess algorithm performance under realistic experimental constraints. Our study shows that it is feasible to achieve chromosome-wide reconstruction of the 3D DNA path in chromatin based on super-resolution microscopy images.

bioinformatics

GARFIELD - GWAS Analysis of Regulatory or Functional Information Enrichment with LD correction

Loci discovered by genome-wide association studies (GWAS) predominantly map outside protein-coding genes. The interpretation of functional consequences of non-coding variants can be greatly enhanced by catalogs of regulatory genomic regions in cell lines and primary tissues. However, robust and readily applicable methods are still lacking to systematically evaluate the contribution of these regions to genetic variation implicated in diseases or quantitative traits. Here we propose a novel approach that leverages GWAS findings with regulatory or functional annotations to classify features relevant to a phenotype of interest. Within our framework, we account for major sources of confounding that current methods do not offer. We further assess enrichment statistics for 27 GWAS traits within regulatory regions from the ENCODE and Roadmap projects. We characterise unique enrichment patterns for traits and annotations, driving novel biological insights. The method is implemented in standalone software and R package to facilitate its application by the research community.

genomics