bioRxiv Science⌕ Search

Biology subjects

Sosinsky, A.

Publications and source records attributed to Sosinsky, A..

4 recordsLinked to original sources

Whole genome sequencing of endometrial cancer identifies novel subgroups, drivers, and actionable alterations

Endometrial cancer (EC) is the most common gynaecological malignancy in high income countries, and is increasing in incidence. While molecular stratification has improved its management, precision care is hampered by incomplete characterization of the EC genome. We address this by analysis of whole genome sequencing (WGS) of 665 ECs generated by the UK Genomics England 100,000 Genome Project (100kGP). 5% of cases were associated with germline pathogenic variants in cancer genes, including BRCA1 which we confirmed predisposes to EC. We identified 107 putative coding driver genes, 35% of which had no prior established role in EC. Novel structural variants included gains of MYCN and loss of its negative regulator NEDD4.1 which were significantly mutually exclusive in copy number (CN) high tumours. Immunogenomic analysis confirmed selection for driver alterations of low immunogenicity based on patient HLA haplotype, and pervasive immune escape through multiple mechanisms. Unsupervised clustering of mutational signatures and genomic alterations identified known and novel molecular subgroups, including a CN-high subset with mutational signatures of homologous recombination deficiency (HRD) and favourable outcome. Independent prognostic value of single nucleotide variant (SNV) burden, CN burden and multiple coding drivers, along with the identification of targetable molecular alterations in over one-third of cases, underscores the promise of WGS for precision medicine in EC.

genomics↗

SAVANA: reliable analysis of somatic structural variants and copy number aberrations in clinical samples using long-read sequencing

Accurate detection of somatic structural variants (SVs) and copy number aberrations (SCNAs) is critical to inform the diagnosis and treatment of human cancers. Here, we describe SAVANA, a computationally efficient algorithm designed for the joint analysis of somatic SVs, SCNAs, tumour purity and ploidy using long-read sequencing data. SAVANA relies on machine learning to distinguish true somatic SVs from artefacts and provide prediction errors for individual SVs. Using high-depth Illumina and nanopore whole-genome sequencing data for 99 human tumours and matched normal samples, we establish best practices for benchmarking SV detection algorithms across the entire genome in an unbiased and data-driven manner using simulated and sequencing replicates of tumour and matched normal samples. SAVANA shows significantly higher sensitivity, and 9- and 59-times higher specificity than the second and third-best performing algorithms, yielding orders of magnitude fewer false positives in comparison to existing long-read sequencing tools across various clonality levels, genomic regions, SV types and SV sizes. In addition, SAVANA harnesses long-range phasing information to detect somatic SVs and SCNAs at single-haplotype resolution. SVs reported by SAVANA are highly consistent with those detected using short-read sequencing, including complex events causing oncogene amplification and tumour suppressor gene inactivation. In summary, SAVANA enables the application of long-read sequencing to detect SVs and SCNAs reliably in clinical samples.

cancer biology↗

Whole genome sequencing of 2,023 colorectal cancers reveals mutational landscapes, new driver genes and immune interactions

To characterise the somatic alterations in colorectal cancer (CRC), we conducted whole-genome sequencing analysis of 2,023 tumours. We provide the most detailed high-resolution map to date of somatic mutations in CRC, and demonstrate associations with clinicopathological features, in particular location in the large bowel. We refined the mutational processes and signatures acting in colorectal tumorigenesis. In analyses across the sample set or restricted to molecular subtypes, we identified 185 CRC driver genes, of which 117 were previously unreported. New drivers acted in various molecular pathways, including Wnt (CTNND1, AXIN1, TCF3), TGF-{beta}/BMP (TGFBR1) and MAP kinase (RASGRF1, RASA1, RAF1, and several MAP2K and MAP3K loci). Non-coding drivers included intronic neo-splice site alterations in APC and SMAD4. Whilst there was evidence of an excess of mutations in functionally active regions of the non-coding genome, no specific drivers were called with high confidence. Novel recurrent copy number changes included deletions of PIK3R1 and PWRN1, as well as amplification of CCND3 and NEDD9. Putative driver structural variants included BRD4 and SOX9 regulatory elements, and ACVR2A and ANKRD11 hotspot deletions. The frequencies of many driver mutations, including somatic Wnt and Ras pathway variants, showed a gradient along the colorectum. The Pks-pathogenic E. coli signature and TP53 mutations were primarily associated with rectal cancer. A set of unreported immune escape driver genes was found, primarily in hypermutated CRCs, most of which showed evidence of genetic evasion of the anti-cancer immune response. About 25% of cancers had a potentially actionable mutation for a known therapy. Thirty-three of the new driver genes were predicted to be essential, 17 possessed a druggable structure, and nine had a bioactive compound available. Our findings provide further insight into the genetics and biology of CRC, especially tumour subtypes defined by genomic instability or clinicopathological features.

genomics↗

Clinical application of tumour in normal contamination assessment from whole genome sequencing

The unexpected contamination of normal samples with tumour cells reduces variant detection sensitivity, compromising downstream analyses in canonical tumour-normal analyses. Leveraging whole-genome sequencing data available at Genomics England, we develop a tool for normal sample contamination assessment, which we validate in silico and against minimal residual disease testing. From a systematic review of 771 patients with haematological malignancies and sarcomas, we find contamination across a range of cancer clinical indications and DNA sources, with highest prevalence in saliva samples from acute myeloid leukaemia patients, and sorted CD3+ T-cells from myeloproliferative neoplasms. Further exploration reveals 108 hotspot mutations in genes associated with haematological cancers at risk of being subtracted by standard variant calling pipelines. Our work highlights the importance of contamination assessment for accurate somatic variants detection in research and clinical settings, especially with large-scale sequencing projects being utilised to deliver accurate data from which to make clinical decisions for patient care.

bioinformatics↗