bioRxiv ScienceSearch

Biology subjects

Shiraishi, Y.

Publications and source records attributed to Shiraishi, Y..

5 recordsLinked to original sources

Comprehensive Analysis of Indels in Whole-genome Microsatellite Regions and Microsatellite Instability across 21 Cancer Types

Microsatellites are repeats of 1-6bp units and [~]10 million microsatellites have been identified across the human genome. Microsatellites are vulnerable to DNA mismatch errors, and have thus been used to detect cancers with mismatch repair deficiency. To reveal the mutational landscape of the microsatellite repeat regions at the genome level, we analyzed approximately 20.1 billion microsatellites in 2,717 whole genomes of pan-cancer samples across 21 tissue types. Firstly, we developed a new insertion and deletion caller (MIMcall) that takes into consideration the error patterns of different types of microsatellites. Among the 2,717 pan-cancer samples, our analysis identified 31 samples, including colorectal, uterus, and stomach cancers, with higher microsatellite mutation rate ([≥] 0.03), which we defined as microsatellite instability (MSI) cancers in genome-wide level. Next, we found 20 highly-mutated microsatellites that can be used to detect MSI cancers with high sensitivity. Third, we found that replication timing and DNA shape were significantly associated with mutation rates of the microsatellites. Analysis of germline variation of the microsatellites suggested that the amount of germline variations and somatic mutation rates were correlated. Lastly, analysis of mutations in mismatch repair genes showed that somatic SNVs and short indels had larger functional impact than germline mutations and structural variations. Our analysis provides a comprehensive picture of mutations in the microsatellite regions, and reveals possible causes of mutations, as well as provides a useful marker set for MSI detection.

cancer biology

A framework for generating interactive reports for cancer genome analysis

SummaryWe introduce paplot, the software for generating dynamic reports that are frequently necessary in the post analytical phases of cancer genome studies. The \"interactive\" nature of the paplot-generated reports enables users to extract much richer information than that obtained from static graphs via most conventional visualization tools.\n\nAvailability and implementationThe python implementation for paplot (MIT license) is available at https://github.com/Genomon-Project/paplot. The documentation is at http://paplot-doc.readthedocs.io/en/latest/.\n\nContactyshira@hgc.jp

bioinformatics

Pan-cancer study of heterogeneous RNA aberrations

We present the most comprehensive catalogue of cancer-associated gene alterations through characterization of tumor transcriptomes from 1,188 donors of the Pan-Cancer Analysis of Whole Genomes project. Using matched whole-genome sequencing data, we attributed RNA alterations to germline and somatic DNA alterations, revealing likely genetic mechanisms. We identified 444 associations of gene expression with somatic non-coding single-nucleotide variants. We found 1,872 splicing alterations associated with somatic mutation in intronic regions, including novel exonization events associated with Alu elements. Somatic copy number alterations were the major driver of total gene and allele-specific expression (ASE) variation. Additionally, 82% of gene fusions had structural variant support, including 75 of a novel class called \"bridged\" fusions, in which a third genomic location bridged two different genes. Globally, we observe transcriptomic alteration signatures that differ between cancer types and have associations with DNA mutational signatures. Given this unique dataset of RNA alterations, we also identified 1,012 genes significantly altered through both DNA and RNA mechanisms. Our study represents an extensive catalog of RNA alterations and reveals new insights into the heterogeneous molecular mechanisms of cancer gene alterations.

genomics

A comprehensive characterization of cis-acting splicing-associated variants in human cancer

Although many driver mutations are thought to promote carcinogenesis via abnormal splicing, the landscape of these splicing-associated variants (SAVs) remains unknown due to the complexity of splicing abnormalities. Here we developed a statistical framework to identify SAVs disrupting or newly creating splice site motifs and applied it to sequencing data from 8,976 samples across 31 cancer types. We constructed a catalog of 14,438 SAVs, approximately 50% of which consist of SAVs disrupting non-canonical splice sites (including the 3rd and 5th intronic bases of donor sites) or newly creating splice sites. Smoking-related signature substantially contributes to SAV generation. As many as 14.7% of samples harbor at least one SAVs in cancer-related genes, particularly in tumor suppressors. Importantly, in addition to previously reported intron retention, exon skipping or alternative splice site usage more frequently affected these genes. Our findings delineate a comprehensive portrait of SAVs, providing a basis for cancer precision medicine.

genomics

A portable system for metagenomic analyses using nanopore-based sequencer and laptop computers can realize rapid on-site determination of bacterial compositions

We developed a portable system for metagenomic analyses consisting of nanopore technology-based sequencer, MinION, and laptop computers, and assessed its potential ability to determine bacterial compositions rapidly. We tested our protocols using mock bacterial community that contained equimolar 16S rDNA and a pleural effusion from a patient with empyema for time effectiveness and accuracy. MinION sequencing targeting 16S rDNA detected all of 20 bacteria present in the mock bacterial community. Time course analysis indicated that sequence data obtained during the first 5-minute sequencing were enough to detect all 20 bacteria species in the mock sample and determine their compositions with sufficient accuracy. Additionally, using a clinical sample extracted from the pleural effusion of a patient with empyema, we could identify major bacteria in a pleural effusion by rapid sequencing and analysis. All of these results are comparable to or even better than the conventional 16S rDNA sequencing results using IonPGM sequencer. Our results suggest that rapid sequencing and bacterial composition determination is possible within 2 hours.Our integrative system is applicable to rapid diagnostic tests for infectious diseases in near future.

genomics