bioRxiv ScienceSearch

Biology subjects

Miyano, S.

Publications and source records attributed to Miyano, S..

8 recordsLinked to original sources

Comprehensive Analysis of Indels in Whole-genome Microsatellite Regions and Microsatellite Instability across 21 Cancer Types

Microsatellites are repeats of 1-6bp units and [~]10 million microsatellites have been identified across the human genome. Microsatellites are vulnerable to DNA mismatch errors, and have thus been used to detect cancers with mismatch repair deficiency. To reveal the mutational landscape of the microsatellite repeat regions at the genome level, we analyzed approximately 20.1 billion microsatellites in 2,717 whole genomes of pan-cancer samples across 21 tissue types. Firstly, we developed a new insertion and deletion caller (MIMcall) that takes into consideration the error patterns of different types of microsatellites. Among the 2,717 pan-cancer samples, our analysis identified 31 samples, including colorectal, uterus, and stomach cancers, with higher microsatellite mutation rate ([≥] 0.03), which we defined as microsatellite instability (MSI) cancers in genome-wide level. Next, we found 20 highly-mutated microsatellites that can be used to detect MSI cancers with high sensitivity. Third, we found that replication timing and DNA shape were significantly associated with mutation rates of the microsatellites. Analysis of germline variation of the microsatellites suggested that the amount of germline variations and somatic mutation rates were correlated. Lastly, analysis of mutations in mismatch repair genes showed that somatic SNVs and short indels had larger functional impact than germline mutations and structural variations. Our analysis provides a comprehensive picture of mutations in the microsatellite regions, and reveals possible causes of mutations, as well as provides a useful marker set for MSI detection.

cancer biology

GIMLET: Identifying Biological Modulators in Context-Specific Gene Regulation Using Local Energy Statistics

The regulation of transcription factor activity dynamically changes across cellular conditions and disease subtypes. The identification of biological modulators contributing to context-specific gene regulation is one of the challenging tasks in systems biology, which is necessary to understand and control cellular responses across different genetic backgrounds and environmental conditions. Previous approaches for identifying biological modulators from gene expression data were restricted to the capturing of a particular type of a three-way dependency among a regulator, its target gene, and a modulator; these methods cannot describe the complex regulation structure, such as when multiple regulators, their target genes, and modulators are functionally related. Here, we propose a statistical method for identifying biological modulators by capturing multivariate local dependencies, based on energy statistics, which is a class of statistics based on distances. Subsequently, our method assigns a measure of statistical significance to each candidate modulator through a permutation test. We compared our approach with that of a leading competitor for identifying modulators, and illustrated its performance through both simulations and real data analysis. Our method, entitled genome-wide identification of modulators using local energy statistical test (GIMLET), is implemented with R ([≥] 3.2.2) and is available from github (https://github.com/tshimam/GIMLET).

bioinformatics

ALPHLARD: a Bayesian method for analyzing HLA genes from whole genome sequence data

Although human leukocyte antigen (HLA) genotyping based on amplicon, whole exome sequence (WES), and RNA sequence data has been achieved in recent years, accurate genotyping from whole genome sequence (WGS) data remains a challenge due to the low depth. Furthermore, there is no method to identify the sequences of unknown HLA types not registered in HLA databases. We developed a Bayesian model, called ALPHLARD, that collects reads potentially generated from HLA genes and accurately determines a pair of HLA types for each of HLA-A, -B, -C, -DPA1, -DPB1, -DQA1, -DQB1, and -DRB1 genes at 6-digit resolution. Furthermore, ALPHLARD can detect rare germline variants not stored in HLA databases and call somatic mutations from paired normal and tumor sequence data. We illustrate the capability of ALPHLARD using 253 WES data and 25 WGS data from Illumina platforms. By comparing the results of HLA genotyping from SBT and amplicon sequencing methods, ALPHLARD achieved 98.8% for WES data and 98.5% for WGS data at 4-digit resolution. We also detected three somatic point mutations and one case of loss of heterozygosity in the HLA genes from the WGS data. ALPHLARD showed good performance for HLA genotyping even from low-coverage data. It also has a potential to detect rare germline variants and somatic mutations in HLA genes. It would help to fill in the current gaps in HLA reference databases and unveil the immunological significance of somatic mutations identified in HLA genes.

bioinformatics

Immuno-genomic PanCancer Landscape Reveals Diverse Immune Escape Mechanisms and Immuno-Editing Histories

Immune reactions in the tumor micro-environment are one of the cancer hallmarks and emerging immune therapies have been proven effective in many types of cancer. To investigate cancer genome-immune interactions and the role of immuno-editing or immune escape mechanisms in cancer development, we analyzed 2,834 whole genomes and RNA-seq datasets across 31 distinct tumor types from the PanCancer Analysis of Whole Genomes (PCAWG) project with respect to key immuno-genomic aspects. We show that selective copy number changes in immune-related genes could contribute to immune escape. Furthermore, we developed an index of the immuno-editing history of each tumor sample based on the information of mutations in exonic regions and pseudogenes. Our immuno-genomic analyses of pan-cancer analyses have the potential to identify a subset of tumors with immunogenicity and diverse background or intrinsic pathways associated with their immune status and immuno-editing history.

genomics

A framework for generating interactive reports for cancer genome analysis

SummaryWe introduce paplot, the software for generating dynamic reports that are frequently necessary in the post analytical phases of cancer genome studies. The \"interactive\" nature of the paplot-generated reports enables users to extract much richer information than that obtained from static graphs via most conventional visualization tools.\n\nAvailability and implementationThe python implementation for paplot (MIT license) is available at https://github.com/Genomon-Project/paplot. The documentation is at http://paplot-doc.readthedocs.io/en/latest/.\n\nContactyshira@hgc.jp

bioinformatics

A comprehensive characterization of cis-acting splicing-associated variants in human cancer

Although many driver mutations are thought to promote carcinogenesis via abnormal splicing, the landscape of these splicing-associated variants (SAVs) remains unknown due to the complexity of splicing abnormalities. Here we developed a statistical framework to identify SAVs disrupting or newly creating splice site motifs and applied it to sequencing data from 8,976 samples across 31 cancer types. We constructed a catalog of 14,438 SAVs, approximately 50% of which consist of SAVs disrupting non-canonical splice sites (including the 3rd and 5th intronic bases of donor sites) or newly creating splice sites. Smoking-related signature substantially contributes to SAV generation. As many as 14.7% of samples harbor at least one SAVs in cancer-related genes, particularly in tumor suppressors. Importantly, in addition to previously reported intron retention, exon skipping or alternative splice site usage more frequently affected these genes. Our findings delineate a comprehensive portrait of SAVs, providing a basis for cancer precision medicine.

genomics

Tumor subclonal progression model for cancer hallmark acquisition

Recent advances in the methods for reconstruction of cancer evolutionary trajectories opened up the prospects of deciphering the subclonal populations and their evolutionary architectures within cancer ecosystems. An important challenge of the cancer evolution studies is how to connect genetic aberrations in subclones to a clinically interpretable and actionable target in the subclones for individual patients. In this study, our aim is to develop a novel method for constructing a model of tumor subclonal progression in terms of cancer hallmark acquisition using multiregional sequencing data. We prepare a subclonal evolutionary tree inferred from variant allele frequencies and estimate pathway alteration probabilities from large-scale cohort genomic data. We then construct an evolutionary tree of pathway alterations that takes into account selectivity of pathway alterations via selectivity score. We show the effectiveness of our method on a dataset of clear cell renal cell carcinomas.

bioinformatics

TOWARDS A GLOBAL SUPPORT OF CORE DATA RESOURCES FOR THE LIFE SCIENCES

On November 18-19, 2016, the Human Frontier Science Program Organization (HFSPO) hosted a meeting of senior managers of key data resources and leaders of several major funding organizations to discuss the challenges associated with sustaining biological and biomedical (i.e., life sciences) data resources and associated infrastructure. A strong consensus emerged from the group that core data resources for the life sciences should be supported through a coordinated international effort(s) that better ensure long-term sustainability and that appropriately align funding with scientific impact. Ideally, funding for such data resources should allow for access at no charge, as is presently the usual (and preferred) mechanism. Below, the rationale for this vision is described, and some important considerations for developing a new international funding model to support core data resources for the life sciences are presented.

scientific communication and education