bioRxiv ScienceSearch

Biology subjects

Gunnar Rätsch

Publications and source records attributed to Gunnar Rätsch.

6 recordsLinked to original sources

Efficient Privacy-Preserving String Search and an Application in Genomics

Motivation: Personal genomes carry inherent privacy risks and protecting privacy poses major social and technological challenges. We consider the case where a user searches for genetic information (e.g., an allele) on a server that stores a large genomic database and aims to receive allele-associated information. The user would like to keep the query and result private and the server the database.\n\nApproach: We propose a novel approach that combines efficient string data structures such as the Burrows-Wheeler transform with cryptographic techniques based on additive homomorphic encryption. We assume that the sequence data is searchable in efficient iterative query operations over a large indexed dictionary, for instance, from large genome collections and employing the (positional) Burrows-Wheeler transform. We use a technique called oblivious transfer that is based on additive homomorphic encryption to conceal the sequence query and the genomic region of interest in positional queries.\n\nResults: We designed and implemented an efficient algorithm for searching sequences of SNPs in large genome databases. During search, the user can only identify the longest match while the server does not learn which sequence of SNPs the user queried. In an experiment based on 2,184 aligned haploid genomes from the 1,000 Genomes Project, our algorithm was able to perform typical queries within {approx}4.6 seconds and {approx}10.8 seconds for client and server side, respectively, on laptop computers. The presented algorithm is at least one order of magnitude faster than an exhaustive baseline algorithm.\n\nAvailability: https://github.com/iskana/PBWT-sec and https://github.com/ratschlab/PBWT-sec.

Genomics

RiboDiff: Detecting Changes of Translation Efficiency from Ribosome Footprints

MotivationDeep sequencing based ribosome footprint profiling can provide novel insights into the regulatory mechanisms of protein translation. However, the observed ribosome profile is fundamentally confounded by transcriptional activity. In order to decipher principles of translation regulation, tools that can reliably detect changes in translation efficiency in case-control studies are needed.\n\nResultsWe present a statistical framework and analysis tool, RiboDiff, to detect genes with changes in translation efficiency across experimental treatments. RiboDiff uses generalized linear models to estimate the over-dispersion of RNA-Seq and ribosome profiling measurements separately, and performs a statistical test for differential translation efficiency using both mRNA abundance and ribosome occupancy.\n\nAvailabilityRiboDiff webpage http://bioweb.me/ribodiff. Source code including scripts for preprocessing the FASTQ data are available at http://github.com/ratschlab/ribodiff.\n\nContactzhongy@cbio.mskcc.org and Gunnar.Ratsch@ratschlab.org.

Bioinformatics

SplAdder: Identification, quantification and testing of alternative splicing events from RNA-Seq data

Motivation: Understanding the occurrence and regulation of alternative splicing (AS) is a key task towards explaining the regulatory processes that shape the complex transcriptomes of higher eukaryotes. With the advent of high-throughput sequencing of RNA (RNA-Seq), the diversity of AS transcripts could be measured at an unprecedented depth. Although the catalog of known AS events has grown ever since, novel transcripts are commonly observed when working with less well annotated organisms, in the context of disease, or within large populations. Whereas an identification of complete transcripts is technically challenging and computationally expensive, focusing on single splicing events as a proxy for transcriptome characteristics is fruitful and sufficient for a wide range of analyses.\n\nResults: We present SplAdder, an alternative splicing toolbox, that takes RNA-Seq alignments and an annotation file as input to i) augment the annotation based on RNA-Seq evidence, ii) identify alternative splicing events present in the augmented annotation graph, iii) quantify and confirm these events based on the RNA-Seq data, and iv) test for significant quantitative differences between samples. Thereby, our main focus lies on performance, accuracy and usability.\n\nAvailability: Source code and documentation are available for download at http://github.com/ratschlab/spladder. Example data, introductory information and a small tutorial are accessible via http://bioweb.me/spladder.\n\nContact: andre.kahles@ratschlab.org, gunnar.ratsch@ratschlab.org

Bioinformatics

MMR: A Tool for Read Multi-Mapper Resolution

MotivationMapping high throughput sequencing data to a reference genome is an essential step for most analysis pipelines aiming at the computational analysis of genome and transcriptome sequencing data. Breaking ties between equally well mapping locations poses a severe problem not only during the alignment phase, but also has significant impact on the results of downstream analyses. We present the multimapper resolution (MMR) tool that infers optimal mapping locations from the coverage density of other mapped reads.\n\nResultsFiltering alignments with MMR can significantly improve the performance of downstream analyses like transcript quantitation and differential testing. We illustrate that the accuracy (Spear-man correlation) of transcript quantification increases by 17% when using reads of length 51. In addition, MMR decreases the alignment file sizes by more than 50% and this leads to a reduced running time of the quantification tool. Our efficient implementation of the MMR algorithm is easily applicable as a post-processing step to existing alignment files in BAM format. Its complexity scales linearly with the number of alignments and requires no further inputs.\n\nSupplementary MaterialSource code and documentation are available for download at github.com/ratschlab/mmr. Supplementary text and figures, comprehensive testing results and further information can be found at bioweb.me/mmr.\n\nContactakahles@cbio.mskcc.org and raetsch@cbio.mskcc.org

Bioinformatics

Integrative Analysis of Transcriptome Variation in Uterine Carcinosarcoma and Comparison to Sarcoma and Endometrial Carcinoma

Large-scale cancer genomics has made a huge impact onto cancer research. It has allowed the characterization of tumor types in an unprecedented depth. More recent studies target the joint analysis of multiple tumor types to gain insight into similarities and differences on a molecular level. Here we present an analysis of Uterine Carcinosarcoma. The histological similarities to sarcomas and carcinomas warrants an in-depth analysis to Uterine Endometrial Carcinoma as well as Sarcomas and we have used data from The Cancer Genome Atlas to understand transcriptome similarities and differences between these tumor types. We have performed a differential transcriptome analysis of Uterine Carinosarcoma to Uterine samples from GTEx to find genes with tumor specific splicing or expression patterns, which may not only be of interest for a deeper mechanistic understanding of the development and progression of Uterine Carcinosarcoma, but may also be potential tumor markers. Similarities and differences to Sarcomas and Endometrial Carcinomas present new opportunities for the development of new and targeted drug therapies. Finally we have also studied genetic determinants of gene expression and splicing changes and identified germline variants that explain expression and splicing differences between individuals. This analysis demonstrates the opportunities of integrative comparative analysis between multiple tumor types.

Cancer Biology

Integrative Genome-wide Analysis of the Determinants of RNA Splicing in Kidney Renal Clear Cell Carcinoma

We present a genome-wide analysis of splicing patterns of 282 kidney renal clear cell carcinoma patients in which we integrate data from whole-exome sequencing of tumor and normal samples, RNA-seq and copy number variation. We proposed a scoring mechanism to compare splicing patterns in tumor samples to normal samples in order to rank and detect tumor-specific isoforms that have a potential for new biomarkers. We identified a subset of genes that show introns only observable in tumor but not in normal samples, ENCODE and GEUVADIS samples. In order to improve our understanding of the underlying genetic mechanisms of splicing variation we performed a large-scale association analysis to find links between somatic or germline variants with alternative splicing events. We identified 915 cis- and trans-splicing quantitative trait loci (sQTL) associated with changes in splicing patterns. Some of these sQTL have previously been associated with being susceptibility loci for cancer and other diseases. Our analysis also allowed us to identify the function of several COSMIC variants showing significant association with changes in alternative splicing. This demonstrates the potential significance of variants affecting alternative splicing events and yields insights into the mechanisms related to an array of disease phenotypes.

Cancer Biology