bioRxiv ScienceSearch

Biology subjects

Shamir, R.

Publications and source records attributed to Shamir, R..

7 recordsLinked to original sources

NEMO: Cancer subtyping by integration of partial multi-omic data

Motivation: Cancer subtypes were usually defined based on molecular characterization of single omic data. Increasingly, measurements of multiple omic profiles for the same cohort are available. Defining cancer subtypes using multi-omic data may improve our understanding of cancer, and suggest more precise treatment for patients.\n\nResults: We present NEMO (NEighborhood based Multi-Omics clustering), a novel algorithm for multiomics clustering. Importantly, NEMO can be applied to partial datasets in which some patients have data for only a subset of the omics, without performing data imputation. In extensive testing on ten cancer datasets spanning 3168 patients, NEMO outperformed nine state-of-the-art multi-omics clustering algorithms on full data and on imputed partial data. On some of the partial data tests, PVC, a multiview algorithm, performed better, but it is limited to two omics and to positive partial data. Finally, we demonstrate the advantage of NEMO in detailed analysis of partial data of AML patients. NEMO is fast and much simpler than existing multi-omics clustering algorithms, and avoids iterative optimization.\n\nAvailability: Code for NEMO and for reproducing all NEMO results in this paper is in github.\n\nContact: rshamir@tau.ac.il\n\nSupplementary information: Supplementary data are available online.

bioinformatics

Multi-omic and multi-view clustering algorithms: review and cancer benchmark

High throughput experimental methods developed in recent years have been used to collect large biomedical omics datasets. Clustering of such datasets has proven invaluable for biological and medical research, and helped reveal structure in data from several domains. Such analysis is often based on investigation of a single omic. The decreasing cost and development of additional high throughput methods now enable measurement of multi-omic data. Clustering multi-omic data has the potential to reveal further systems-level insights, but raises computational and biological challenges. Here we review algorithms for multi-omics clustering, and discuss key issues in applying these algorithms. Our review covers methods developed specifically for multi-omic data as well as generic multi-view methods developed in the machine learning community for joint clustering of multiple data types.\n\nIn addition, using cancer data from TCGA, we perform an extensive benchmark spanning ten different cancer types, providing the first systematic benchmark comparison of leading multi-omics and multiview clustering algorithms. The results highlight several key questions regarding the use of single-vs. multi-omics, the choice of clustering strategy, the power of generic multi-view methods and the use of approximated p-values for gauging solution quality. Due to the rapidly increasing use of multi-omics data, these issues may be important for future progress in the field.

bioinformatics

Genomic meta-analysis of the interplay between 3D chromatin organization and gene expression programs under basal and stress conditions

BackgroundOur appreciation of the critical role of the 3D organization of the genome in gene regulation is steadily increasing. Recent 3C-based deep sequencing techniques elucidated a hierarchy of structures that underlie the spatial organization of the genome in the nucleus. At the top of this hierarchical organization are chromosomal territories and the megabase-scale A/B compartments that correlate with transcriptional activity within cells. Below them are the relatively cell-type invariant topologically associated domains (TADs), characterized by high frequency of physical contacts between loci within the same TAD and are assumed to function as regulatory units. Within TADs, chromatin loops bring enhancers and target promoters to close spatial proximity. Yet, we still have only rudimentary understanding how differences in chromatin organization between different cell types affect cell-type specific gene expression programs that are executed under basal and challenged conditions.\n\nResultsHere, we carried out a large-scale meta-analysis that integrated Hi-C data from thirteen different cell lines and dozens of ChIP-seq and RNA-seq datasets measured on these cells, either under basal conditions or after treatment. Pairwise comparisons between cell lines demonstrated the strong association between modulation of A/B compartmentalization, differential gene expression and transcription factor (TF) binding events. Furthermore, integrating the analysis of transcriptomes of different cell lines in response to various challenges, we show that 3D organization of cells under basal conditions constrains not only gene expression programs and TF binding profiles that are active under the basal condition but also those induced in response to treatment.\n\nConclusionsOur results further elucidate the role of dynamic genome organization in regulation of differential gene expression between different cell types, and indicate the impact of intra-TAD enhancer-promoter interactions that are established under basal conditions on both the basal and treatment-induced gene expression programs.

genomics

An extensive enhancer-promoter map generated by genome-scale analysis of enhancer and gene activity patterns

Massive efforts have documented hundreds of thousands of putative enhancers in the human genome. A pressing genomic challenge is to identify which of these enhancers are functional and map them to the genes they regulate. We developed a novel method for inferring enhancer-promoter (E-P) links based on correlated activity patterns across many samples. Our method, called FOCS, uses rigorous statistical validation tailored for zero-inflated data, identifying the most important E-P links in each gene model. We applied FOCS to the wide epigenomic and transcriptomic datasets recorded by the ENCODE, Roadmap Epigenomics and FANTOM5 projects, together covering 2,630 samples of human primary cells, tissues and cell lines. In addition, building on expression of enhancer RNAs (eRNAs) as an exquisite mark of enhancer activity and on the robust detection of eRNAs by the GRO-seq technique, we compiled a compendium of eRNA and gene expression profiles based on public GRO-seq data from 245 samples and 23 human cell types. Applying FOCS to this compendium further expanded the coverage of our inferred E-P map. Benchmarking against gold standard E-P links from ChIA-PET and eQTL data, we demonstrate that FOCS prediction of E-P links outperforms extant methods. Collectively, we inferred >300,000 cross-validated E-P links spanning ~16K known genes. Our study presents an improved method for inferring regulatory links between enhancers and promoters, and provides an extensive resource of E-P maps that could greatly assist the functional interpretation of the noncoding regulatory genome. FOCS and our predicted E-P map are publicly available at http://acgt.cs.tau.ac.il/focs.

bioinformatics

Reconstructing cancer karyotypes from short read data: the half full and half empty glass

BackgroundDuring cancer progression genomes undergo point mutations as well as larger segmental changes. The latter include, among others, segmental deletions duplications, translocations and inversions. The result is a highly complex, patient-specific cancer karyotype. Using high-throughput technologies of deep sequencing and microarrays it is possible to interrogate a cancer genome and produce chromosomal copy number profiles and a list of breakpoints (\"jumps\") relative to the normal genome. This information is very detailed but local, and does not give the overall picture of the cancer genome. One of the basic challenges in cancer genome research is to use such information to infer the cancer karyotype.\n\nWe present here an algorithmic approach, based on graph theory and integer linear programming, that receives segmental copy number and breakpoint data as input and produces a cancer karyotype that is most concordant with them. We used simulations to evaluate the utility of our approach, and applied it to real data.\n\nResultsBy using a simulation model, we were able to estimate the correctness and robustness of the algorithm in a spectrum of scenarios. Under our base scenario, designed according to observations in real data, the algorithm correctly inferred 69% of the karyotypes. However, when using less stringent correctness metrics that account for incomplete and noisy data, 87% of the reconstructed karyotypes were correct. Furthermore, in scenarios where the data were very clean and complete, accuracy rose to 90%-100%. Some examples of analysis of real data, and the karyotypes reconstructed by our algorithm, are also presented.\n\nConclusionWhile reconstruction of complete, perfect karyotype based on short read data is very hard, a large portion of the reconstruction will still be correct and can provide useful information.

bioinformatics

Faucet: streaming de novo assembly graph construction

MotivationWe present Faucet, a 2-pass streaming algorithm for assembly graph construction. Faucet builds an assembly graph incrementally as each read is processed. Thus, reads need not be stored locally, as they can be processed while downloading data and then discarded. We demonstrate this functionality by performing streaming graph assembly of publicly available data, and observe that the ratio of disk use to raw data size decreases as coverage is increased.\n\nResultsFaucet pairs the de Bruijn graph obtained from the reads with additional meta-data derived from them. We show these metadata - coverage counts collected at junction k-mers and connections bridging between junction pairs - contain most salient information needed for assembly, and demonstrate they enable cleaning of metagenome assembly graphs, greatly improving contiguity while maintaining accuracy. We compared Faucets resource use and assembly quality to state of the art metagenome assemblers, as well as leading resource-efficient genome assemblers. Faucet used orders of magnitude less time and disk space than the specialized metagenome assemblers MetaSPAdes and Megahit, while also improving on their memory use; this broadly matched performance of other assemblers optimizing resource efficiency - namely, Minia and LightAssembler. However, on metagenomes tested, Faucets outputs had 14-110% higher mean NGA50 lengths compared to Minia, and 2-11-fold higher mean NGA50 lengths compared to LightAssembler, the only other streaming assembler available.\n\nAvailabilityFaucet is available at https://github.com/Shamir-Lab/Faucet\n\nContactrshamir@tau.ac.il,eranhalperin@gmail.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Improving the performance of minimizers andwinnowing schemes

The minimizers scheme is a method for selecting k-mers from sequences. It is used in many bioinformatics software tools to bin comparable sequences or to sample a sequence in a deterministic fashion at approximately regular intervals, in order to reduce memory consumption and processing time. Although very useful, the minimizers selection procedure has undesirable behaviors (e.g., too many k-mers are selected when processing certain sequences). Some of these problems were already known to the authors of the minimizers technique, and the natural lexicographic ordering of k-mers used by minimizers was recognized as their origin. Many software tools using minimizers employ ad hoc variations of the lexicographic order to alleviate those issues.\n\nWe provide an in-depth analysis of the effect of k-mer ordering on the performance of the minimizers technique. By using small universal hitting sets (a recently defined concept), we show how to significantly improve the performance of minimizers and avoid some of its worse behaviors. Based on these results, we encourage bioinformatics software developers to use an ordering based on a universal hitting set or, if not possible, a randomized ordering, rather than the lexicographic order. This analysis also settles negatively a conjecture (by Schleimer et al.) on the expected density of minimizers in a random sequence.\n\nThe software used for this analysis is available on GitHub: https://github.com/gmarcais/minimizers.git.\n\nContact: gmarcais@cs.cmu.edu

bioinformatics