bioRxiv ScienceSearch

Biology subjects

Harismendy, O.

Publications and source records attributed to Harismendy, O..

6 recordsLinked to original sources

Ad-Seq, a genome-wide DNA-adduct profiling assay

SummaryCarcinogens form adducts with the DNA which, when not properly repaired, can lead to mutations and drive oncogenesis. The identity, sequence specificity and mutagenicity of most DNA-adducts is however poorly understood and current molecular assays are limited in their scope and scalability. We present a novel genome-wide DNA adduct sequencing (Ad-Seq) assay to map the location of DNA-adducts at single-nucleotide resolution. Ad-Seq enriches for DNA fragments containing nuclease digestion resistant DNA-adducts. The genomic location of the resulting reads is aggregated in a quantitative profile showing the DNA-adduct sequence context. Ad-Seq is quantitative and confirms known specificity of damages from Ultra-Violet light (di-pyrimidine) and cisplatin (AG and GG di-purines). Furthermore, in cells, Ad-Seq profile can be compared to chromatin segments to show that cisplatin associated adducts are depleted in open and active chromatin regions. The Ad-Seq assay can therefore generate a broad DNA signature of DNA damage and, by comparing to mutagen exposure or downstream mutational profile and signatures, be used to improve our understanding of cancer molecular etiology.

genomics

Evolution and characterization of carboplatin resistance at single cell resolution

Acquired resistance to carboplatin is a major obstacle to the cure of ovarian cancer, but its molecular underpinnings are still poorly understood and often inconsistent between in vitro modeling studies. Using sequential treatment cycles, multiple clones derived from a single ovarian cancer cell reached similar levels of resistance. The resistant clones showed significant transcriptional heterogeneity, with shared repression of cell cycle processes and induction of IFN response signaling, and subsequent pharmacological inhibition of the JAK/STAT pathway led to a general increase in carboplatin sensitivity. Gene-expression based virtual synchronization of 26,772 single cells from 2 treatment steps and 4 resistant clones was used to evaluate the activity of Hallmark gene sets in proliferative (P) and quiescent (Q) phases. Two behaviors were associated with resistance: 1) broad repression in the P phase observed in all clones in early resistant steps and 2) prevalent induction in Q phase observed in the late treatment step of one clone. Furthermore, the induction of IFN response in P phase or Wnt-signaling in Q phase were observed in distinct resistant clones. These observations suggest a model of resistance hysteresis, where functional alterations of the P and Q phase states affect the dynamics of the successive transitions between drug exposure and recovery, and prompts for a precise monitoring of single-cell states to develop more effective schedules for, or combination of, chemotherapy treatments.

cancer biology

Sharing genetic admixture and diversity of public biomedical datasets

Genetic ancestry and admixture are critical co-factors to study phenotype-genotype associations using cohorts of human subjects. Most publically available molecular datasets - genomes, exomes or transcriptomes - are however missing this information or only share self-reported ancestry. This represents a limitation to identify and re-purpose datasets to investigate the contribution of race and ethnicity to diseases and traits. we propose an analytical framework to enrich the meta-data from publically available cohorts with admixture information and a resulting diversity score at continental resolution, calculated directly from the data. We illustrate the utility and versatility of the framework using The Cancer Genome Atlas datasets indexed and searched through the DataMed Data Discovery Index. Data repositories or data contributors can use this framework to provide, as metadata, admixture for controlled access datasets, minimizing the work involved in requesting a dataset that may ultimately prove inadequate for a researchers purpose. With the increasingly global scale of human genetics research, research on disease risk and susceptibility would benefit greatly from the adequate estimation and sharing of admixture data following a framework such as the one presented.

genetics

Large-Scale Uniform Analysis of Cancer Whole Genomes in Multiple Computing Environments

The International Cancer Genome Consortium (ICGC)s Pan-Cancer Analysis of Whole Genomes (PCAWG) project aimed to categorize somatic and germline variations in both coding and non-coding regions in over 2,800 cancer patients. To provide this dataset to the research working groups for downstream analysis, the PCAWG Technical Working Group marshalled ~800TB of sequencing data from distributed geographical locations; developed portable software for uniform alignment, variant calling, artifact filtering and variant merging; performed the analysis in a geographically and technologically disparate collection of compute environments; and disseminated high-quality validated consensus variants to the working groups. The PCAWG dataset has been mirrored to multiple repositories and can be located using the ICGC Data Portal. The PCAWG workflows are also available as Docker images through Dockstore enabling researchers to replicate our analysis on their own data.

genomics

PinAPL-Py: a web-service for the analysis of CRISPR-Cas9 Screens

BackgroundLarge-scale genetic screens using CRISPR/Cas9 technology have emerged as a major tool for functional genomics. With its increased popularity, experimental biologists frequently acquire large sequencing datasets for which they often do not have an easy analysis option. While a few bioinformatic tools have been developed for this purpose, their utility is still hindered either due to limited functionality or the requirement of bioinformatic expertise.\n\nResultsTo make sequencing data analysis of CRISPR/Cas9 screens more accessible to a wide range of scientists, we developed a Platform-independent Analysis of Pooled Screens using Python (PinAPL-Py), which is operated as an intuitive web-service. PinAPL-Py implements state-of-the-art tools and statistical models, assembled in a comprehensive workflow covering sequence quality control, automated sgRNA sequence extraction, alignment, sgRNA enrichment/depletion analysis and gene ranking. The workflow is set up to use a variety of popular sgRNA libraries as well as custom libraries that can be easily uploaded. Various analysis options are offered, suitable to analyze a large variety of CRISPR/Cas9 screening experiments. Analysis output includes ranked lists of sgRNAs and genes, and publication-ready plots.\n\nConclusionsPinAPL-Py helps to advance genome-wide screening efforts by combining comprehensive functionality with user-friendly implementation. PinAPL-Py is freely accessible at http://pinapl-py.ucsd.edu with instructions, documentation and test datasets. The source code is available at https://github.com/LewisLabUCSD/PinAPL-Py

bioinformatics

Pan-Cancer Analysis Reveals Technical Artifacts in The Cancer Genome Atlas (TCGA) Germline Variant Calls

The degree to which germline variation drives cancer development and shapes tumor phenotypes remains largely unexplored, possibly due to a lack of large scale publicly available germline data for a cancer cohort. Here we called germline variants on 9,618 cases from The Cancer Genome Atlas (TCGA) database representing 31 cancer types. We identified batch effects affecting loss of function (LOF) variant calls that can be traced back to differences in the way the sequence data were generated both within and across cancer types. Overall, LOF indel calls were more sensitive to technical artifacts than LOF Single Nucleotide Variant (SNV) calls. In particular, whole genome amplification of DNA prior to sequencing led to an artificially increased burden of LOF indel calls, which confounded association analyses relating germline variants to tumor type despite stringent indel filtering strategies. Due to the inherent noise we chose to remove all 614 amplified DNA samples, including all acute myeloid leukemia and virtually all ovarian cancer samples, from the final dataset. This study demonstrates how insufficient quality control can lead to false positive germlinetumor type associations and draws attention to the need to be sensitive to problems associated with a lack of uniformity in data generation in TCGA data.\n\nAuthor SummaryCancer research to date has largely focused on genetic aberrations specific to tumor tissue. In contrast, the degree to which germline, or inherited, variation contributes to tumorigenesis remains unclear, possibly due to a lack of accessible germline variant data. In this study we identify germline variants in 9,618 samples using raw germline exome data from The Cancer Genome Atlas (TCGA). There are substantial differences in the way exome sequence data was generated both across and within cancer types in TCGA. We observe that differences in sequence data generation introduced batch effects, or variation that is due to technical factors not true biological variation, in our variant data. Most notably, we observe that amplification of DNA prior to sequencing resulted in an excess of predicted damaging indel variants. We show how these batch effects can confound germline association analyses if not properly addressed. Our study highlights the difficulties of working with large public genomic datasets like TCGA where samples are collected over time and across data centers, and particularly cautions the use of amplified DNA samples for genetic association analyses.

genomics