bioRxiv ScienceSearch

Biology subjects

Shoresh, N.

Publications and source records attributed to Shoresh, N..

2 recordsLinked to original sources

A comprehensive analysis of RNA sequences reveals macroscopic somatic clonal expansion across normal tissues

Cancer genome studies have significantly advanced our knowledge of somatic mutations. However, how these mutations accumulate in normal cells and whether they promote pre-cancerous lesions remains poorly understood. Here we perform a comprehensive analysis of normal tissues by utilizing RNA sequencing data from [~]6,700 samples across 29 normal tissues collected as part of the Genotype-Tissue Expression (GTEx) project. We identify somatic mutations using a newly developed pipeline, RNA-MuTect, for calling somatic mutations directly from RNA-seq samples and their matched-normal DNA. When applied to the GTEx dataset, we detect multiple variants across different tissues and find that mutation burden is associated with both the age of the individual and tissue proliferation rate. We also detect hotspot cancer mutations that share tissue specificity with their matched cancer type. This study is the first to analyze a large number of samples across multiple normal tissues, identifying clones with genomic aberrations observed in cancer.

genomics

Defining the core essential genome of Pseudomonas aeruginosa

Genomics offered the promise of transforming antibiotic discovery by revealing many new essential genes as good targets, but the results fell short of the promise. It is becoming clear that a major limitation was that essential genes for a bacterial species were often defined based on a single or limited number of strains grown under a single or limited number of in vitro laboratory conditions. In fact, the essentiality of a gene can depend on both genetic background and growth condition. We thus developed a strategy for more rigorously defining the core essential genome of a bacterial species by studying many pathogen strains and growth conditions. We assessed how many strains must be examined to converge on a set of core essential genes for a species. We used transposon insertion sequencing (Tn-Seq) to define essential genes in nine strains of Pseudomonas aeruginosa on five different media and developed a novel statistical model, FiTnEss, to classify genes as essential versus non-essential across all strain-media combinations. We defined a set of 321 core essential genes, representing 6.6% of the genome. We determined that analysis of 4 strains was typically sufficient in P. aeruginosa to converge on a set of core essential genes likely to be essential across the species across a wide range of conditions relevant to in vivo infection, and thus to represent attractive targets for novel drug discovery.

microbiology