bioRxiv Science⌕ Search

Biology subjects

Neff, S. L.

Publications and source records attributed to Neff, S. L..

3 recordsLinked to original sources

Compendium-wide analysis of P. aeruginosa core and accessory genes reveal more nuanced transcriptional patterns

Strains of Pseudomonas aeruginosa, an opportunistic pathogen that causes difficult to treat infections, have significant genomic heterogeneity including the presence of diverse accessory genes that are only present in some strains or clades. Both core genes, which are conserved across strains, and accessory genes have been associated with traits such as biofilm formation and virulence. Much of what we know about core and accessory gene content comes from genome analyses. Here, we use a newly assembled transcriptome compendium to analyze the transcriptional patterns of core and accessory gene expression in PAO1 and PA14 strains across thousands of samples from hundreds of distinct experiments. We found that a subset of core genes were stable, having consistent correlated expression patterns across samples regardless of strain background, with a focus on strains PAO1 and PA14. These stable core genes had fewer co-expressed neighbors that were accessory genes. SignificancePseudomonas aeruginosa is a ubiquitous pathogen. There is a lot of diversity amongst P. aeruginosa strains, some which are clinically relevant. Understanding how these different strain-level traits manifest is important for identifying targets that regulate different traits of interest. With the availability of a PAO1-mapped and PA14-mapped RNA-seq compendium, which contain hundreds of strains, it is now possible to examine the effect of different strains on expression, which can mediate different traits. In this study we developed an approach to compare expression profiles across different P. aeruginosa gene groups - core and accessory genes. This approach revealed a subset of core genes with different transcriptional patterns across strains, which could contribute to trait differences.

bioinformatics↗

CF-Seq, An Accessible Web Application for Rapid Re-Analysis of Cystic Fibrosis Pathogen RNA Sequencing Studies

Researchers studying cystic fibrosis (CF) pathogens have produced numerous RNA-seq datasets which are available in the gene expression omnibus (GEO). Although these studies are publicly available, substantial computational expertise and manual effort are required to compare similar studies, visualize gene expression patterns within studies, and use published data to generate new experimental hypotheses. Furthermore, it is difficult to filter available studies by domain-relevant attributes such as strain, treatment, or media, or for a researcher to assess how a specific gene responds to various experimental conditions across studies. To reduce these barriers to data re-analysis, we have developed an R Shiny application called CF-Seq, which works with a compendium of 147 studies and 1,446 individual samples from 13 clinically relevant CF pathogens. The application allows users to filter studies by experimental factors and to view complex differential gene expression analyses at the click of a button. Here we present a series of use cases that demonstrate the application is a useful and efficient tool for new hypothesis generation. (CFSeq: http://scangeo.dartmouth.edu/CFSeq/)

bioinformatics↗

Computationally efficient assembly of a Pseudomonas aeruginosa gene expression compendium

Over the past two decades, thousands of RNA sequencing (RNA-seq) gene expression profiles of Pseudomonas aeruginosa have been made publicly available via the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA). In the work we present here, we draw on over 2,300 P. aeruginosa transcriptomes from hundreds of studies performed by over seventy-five different research groups. We first developed a pipeline, using the Salmon pseudo-aligner and two different P. aeruginosa reference genomes (strains PAO1 and PA14), that transformed raw sequence data into a uniformly processed data in the form of sample-wise normalized counts. In this workflow, P. aeruginosa RNA-seq data are filtered using technically and biologically driven criteria with characteristics tailored to bacterial gene expression and that account for the effects of alignment to different reference genomes. The filtered data are then normalized to enable cross experiment comparisons. Finally, annotations are programmatically collected for those samples with sufficient meta-data and expression-based metrics are used to further enhance strain assignment for each sample. Our processing and quality control methods provide a scalable framework for taking full advantage of the troves of biological information hibernating in the depths of microbial gene expression data. The re-analysis of these data in aggregate is a powerful approach for hypothesis generation and testing, and this approach can be applied to transcriptome datasets in other species. SignificancePseudomonas aeruginosa causes a wide range of infections including chronic infections associated with cystic fibrosis. P. aeruginosa infections are difficult to treat and people with CF-associated P. aeruginosa infections often have poor clinical outcomes. To aid the study of this important pathogen, we developed a methodology that facilitates analyses across experiments, strains, and conditions. We aligned, filtered for quality and normalized thousands of P. aeruginosa RNA-seq gene expression profiles that were publicly available via the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA). The workflow that we present can be efficiently scaled to incorporate new data and applied to the analysis of other species.

microbiology↗