bioRxiv ScienceSearch

Biology subjects

Mell, J. C.

Publications and source records attributed to Mell, J. C..

5 recordsLinked to original sources

Species-level bacterial community profiling of the healthy sinonasal microbiome using Pacific Biosciences sequencing of full-length 16S rRNA genes

BackgroundPan-bacterial 16S rRNA microbiome surveys performed with massively parallel DNA sequencing technologies have transformed community microbiological studies. Current 16S profiling methods, however, fail to provide sufficient taxonomic resolution and accuracy to adequately perform species-level associative studies for specific conditions. This is due to the amplification and sequencing of only short 16S rRNA gene regions, typically providing for only family- or genus-level taxonomy. Moreover, sequencing errors often inflate the number of taxa present. Pacific Biosciences (PacBios) long-read technology in particular suffers from high error rates per base. Herein we present a microbiome analysis pipeline that takes advantage of PacBio circular consensus sequencing (CCS) technology to sequence and error correct full-length bacterial 16S rRNA genes, which provides high-fidelity species-level microbiome data\n\nResultsAnalysis of a mock community with 20 bacterial species demonstrated 100% specificity and sensitivity. Examination of a 250-plus species mock community demonstrated correct species-level classification of >90% of taxa and relative abundances were accurately captured. The majority of the remaining taxa were demonstrated to be multiply, incorrectly, or incompletely classified. Using this methodology, we examined the microgeographic variation present among the microbiomes of six sinonasal sites, by both swab and biopsy, from the anterior nasal cavity to the sphenoid sinus from 12 subjects undergoing trans-sphenoidal hypophysectomy. We found greater variation among subjects than among sites within a subject, although significant within-individual differences were also observed. Propiniobacterium acnes (recently renamed Cutibacterium acnes [1]) was the predominant species throughout, but was found at distinct relative abundances by site.\n\nConclusionsOur microbial composition analysis pipeline for single-molecule real-time 16S rRNA gene sequencing (MCSMRT, https://github.com/jpearl01/mcsmrt) overcomes deficits of standard marker gene based microbiome analyses by using CCS of entire 16S rRNA genes to provide increased taxonomic and phylogenetic resolution. Extensions of this approach to other marker genes could help refine taxonomic assignments of microbial species and improve reference databases, as well as strengthen the specificity of associations between microbial communities and dysbiotic states.

bioinformatics

A competence-regulated toxin-antitoxin system in Haemophilus influenzae

Natural competence allows bacteria to respond to environmental and nutritional cues by taking up free DNA from their surroundings, thus gaining nutrients and genetic information. In the Gram-negative bacterium Haemophilus influenae, the DNA uptake machinery is induced by the CRP and Sxy transcription factors in response to lack of preferred carbon sources and nucleotide precursors. Here we show that HI0659--which is absolutely required for DNA uptake-- encodes the antitoxin of a competence-regulated toxin-antitoxin operon ( toxTA), likely acquired by horizontal gene transfer from a Streptococcus species. Deletion of the toxin restores uptake to the antitoxin mutant. In addition to the expected Sxy-and CRP-dependent-competence promoter, transcript analysis using RNA-seq identified an internal antitoxin-repressed promoter whose transcription starts within toxT and will yield nonfunctional protein. We present evidence that the most likely effect of unopposed toxin expression is non-specific cleavage of mRNAs and arrest or death of competent cells in the culture, and we show that the toxin gene has been inactivated by deletion in many H. influenzae strains. We suggest that this competence-regulated toxin-antitoxin system may facilitate downregulation of protein synthesis and recycling of nucleotides under starvation conditions, or alternatively be a simple genetic parasite.

microbiology

Microbiome-TP53 Gene Interaction in Human Lung Cancer

BackgroundLung cancer is the leading cancer diagnosis worldwide and the number one cause of cancer deaths. Exposure to cigarette smoke, the primary risk factor in lung cancer, reduces epithelial barrier integrity and increases susceptibility to infections. Herein, we hypothesized that somatic mutations together with cigarette smoke generate a dysbiotic microbiota that is associated with lung carcinogenesis. Using lung tissue from controls (n=33) and cancer cases (n=143), we conducted 16S rRNA bacterial gene sequencing, with RNA-seq data from lung cancer cases in The Cancer Genome Atlas (n=1112) serving as the validation cohort.\n\nResultsOverall, we demonstrate a lower alpha diversity in normal lung as compared to non-tumor adjacent or tumor tissue. In squamous cell carcinoma (SCC) specifically, a separate group of taxa were identified, in which Acidovorax was enriched in smokers (P =0.0013). Acidovorax temporans was identified by fluorescent in situ hybridization within tumor sections, and confirmed by two separate 16S rRNA strategies. Further, these taxa, including Acidovorax, exhibited higher abundance among the subset of SCC cases with TP53 mutations, an association not seen in adenocarcinomas (AD).\n\nConclusionsThe results of this comprehensive study show both a microbiome-gene and microbiome-exposure interactions in SCC lung cancer tissue. Specifically, tumors harboring TP53 mutations, which can damage epithelial function, have a unique bacterial consortia which is higher in relative abundance in smoking-associated SCC. Given the significant need for clinical diagnostic tools in lung cancer, this study may provide novel biomarkers for early detection.

cancer biology

Evaluating a topic model approach for parsing microbiome data structure

The increasing availability of microbiome survey data has led to the use of complex machine learning and statistical approaches to measure taxonomic diversity and extract relationships between taxa and their host or environment. However, many approaches inadequately account for the difficulties inherent to microbiome data. These difficulties include (1) insufficient sequencing depth resulting in sparse count data, (2) a large feature space relative to sample space, resulting in data prone to overfitting, (3) library size imbalance, requiring normalization strategies that lead to compositional artifacts, and (4) zero-inflation. Recent work has used probabilistic topics models to more appropriately model microbiome data, but a thorough inspection of just how well topic models capture underlying microbiome signal is lacking. Also, no work has determined whether library size or variance normalization improves model fitting. Here, we assessed a topic model approach on 16S rRNA gene survey data. Through simulation, we show, for small sample sizes, library-size or variance normalization is unnecessary prior to fitting the topic model. In addition, by exploiting topic-to-topic correlations, the topic model successfully captured dynamic time-series behavior of simulated taxonomic subcommunities. Lastly, when the topic model was applied to the David et al. time-series dataset, three distinct gut configurations emerged. However, unlike the David et al. approach, we characterized the events in terms of topics, which captured taxonomic co-occurrence, and posterior uncertainty, which facilitated the interpretation of how the taxonomic configurations evolved over time.

bioinformatics

Exploring thematic structure in 16S rRNA marker gene surveys

BackgroundAnalysis of microbiome data involves identifying co-occurring groups of taxa associated with sample features of interest (e.g., disease state). But elucidating key associations is often difficult since microbiome data are compositional, high dimensional, and sparse. Also, the configuration of co-occurring taxa may represent overlapping subcommunities that contribute to, for example, host status. Preserving the configuration of co-occurring microbes rather than detecting specific indicator species is more likely to facilitate biologically meaningful interpretations. In addition, analyses that utilize both taxonomic and predicted functional abundances typically independently characterize the taxonomic and functional profiles before linking them to sample information. This prevents investigators from identifying the specific functional components associate with which subsets of co-occurring taxa.\n\nResultsWe provide an approach to explore co-occurring taxa using \"topics\" generated via a topic model and then link these topics to specific sample classes (e.g., diseased versus healthy). Rather than inferring predicted functional content independently from taxonomic abundances, we instead focus on inference of functional content within topics, which we parse by estimating pathway-topic interactions through a multilevel, fully Bayesian regression model. We apply our methods to two large publically available 16S amplicon sequencing datasets: an inflammatory bowel disease (IBD) dataset from Gevers et al. and data from the American Gut (AG) project. When applied to the Gevers et al. IBD study, we demonstrate that a topic highly associated with Crohns disease (CD) diagnosis is (1) dominated by a cluster of bacteria known to be linked with CD and (2) uniquely enriched for a subset of lipopolysaccharide (LPS) synthesis genes. In the AG data, our approach found that individuals with plant-based diets were enriched with Lachnospiraceae, Roseburia, Blautia, and Ruminococcaceae, as well as fluorobenzoate degradation pathways, whereas pathways involved in LPS biosynthesis were depleted.\n\nConclusionsWe introduce an approach for uncovering latent thematic structure in the context of sample features for 16S rRNA surveys. Using our topic-model approach, investigators can (1) capture groups of co-occurring taxa termed topics, (2) uncover within-topic functional potential, and (3) identify gene sets that may guide future inquiry. These methods have been implemented in a freely available R package https://github.com/EESI/themetagenomics.

bioinformatics