bioRxiv ScienceSearch

Biology subjects

Brown, C. T.

Publications and source records attributed to Brown, C. T..

7 recordsLinked to original sources

Comparing faster evolving rplB and rpsC versus SSU rRNA for improved microbial community resolution

Many conserved protein-coding core genes are single copy and evolve faster, and thus are more resolving phylogenetic markers than the standard SSU rRNA gene but their use has been precluded by the lack of universal primers. Recent advances in gene targeted assembly methods for large shotgun metagenomes make their use feasible. To evaluate this approach, we compared the variation of two single copy ribosomal protein genes, rplB and rpsC, with the SSU rRNA gene for all completed bacterial genomes in NCBI RefSeq. As expected, among pairwise comparisons of all species that belong to the same genus, 94.9% and 91.0% of the pairs of rplB and rpsC, respectively, showed more variation than did their SSU rRNA gene sequences. We used a gene-targeted assembler, Xander, to assemble rplB and rpsC from shotgun metagenomic data from rhizosphere samples of three crops: corn (annual), and Miscanthus and switchgrass (both perennials). Both protein-coding genes separated all three communities whereas the SSU rRNA gene could only separate the annual from the two perennial communities in ordination analyses. Furthermore, assembled rplB and rpsC yielded significantly higher numbers of OTUs (alpha diversity) than the SSU rRNA gene. These results confirm these faster evolving marker genes offer increased resolution of for comparative microbiome studies.

microbiology

Insights into the evolution of oxygenic photosynthesis from a phylogenetically novel, low-light cyanobacterium

Atmospheric oxygen level rose dramatically around 2.4 billion years ago due to oxygenic photosynthesis by the Cyanobacteria. The oxidation of surface environments permanently changed the future of life on Earth, yet the evolutionary processes leading to oxygen production are poorly constrained. Partial records of these evolutionary steps are preserved in the genomes of organisms phylogenetically placed between non-photosynthetic Melainabacteria, crown-group Cyanobacteria, and Gloeobacter, representing the earliest-branching Cyanobacteria capable of oxygenic photosynthesis. Here, we describe nearly complete, metagenome assembled genomes of an uncultured organism phylogenetically placed between the Melainabacteria and crown-group Cyanobacteria, for which we propose the name Candidatus Aurora vandensis {au.rora Latin noun dawn and vand.ensis, originating from Vanda}.\n\nThe metagenome assembled genome of A. vandensis contains homologs of most genes necessary for oxygenic photosynthesis including key reaction center proteins. Many extrinsic proteins associated with the photosystems in other species are, however, missing or poorly conserved. The assembled genome also lacks homologs of genes associated with the pigments phycocyanoerethrin, phycoeretherin and several structural parts of the phycobilisome. Based on the content of the genome, we propose an evolutionary model for increasing efficiency of oxygenic photosynthesis through the evolution of extrinsic proteins to stabilize photosystem II and I reaction centers and improve photon capture. This model suggests that the evolution of oxygenic photosynthesis may have significantly preceded oxidation of Earths atmosphere due to low net oxygen production by early Cyanobacteria.

microbiology

Re-assembly, quality evaluation, and annotation of 678 microbial eukaryotic reference transcriptomes

BackgroundDe novo transcriptome assemblies are required prior to analyzing RNAseq data from a species without an existing reference genome or transcriptome. Despite the prevalence of transcriptomic studies, the effects of using different workflows, or \"pipelines\", on the resulting assemblies are poorly understood. Here, a pipeline was programmatically automated and used to assemble and annotate raw transcriptomic short read data collected by the Marine Microbial Eukaryotic Transcriptome Sequencing Project (MMETSP). The resulting transcriptome assemblies were evaluated and compared against assemblies that were previously generated with a different pipeline developed by the National Center for Genome Research (NCGR).\n\nResultsNew transcriptome assemblies contained the majority of previous contigs as well as new content. On average, 7.8% of the annotated contigs in the new assemblies were novel gene names not found in the previous assemblies. Taxonomic trends were observed in the assembly metrics, with assemblies from the Dinoflagellata and Ciliophora phyla showing a higher percentage of open reading frames and number of contigs than transcriptomes from other phyla.\n\nConclusionsGiven current bioinformatics approaches, there is no single best reference transcriptome for a particular set of raw data. As the optimum transcriptome is a moving target, improving (or not) with new tools and approaches, automated and programmable pipelines are invaluable for managing the computationally-intensive tasks required for re-processing large sets of samples with revised pipelines and ensuring a common evaluation workflow is applied to all samples. Thus, re-assembling existing data with new tools using automated and programmable pipelines may yield more accurate identification of taxon-specific trends across samples in addition to novel and useful products for the community.\n\nKey PointsO_LIRe-assembly with new tools can yield new results\nC_LIO_LIAutomated and programmable pipelines can be used to process arbitrarily many samples.\nC_LIO_LIAnalyzing many samples using a common pipeline identifies taxon-specific trends.\nC_LI

bioinformatics

Hospitalized premature infants are colonized by related bacterial strains with distinct proteomic profiles

During the first weeks of life, microbial colonization of the gut impacts human immune system maturation and other developmental processes. In premature infants, aberrant colonization has been implicated in the onset of necrotizing enterocolitis (NEC), a life-threatening intestinal disease. To study the premature infant gut colonization process, genome-resolved metagenomics was conducted on 343 fecal samples collected during the first three months of life from 35 premature infants housed in a neonatal intensive care unit, 14 of which developed NEC, and metaproteomic measurements were made on 87 samples. Microbial community composition and proteomic profiles remained relatively stable on the time scale of a week, but the proteome was more variable. Although genetically similar organisms colonized many infants, most infants were colonized by distinct strains with metabolic profiles that could be distinguished using metaproteomics. Microbiome composition correlated with infant, antibiotics administration, and NEC diagnosis. Communities were found to cluster into seven primary types, and community type switched within infants, sometimes multiple times. Interestingly, some communities sampled from the same infant at subsequent time points clustered with those of other infants. In some cases, switches preceded onset of NEC; however, no species or community type could account for NEC across the majority of infants. In addition to a correlation of protein abundances with organism replication rates, we found that organism proteomes correlated with overall community composition. Thus, this genome-resolved proteomics study demonstrates that the contributions of individual organisms to microbiome development depend on microbial community context.\n\nImportance\n\nHumans are colonized by microbes at birth, a process that is important to health and development. However, much remains to be known about the fine-scale microbial dynamics that occur during the colonization period. We conducted a genome-resolved study of microbial community composition, replication rates, and proteomes during the first three months of life of both healthy and sick premature infants. Infants were found to be colonized by similar microbes, but each underwent a distinct colonization trajectory.\n\nInterestingly, related microbes colonizing different infants were found to have distinct proteomes, indicating that microbiome function is not only driven by which organisms are present, but also largely depends on microbial responses to the unique set of physiological conditions in the infant gut.

microbiology

Evaluating Metagenome Assembly on a Simple Defined Community with Many Strain Variants

We evaluate the performance of three metagenome assemblers, IDBA, MetaSPAdes, and MEGAHIT, on short-read sequencing of a defined \"mock\" community containing 64 genomes (Shakya et al. (2013)). We update the reference metagenome for this mock community and detect several additional genomes in the read data set. We show that strain confusion results in significant loss in assembly of reference genomes that are otherwise completely present in the read data set. In agreement with previous studies, we find that MEGAHIT performs best computationally; we also show that MEGAHIT tends to recover larger portions of the strain variants than the other assemblers.

bioinformatics

FGF4 Retrogene On CFA12 Is Responsible For Chondrodystrophy And Intervertebral Disc Disease In Dogs

Chondrodystrophy in dogs is defined by dysplastic, shortened long bones and premature degeneration and calcification of intervertebral discs. Independent genome-wide association analyses for skeletal dysplasia (short limbs) within a single breed (pBonferroni=0.0072) and intervertebral disc disease (IVDD) across breeds (pBonferroni=4.02x10-10) both identified a significant association to the same region on CFA12. Whole genome sequencing identified a highly expressed FGF4 retrogene within this shared region. The FGF4 retrogene segregated with limb length and had an odds ratio of 51.23 (95% CI = 46.69, 56.20) for IVDD. Long bone length in dogs is a unique example of multiple disease-causing retrocopies of the same parental gene in a mammalian species. FGF signaling abnormalities have been associated with skeletal dysplasia in humans, and our findings present opportunities for both selective elimination of a medically and financially devastating disease in dogs and further understanding of the ever-growing complexity of retrogene biology.

genetics

dRep: A tool for fast and accurate genome de-replication that enables tracking of microbial genotypes and improved genome recovery from metagenomes

The number of microbial genomes sequenced each year is expanding rapidly, in part due to genome-resolved metagenomic studies that routinely recover hundreds of draft-quality genomes. Rapid algorithms have been developed to comprehensively compare large genome sets, but they are not accurate with draft-quality genomes. Here we present dRep, a program that sequentially applies a fast, inaccurate estimation of genome distance and a slow but accurate measure of average nucleotide identity to reduce the computational time for pair-wise genome set comparisons by orders of magnitude. We demonstrate its use in a study where we separately assembled each metagenome from time series datasets. Groups of essentially identical genomes were identified with dRep, and the best genome from each set was selected. This resulted in recovery of significantly more and higher-quality genomes compared to the set recovered using the typical co-assembly method. Documentation is available at http://drep.readthedocs.io/en/master/ and source code is available at https://github.com/MrOlm/drep.

bioinformatics