bioRxiv ScienceSearch

Biology subjects

Katherine S Pollard

Publications and source records attributed to Katherine S Pollard.

5 recordsLinked to original sources

Proteobacteria drive significant functional variability in the human gut microbiome

While human gut microbiomes vary significantly in taxonomic composition, biological pathway abundance is surprisingly invariable across hosts. We hypothesized that healthy microbiomes appear functionally redundant due to factors that obscure differences in gene abundance across hosts. To account for these biases, we developed a powerful test of gene variability, applicable to shotgun metagenomes from any environment. Our analysis of healthy stool metagenomes reveals thousands of genes whose abundance differs signifi-cantly between people consistently across studies, including glycolytic enzymes, lipopolysac-charide biosynthetic genes, and secretion systems. Even housekeeping pathways contain a mix of variable and invariable genes, though most deeply conserved genes are significantly invariable. Variable genes tend to be associated with Proteobacteria, as opposed to taxa used to define enterotypes or the dominant phyla Bacteroidetes and Firmicutes. These re-sults establish limits on functional redundancy and predict specific genes and taxa that may drive physiological differences between gut microbiomes.\n\nImpact StatementA statistical test for gene variability reveals extensive functional differences between healthy humanmicrobiomes.

Bioinformatics

Features of ChIP-seq data peak calling algorithms with good operating characteristics

Author descriptionReuben Thomas is a Staff Research Scientist in the Bioinformatics Core at Gladstone Institutes\n\nSean Thomas is a Staff Research Scientist in the Bioinformatics Core at Gladstone Institutes\n\nAlisha K Holloway is the Director of Bioinformatics at Phylos Biosciences, visiting scientist at Gladstone Institutes and Adjunct Assistant Professor in Biostatistics at the University of California, San Francisco.\n\nKatherine S Pollard is a Senior Investigator at Gladstone Institutes and Professor of Biostatistics at University of California, San Francisco.\n\nKey PointsO_LIPeak-calling using Chip-seq data consists of two sub-problems: identifying candidate peaks and testing candidate peaks for statistical significance.\nC_LIO_LITwelve features of the two sub-problems of peak-calling methods are identified.\nC_LIO_LIMethods that explicitly combine the signals from ChIP and input samples are less powerful than methods that do not.\nC_LIO_LIMethods that use windows of different sizes to scan the genome for potential peaks are more powerful than ones that do not.\nC_LIO_LIMethods that use a Poisson test to rank their candidate peaks are more powerful than those that use a Binomial test.\nC_LI\n\nAbstractChromatin immunoprecipitation followed by sequencing (ChIP-seq) is an important tool for studying gene regulatory proteins, such as transcription factors and histones. Peak calling is one of the first steps in analysis of these data. Peak-calling consists of two sub-problems: identifying candidate peaks and testing candidate peaks for statistical significance. We surveyed 30 methods and identified 12 features of the two sub-problems that distinguish methods from each other. We picked six methods (GEM, MACS2, MUSIC, BCP, TM and ZINBA) that span this feature space and used a combination of 300 simulated ChIP-seq data sets, 3 real data sets and mathematical analyses to identify features of methods that allow some to perform better than others. We prove that methods that explicitly combine the signals from ChIP and input samples are less powerful than methods that do not. Methods that use windows of different sizes are more powerful than ones that do not. For statistical testing of candidate peaks, methods that use a Poisson test to rank their candidate peaks are more powerful than those that use a Binomial test. BCP and MACS2 have the best operating characteristics on simulated transcription factor binding data. GEM has the highest fraction of the top 500 peaks containing the binding motif of the immunoprecipitated factor, with 50% of its peaks within 10 base pairs (bp) of a motif. BCP and MUSIC perform best on histone data. These findings provide guidance and rationale for selecting the best peak caller for a given application.

Bioinformatics

Bat Accelerated Regions Identify a Bat Forelimb Specific Enhancer in the HoxD Locus

The molecular events leading to the development of the bat wing remain largely unknown, and are thought to be caused, in part, by changes in gene expression during limb development. These expression changes could be instigated by variations in gene regulatory enhancers. Here, we used a comparative genomics approach to identify regions that evolved rapidly in the bat ancestor but are highly conserved in other vertebrates. We discovered 166 bat accelerated regions (BARs) that overlap H3K27ac and p300 ChIP-seq peaks in developing mouse limbs. Using a mouse enhancer assay, we show that five Myotis lucifugus BARs drive gene expression in the developing mouse limb, with the majority showing differential enhancer activity compared to the mouse orthologous BAR sequences. These include BAR116, which is located telomeric to the HoxD cluster and had robust forelimb expression for the M. lucifugus sequence and no activity for the mouse sequence at embryonic day 12.5. Developing limb expression analysis of Hoxd10-Hoxd13 in Miniopterus natalensis bats showed a high-forelimb weak-hindlimb expression for Hoxd10-Hoxd11, similar to the expression trend observed for M. lucifugus BAR116 in mice, suggesting that it could be involved in the regulation of the bat HoxD complex. Combined, our results highlight novel regulatory regions that could be instrumental for the morphological differences leading to the development of the bat wing.\n\nAuthor SummaryThe limb is a classic example of vertebrate homology and is represented by a large range of morphological structures such as fins, legs and wings. The evolution of these structures could be driven by alterations in gene regulatory elements that have critical roles during development. To identify elements that may contribute to bat wing development, we characterized sequences that are conserved between vertebrates, but changed significantly in the bat lineage. We then overlapped these sequences with predicted developing limb enhancers as determined by ChIP-seq, finding 166 bat accelerated sequences (BARs). Five BARs that were tested for enhancer activity in mice all drove expression in the limb. Testing the mouse orthologous sequence showed that three had differences in their limb enhancer activity as compared to the bat sequence. Of these, BAR116 was of particular interest as it is located near the HoxD locus, an essential gene complex required for proper spatiotemporal patterning of the developing limb. The bat BAR116 sequence drove robust forelimb expression but the mouse BAR116 sequence did not show enhancer activity. These experiments correspond to analyses of HoxD gene expressions in developing bat limbs, which had strong forelimb versus weak hindlimb expression for Hoxd10-11. Combined, our studies highlight specific genomic regions that could be important in shaping the morphological differences that led to the development of the bat wing.

Evolutionary Biology

An integrated metagenomics pipeline for strain profiling reveals novel patterns of transmission and global biogeography of bacteria

We present the Metagenomic Intra-species Diversity Analysis System (MIDAS), which is an integrated computational pipeline for quantifying bacterial species abundance and strain-level genomic variation, including gene content and single nucleotide polymorphisms, from shotgun metagenomes. Our method leverages a database of >30,000 bacterial reference genomes which we clustered into species groups. These cover the majority of abundant species in the human microbiome but only a small proportion of microbes in other environments, including soil and seawater. We applied MIDAS to stool metagenomes from 98 Swedish mothers and their infants over one year and used rare single nucleotide variants to reveal extensive vertical transmission of strains at birth but colonization with strains unlikely to derive from the mother at later time points. This pattern was missed with species-level analysis, because the infant gut microbiome composition converges towards that of an adult over time. We also applied MIDAS to 198 globally distributed marine metagenomes and used gene content to show that many prevalent bacterial species have population structure that correlates with geographic location. Strain-level genetic variants present in metagenomes clearly reveal extensive structure and dynamics that are obscured when data is analyzed at a higher taxonomic resolution.

Genomics

Average genome size estimation enables accurate quantification of gene family abundance and sheds light on the functional ecology of the human microbiome

Average genome size (AGS) is an important, yet often overlooked property of microbial communities. We developed MicrobeCensus to rapidly and accurately estimate AGS from short-read metagenomics data and applied our tool to over 1,300 human microbiome samples. We found that AGS differs significantly within and between body sites and tracks with major functional and taxonomic differences. For example, in the gut, AGS ranges from 2.5 to 5.8 megabases and is positively correlated with the abundance of Bacteroides and polysaccharide metabolism. Furthermore, we found that AGS variation can bias comparative analyses, and that normalization improves detection of differentially abundant genes.

Bioinformatics