bioRxiv Science⌕ Search

Biology subjects

Freel, E. B.

Publications and source records attributed to Freel, E. B..

3 recordsLinked to original sources

A high-resolution diel survey of surface ocean metagenomes, metatranscriptomes, and transfer RNA transcripts

The roles of marine microbes in ecosystem processes are inherently linked to their ability to sense, respond, and ultimately adapt to environmental change. Capturing the nuances of this perpetual dialogue and its long-term implications requires insight into the subtle drivers of microbial responses to environmental change that are most accessible at the shortest scales of time. Here, we present a multi-omics dataset comprising surface ocean metagenomes, metatranscriptomes, tRNA transcripts, and biogeochemical measurements, collected every 1.5 hours for 48 hours at two stations within coastal and adjacent offshore waters of the tropical Pacific Ocean. We expect that this integrated dataset of multiple sequence types and environmental parameters will facilitate novel insights into microbial ecology, microbial physiology, and ocean biogeochemistry and help investigate the different mechanisms of adaptation that drive microbial responses to environmental change.

microbiology↗

New isolate genomes and global marine metagenomes resolve ecologically relevant units of SAR11

The bacterial order Pelagibacterales (SAR11) is widely distributed across the global surface ocean, where its activities are integral to the marine carbon cycle. High-quality genomes from isolates that can be propagated and phenotyped are needed to unify perspectives on the ecology and evolution of this complex group. Here, we increase the number of complete SAR11 isolate genomes threefold by describing 81 new SAR11 strains from coastal and offshore surface seawater of the tropical Pacific Ocean. Our analyses of the genomes and their spatiotemporal distributions support the existence of 29 monophyletic, discrete Pelagibacterales ecotypes that we define as genera. The spatiotemporal distributions of genomes within genera were correlated at fine scales with variation in ecologically-relevant gene content, supporting generic assignments and providing indications of speciation. We provide a hierarchical system of classification for SAR11 populations that is meaningfully correlated with evolution and ecology, providing a valid and utilitarian systematic nomenclature for this clade.

microbiology↗

assessPool: a fexible pipeline for population genomic analyses of pooled sequencing data

Despite the dramatic decrease in high-throughput sequencing costs over time, sequencing the ideal number of individuals for population genetic inference remains prohibitively expensive. When research questions require only population-level resolution, pooling individual samples before sequencing (pool-seq) can substantially reduce costs while still providing allele frequencies of Single Nucleotide Polymorphisms (SNPs). However, analyzing pooled data is comparatively difficult and less standardized than individual-based analyses. Although several programs have been developed to handle pool-seq data, most require extensive formatting or programming skills to operate. Here we introduce assessPool, an open-source R and Bash pipeline for pool- seq analyses with a focus on population structure. AssessPool accepts a Variant-Call Format (VCF) file and a FASTA-formatted reference, providing a straightforward transition from commonly used pipelines such as Stacks or dDocent. AssessPool handles varying numbers of pools and utilizes PoPoolation2 to generate locus-by-locus pairwise FST values and associated Fisher T-test values as measures of population structure. Starting with a VCF file containing all identified SNPs, assessPool facilitates several key functionalities for population genetic analyses: i) filtering SNPs based on adjustable criteria with parameter suggestions for pool-seq data, ii) organizing data structures for analysis based on input pools, iii) creating customizable run scripts for FST calculations using PoPoolation2 and/or the {poolfstat} R package, for all pairwise comparisons, iv) calculating locus-specific FST values using PoPoolation2 and/or {poolfstat}, v) importing FST output into a format compatible with R, vi) producing population genomic summary statistics, and vii) generating interactive plots to visualize and explore data. A pooled dataset generated from wild populations is used here to showcase the features of the assessPool pipeline for population genomic analyses.

bioinformatics↗