bioRxiv ScienceSearch

Biology subjects

Beck, K. L.

Publications and source records attributed to Beck, K. L..

6 recordsLinked to original sources

Multi-omics profiling of Earth's biomes reveals that microbial and metabolite composition are shaped by the environment

As our understanding of the structure and diversity of the microbial world grows, interpreting its function is of critical interest for understanding and managing the many systems microbes influence. Despite advances in sequencing, lack of standardization challenges comparisons among studies that could provide insight into the structure and function of microbial communities across multiple habitats on a planetary scale. Technical variation among distinct studies without proper standardization of approaches prevents robust meta-analysis. Here, we present a multi-omics, meta-analysis of a novel, diverse set of microbial community samples collected for the Earth Microbiome Project. We include amplicon (16S, 18S, ITS) and shotgun metagenomic sequence data, and untargeted metabolomics data (liquid chromatography-tandem mass spectrometry and gas chromatography mass spectrometry), centering our description on relationships and co-occurrences of microbially-related metabolites and microbial taxa across environments. Standardized protocols and analytical methods for characterizing microbial communities, including assessment of molecular diversity using untargeted metabolomics, facilitate identification of shared microbial and metabolite features, permitting us to explore diversity at extraordinary scale. In addition to a reference database for metagenomic and metabolomic data, we provide a framework for incorporating additional studies, enabling the expansion of existing knowledge in the form of a community resource that will become more valuable with time. To provide examples of applying this database, we outline important ecological questions that can be addressed, and test the hypotheses that every microbe and metabolite is everywhere, but the environment selects. Our results show that metabolite diversity exhibits turnover and nestedness related to both microbial communities and the environment. The relative abundances of microbially-related metabolites vary and co-occur with specific microbial consortia in a habitat-specific manner, and highlight the power of certain chemistry - in particular terpenoids - in distinguishing Earths environments.

ecology

Semi-supervised identification of SARS-CoV-2 molecular targets

SARS-CoV-2 genomic sequencing efforts have scaled dramatically to address the current global pandemic and aid public health. In this work, we analyzed a corpus of 66,000 SARS-CoV-2 genome sequences. We developed a novel semi-supervised pipeline for automated gene, protein, and functional domain annotation of SARS-CoV-2 genomes that differentiates itself by not relying on use of a single reference genome and by overcoming atypical genome traits. Using this method, we identified the comprehensive set of known proteins with 98.5% set membership accuracy and 99.1% accuracy in length prediction compared to proteome references including Replicase polyprotein 1ab (with its transcriptional slippage site). Compared to other published tools such as Prokka (base) and VAPiD, we yielded an 6.4- and 1.8-fold increase in protein annotations. Our method generated 13,000,000 molecular target sequences-- some conserved across time and geography while others represent emerging variants. We observed 3,362 non-redundant sequences per protein on average within this corpus and describe key D614G and N501Y variants spatiotemporally. For spike glycoprotein domains, we achieved greater than 97.9% sequence identity to references and characterized Receptor Binding Domain variants. Here, we comprehensively present the molecular targets to refine biomedical interventions for SARS-CoV-2 with a scalable high-accuracy method to analyze newly sequenced infections.

bioinformatics

Analysis and Forecasting of Global of RT-PCR Primers for SARS-CoV-2

Rapid tests for active SARS-CoV-2 infections rely on reverse transcription polymerase chain reaction (RT-PCR). RT-PCR uses reverse transcription of RNA into complementary DNA (cDNA) and amplification of specific DNA (primer and probe) targets using polymerase chain reaction (PCR). The technology makes rapid and specific identification of the virus possible based on sequence homology of nucleic acid sequence and is much faster than tissue culture or animal cell models. However the technique can lose sensitivity over time as the virus evolves and the target sequences diverge from the selective primer sequences. Different primer sequences have been adopted in different geographic regions. As we rely on these existing RT-PCR primers to track and manage the spread of the Coronavirus, it is imperative to understand how SARS-CoV-2 mutations, over time and geographically, diverge from existing primers used today. In this study, we analyze the performance of the SARS-CoV-2 primers in use today by measuring the number of mismatches between primer sequence and genome targets over time and spatially. We find that there is a growing number of mismatches, an increase by 2% per month, as well as a high specificity of virus based on geographic location.

bioinformatics

EMPress enables tree-guided, interactive, and exploratory analyses of multi-omic datasets

Standard workflows for analyzing microbiomes often include the creation and curation of phylogenetic trees. Here we present EMPress, an interactive tool for visualizing trees in the context of microbiome, metabolome, etc. community data scalable beyond modern large datasets like the Earth Microbiome Project. EMPress provides novel functionality--including ordination integration and animations--alongside many standard tree visualization features, and thus simplifies exploratory analyses of many forms of omic data.

bioinformatics

DNA extraction and host depletion methods significantly impact and potentially bias bacterial detection in a biological fluid

Untargeted sequencing of nucleic acids present in food can inform the detection of food safety and origin, as well as product tampering and mislabeling issues. The application of such technologies to food analysis could reveal valuable insights that are simply unobtainable by targeted testing, leading to the efforts of applying such technologies in the food industry. However, before these approaches can be applied, it is imperative to verify that the most appropriate methods are used at every step of the process: gathering primary material, laboratory methods, data analysis, and interpretation. The focus of this study is in gathering the primary material, in this case, DNA. We used bovine milk as a model to 1) evaluate commercially available kits for their ability to extract nucleic acids from inoculated bovine milk; 2) evaluate host DNA depletion methods for use with milk, and 3) develop and evaluate a selective lysis-PMA based protocol for host DNA depletion in milk. Our results suggest that magnetic-based nucleic acid extraction methods are best for nucleic acid isolation of bovine milk. Removal of host DNA remains a challenge for untargeted sequencing of milk, highlighting that the individual matrix characteristics should always be considered in food testing. Some reported methods introduce bias against specific types of microbes, which may be particularly problematic in food safety where the detection of Gram-negative pathogens and indicators is essential. Continuous efforts are needed to develop and validate new approaches for untargeted metagenomics in samples with large amounts of DNA from a single host. ImportanceTracking the bacterial communities present in our food has the potential to inform food safety and product origin. To do so, the entire genetic material present in a sample is extracted using chemical methods or commercially available kits and sequenced using next-generation platforms to provide a snapshot of what the relative composition looks like. Because the genetic material of higher organisms present in food (e.g., cow in milk or beef, wheat in flour) is around one thousand times larger than the bacterial content, challenges exist in gathering the information of interest. Additionally, specific bacterial characteristics can make them easier or harder to detect, adding another layer of complexity to this issue. In this study, we demonstrate the impact of using different methods in the ability of detecting specific bacteria and highlight the need to ensure that the most appropriate methods are being used for each particular sample.

molecular biology

Monitoring the microbiome for food safety and quality using deep shotgun sequencing

In this work, we hypothesized that shifts in the food microbiome can be used as an indicator of unexpected contaminants or environmental changes. To test this hypothesis, we sequenced total RNA of 31 high protein powder (HPP) samples of poultry meal pet food ingredients. We developed a microbiome analysis pipeline employing a key eukaryotic matrix filtering step that improved microbe detection specificity to >99.96% during in silico validation. The pipeline identified 119 microbial genera per HPP sample on average with 65 genera present in all samples. The most abundant of these were Bacteroides, Clostridium, Lactococcus, Aeromonas, and Citrobacter. We also observed shifts in the microbial community corresponding to ingredient composition differences. When comparing culture-based results for Salmonella with total RNA sequencing, we found that Salmonella growth did not correlate with multiple sequence analyses. We conclude that microbiome sequencing is useful to characterize complex food microbial communities, while additional work is required for predicting specific species viability from total RNA sequencing.

microbiology