bioRxiv ScienceSearch

Biology subjects

Arkin, A. P.

Publications and source records attributed to Arkin, A. P..

16 recordsLinked to original sources

Massively parallel fitness profiling reveals multiple novel enzymes in Pseudomonas putida lysine metabolism

The lysine metabolism of Pseudomonas putida can produce multiple important commodity chemicals and is implicated in rhizosphere colonization. However, despite intensive study, the biochemical and genetic links between lysine metabolism and central metabolism remain unresolved in P. putida. Here, we leverage Random Barcode Transposon Sequencing (RB-TnSeq), a genome-wide assay measuring the fitness of thousands of genes in parallel, to identify multiple novel enzymes in both L- and D-lysine metabolism. We first describe three pathway enzymes that catabolize 2-aminoadipate (2AA) to 2-ketoglutarate (2KG) connecting D-lysine to the TCA cycle. One of these enzymes, PP_5260, contains a DUF1338 domain, a family without a previously described biological function. We demonstrate PP_5260 converts 2-oxoadipate (2OA) to 2-hydroxyglutarate (2HG), a novel biochemical reaction. We expand on recent work showing that the glutarate hydroxylase, CsiD, can co-utilize both 2OA and 2KG as a co-substrate in the hydroxylation of glutarate. Finally we demonstrate that the cellular abundance of D- and L-lysine pathway proteins are highly sensitive to pathway specific substrates. This work demonstrates the utility of RB-TnSeq for discovering novel metabolic pathways in even well-studied bacteria.\n\nImportanceO_LIP. putida is an attractive host for metabolic engineering as its lysine metabolism can be utilized for the production of multiple important commodity chemicals.\nC_LIO_LIWe demonstrate the first biochemical evidence of a bacterial 2OA catabolic pathway to central metabolites.\nC_LIO_LIDUF1338 proteins are widely dispersed across many kingdoms of life. Here we demonstrate the first biochemical evidence of function for a member of this protein family.\nC_LI

microbiology

A versatile platform strain for high-fidelity multiplex genome editing

Precision genome editing accelerates the discovery of the genetic determinants of phenotype and the engineering of novel behaviors in organisms. Advances in DNA synthesis and recombineering have enabled high-throughput engineering of genetic circuits and biosynthetic pathways via directed mutagenesis of bacterial chromosomes. However, the highest recombination efficiencies have to date been reported in persistent mutator strains, which suffer from reduced genomic fidelity. The absence of inducible transcriptional regulators in these strains also prevents concurrent control of genome engineering tools and engineered functions. Here, we introduce a new recombineering platform strain, BioDesignER, which incorporates (1) a refactored {lambda}-Red recombination system that reduces toxicity and accelerates multi-cycle recombination, (2) genetic modifications that boost recombination efficiency, and (3) four independent inducible regulators to control engineered functions. These modifications resulted in single-cycle recombineering efficiencies of up to 25% with a seven-fold increase in recombineering fidelity compared to the widely used recombineering strain EcNR2. To facilitate genome engineering in BioDesignER, we have curated eight context-neutral genomic loci, termed Safe Sites, for stable gene expression and consistent recombination efficiency. BioDesignER is a platform to develop and optimize engineered cellular functions and can serve as a model to implement comparable recombination and regulatory systems in other bacteria.

synthetic biology

Defining the toxicity limits on microbial range in a metal-contaminated aquifer

In extreme environments, toxic compounds restrict which microorganisms persist. However, in complex mixtures of inhibitory compounds, it is challenging to determine which specific compounds cause changes in abundance and prevent some microorganisms from growing. We focused on a contaminated aquifer in Oak Ridge, Tennessee, U.S.A. that has low pH and high concentrations of uranium, nitrate and many other inorganic ions. In the most contaminated wells, the microbial community is enriched in the Rhodanobacter genus. Rhodanobacter relative abundance is positively correlated with low pH and high concentrations of U, Mn, Al, Cd, Zn, Ni, Co, Ca, NO3-, Mg, Cl, SO42-, Sr, K and Ba and we sought to determine which of these correlated parameters are selective pressures that favor the growth of Rhodanobacter over other taxa. Using high-throughput cultivation, we determined that of the ions correlated high Rhodanobacter abundance, only low pH and high U, Mn, Al, Cd, Zn, Co and Ni (a) are selectively inhibitory of a sensitive Pseudomonas isolate from a background well versus a representative resistant Rhodanobacter isolate from a contaminated well, and (b) reach toxic concentrations in the most contaminated wells that can inhibit the sensitive Pseudomonas isolate. We prepared mixtures of inorganic ions representative of the most contaminated wells and verified that few other isolates aside from Rhodanobacter can tolerate these 8 parameters. These results clarify which toxic inorganic ions are causal factors that impact the microbial community at this field site and are not merely correlated with taxonomic shifts.

microbiology

Dub-seq: dual-barcoded shotgun expression library sequencing for high-throughput characterization of functional traits

A major challenge in genomics is the knowledge gap between sequence and its encoded function. Gain-of-function methods based on gene overexpression are attractive avenues for phenotype-based functional screens, but are not easily applied in high-throughput across many experimental conditions. Here, we present Dual Barcoded Shotgun Expression Library Sequencing (Dub-seq), a method that greatly increases the throughput of genome-wide overexpression assays. In Dub-seq, a shotgun expression library is cloned between dual random DNA barcodes and the precise breakpoints of DNA fragments are associated to the barcode sequences prior to performing assays. To assess the fitness of individual strains carrying these plasmids, we use DNA barcode sequencing (BarSeq), which is amenable to large-scale sample multiplexing. As a demonstration of this approach, we constructed a Dub-seq library with total Escherichia coli genomic DNA, performed 155 genome-wide fitness assays in 52 experimental conditions, and identified 813 genes with high-confidence overexpression phenotypes across 4,151 genes assayed. We show that Dub-seq data is reproducible, accurately recapitulates known biology, and identifies hundreds of novel gain-of-function phenotypes for E. coli genes, a subset of which we verified with assays of individual strains. Dub-seq provides complementary information to loss-of-function approaches such as transposon site sequencing or CRISPRi and will facilitate rapid and systematic functional characterization of microbial genomes.\n\nImportanceMeasuring the phenotypic consequences of overexpressing genes is a classic genetic approach for understanding protein function; for identifying drug targets, antibiotic and metal resistance mechanisms; and for optimizing strains for metabolic engineering. In microorganisms, these gain-of-function assays are typically done using laborious protocols with individually archived strains or in low-throughput following qualitative selection for a phenotype of interest, such as antibiotic resistance. However, many microbial genes are poorly characterized and the importance of a given gene may only be apparent under certain conditions. Therefore, more scalable approaches for gain-of-function assays are needed. Here, we present Dual Barcoded Shotgun Expression Library Sequencing (Dub-seq), a strategy that couples systematic gene overexpression with DNA barcode sequencing for large-scale interrogation of gene fitness under many experimental conditions at low cost. Dub-seq can be applied to many microorganisms and is a valuable new tool for large-scale gene function characterization.

microbiology

Kluyveromyces marxianus as a robust synthetic biology platform host

Throughout history, the yeast Saccharomyces cerevisiae has played a central role in human society due to its use in food production and more recently as a major industrial and model microorganism, because of the many genetic and genomic tools available to probe its biology. However S. cerevisiae has proven difficult to engineer to expand the carbon sources it can utilize, the products it can make, and the harsh conditions it can tolerate in industrial applications. Other yeasts that could solve many of these problems remain difficult to manipulate genetically. Here, we engineer the thermotolerant yeast Kluyveromyces marxianus to create a new synthetic biology platform. Using CRISPR-Cas9 mediated genome editing, we show that wild isolates of K. marxianus can be made heterothallic for sexual crossing. By breeding two of these mating-type engineered K. marxianus strains, we combined three complex traits- thermotolerance, lipid production, and facile transformation with exogenous DNA-into a single host. The ability to cross K. marxianus strains with relative ease, together with CRISPR-Cas9 genome editing, should enable engineering of K. marxianus isolates with promising lipid production at temperatures far exceeding those of other fungi under development for industrial applications. These results establish K. marxianus as a synthetic biology platform comparable to S. cerevisiae, with naturally more robust traits that hold potential for the industrial production of renewable chemicals.

synthetic biology

Deciphering microbial interactions in synthetic human gut microbiome communities

The human gut microbiota comprises a dynamic ecological system that contributes significantly to human health and disease. The ecological forces that govern community assembly and stability in the gut microbiota remain unresolved. We developed a generalizable model-guided framework to predict higher-order consortia from time-resolved measurements of lower-order assemblages. This method was employed to decipher microbial interactions in a diverse 12-member human gut microbiome synthetic community. We show that microbial growth parameters and pairwise interactions are the major drivers of multi-species community dynamics, as opposed to context-dependent (conditional) interactions. The inferred microbial interaction network as well as a top-down approach to community assembly pinpointed both ecological driver and responsive species that were significantly modulated by microbial inter-relationships. Our model demonstrated that negative pairwise interactions could generate history-dependent responses of initial species proportions on physiological timescales that frequently does not originate from bistability. The model elucidated a topology for robust coexistence in pairwise assemblages consisting of a negative feedback loop that balances disparities in monospecies fitness levels. Bayesian statistical methods were used to evaluate the constraint of model parameters by the experimental data. Measurements of extracellular metabolites illuminated the metabolic capabilities of monospecies and potential molecular basis for competitive and cooperative interactions in the community. However, these data failed to predict influential organisms shaping community assembly. In sum, these methods defined the ecological roles of key species shaping community assembly and illuminated network design principles of microbial communities.

systems biology

The polygenic basis of an ancient divergence in yeast thermotolerance

Some of the most unique and compelling survival strategies in the natural world are fixed in isolated species. To date, molecular insight into these ancient adaptations has been limited, as classic experimental genetics has focused on interfertile individuals in populations. Here we use a new mapping approach, which screens mutants in a sterile interspecific hybrid, to identify eight housekeeping genes that underlie the growth advantage of Saccharomyces cerevisiae over its distant relative S. paradoxus at high temperature. Pro-thermotolerance alleles at these mapped loci were required for the adaptive trait in S. cerevisiae and sufficient for its partial reconstruction in S. paradoxus. The emerging picture is one in which S. cerevisiae improved the heat resistance of multiple components of the fundamental growth machinery in response to selective pressure. This study lays the groundwork for the mapping of genotype to phenotype in clades of sister species across Eukarya.

evolutionary biology

Massive Factorial Design Untangles Coding Sequences Determinants Of Translation Efficacy

Comparative analyses of natural sequences or variant libraries are often used to infer mechanisms of expression, activity and evolution. Contingent selective histories and small sample sizes can profoundly bias such approaches. Both limitations can be lifted using precise design of large-scale DNA synthesis. Here, we precisely design 5 E. coli genomes worth of synthetic DNA to untangle the relative contributions of 8 interlaced sequence properties described independently as major determinants of translation in Escherichia coli. To expose hierarchical effects, we engineer an inducible translational coupling device enabling epigenetic disruption of mRNA secondary structures. We find that properties commonly believed to modulate translation generally explain less than a third of the variation in protein production. We describe dominant effects of mRNA structures over codon composition on both initiation and elongation, and previously uncharacterized relationships among factors controlling translation. These results advance our understanding of translation efficiency and expose critical design challenges.

synthetic biology

Massive Phenotypic Measurements Reveal Complex Physiological Consequences Of Differential Translation Efficacies

AO_SCPCAPBSTRACTC_SCPCAPControl of protein biosynthesis is at the heart of resource allocation and cell adaptation to fluctuating environments. One genes translation often occurs at the expense of anothers, resulting in global energetic and fitness trade-offs during differential expression of various functions. Patterns of ribosome utilization--as controlled by initiation, elongation and release rates--are central to this balance. To disentangle their respective determinants and physiological impacts, we complemented measurements of protein production with highly parallelized quantifications of transcripts abundance and decay, ribosome loading and cellular growth rate for 244,000 precisely designed sequence variants of an otherwise standard reporter. We find highly constrained, non-monotonic relationships between measured phenotypes. We show that fitness defects derive either from protein overproduction, with efficient translation initiation and heavy ribosome flows; or from unproductive ribosome sequestration by highly structured, slowly initiated and overly stabilized transcripts. These observations demonstrate physiological impacts of key sequence features in natural and designed transcripts.

synthetic biology

An Oxidative Pathway of Deoxyribose Catabolism

Using genome-wide mutant fitness assays in diverse bacteria, we identified novel oxidative pathways for the catabolism of 2-deoxy-D-ribose and 2-deoxy-D-ribonate. We propose that deoxyribose is oxidized to deoxyribonate, oxidized to ketodeoxyribonate, and cleaved to acetyl-CoA and glyceryl-CoA. We have genetic evidence for this pathway in three genera of bacteria, and we confirmed the oxidation of deoxyribose to ketodeoxyribonate in vitro. In Pseudomonas simiae, the expression of enzymes in the pathway is induced by deoxyribose or deoxyribonate, while in Paraburkholderia bryophila and in Burkholderia phytofirmans, the pathway proceeds in parallel with the known deoxyribose 5-phosphate aldolase pathway. We identified another oxidative pathway for the catabolism of deoxyribonate, with acyl-CoA intermediates, in Klebsiella michiganensis. Of these four bacteria, only P. simiae relies entirely on an oxidative pathway to consume deoxyribose. The deoxyribose dehydrogenase of P. simiae is either non-specific or evolved recently, as this enzyme is very similar to a novel vanillin dehydrogenase from Pseudomonas putida that we identified. So, we propose that these oxidative pathways evolved primarily to consume deoxyribonate, which is a waste product of metabolism.\n\nImportanceDeoxyribose is one of the building blocks of DNA and is released when cells die and their DNA degrades. We identified a bacterium that can grow with deoxyribose as its sole source of carbon even though its genome does not encode any of the known genes for breaking down deoxyribose. By growing many mutants of this bacterium together on deoxyribose and using DNA sequencing to measure the change in the mutants abundance, we identified multiple protein-coding genes that are required for growth on deoxyribose. Based on the similarity of these proteins to enzymes of known function, we propose a 6-step pathway in which deoxyribose is oxidized and then cleaved. Diverse bacteria use a portion of this pathway to break down a related compound, deoxyribonate, which is a waste product of human metabolism and is present in urine. Our study illustrates the utility of large-scale bacterial genetics to identify previously unknown metabolic pathways.

microbiology

An interventional Soylent diet increases the Bacteroidetes to Firmicutes ratio in human gut microbiome communities: a randomized controlled trial

Our knowledge of the relationship between the gut microbiome and health has rapidly expanded in recent years. Diet has been shown to have causative effects on microbiome composition, which can have subsequent implications on health. Soylent 2.0 is a liquid meal replacement drink that satisfies nearly 20% of all recommended daily intakes per serving. This study aims to characterize the changes in gut microbiota composition resulting from a short-term Soylent diet. Fourteen participants were separated into two groups: 5 in the regular diet group and 9 in the Soylent diet group. The regular diet group maintained a diet closely resembling self-reported regular diets. The Soylent diet group underwent three dietary phases: A) a regular diet for 2 days, B) a Soylent-only diet (five servings of Soylent daily and water as needed) for 4 days, and C) a regular diet for 4 days. Daily logs self-reporting diet, Bristol stool ratings, and any abdominal discomfort were electronically submitted. Eight fecal samples per participant were collected using fecal sampling kits, which were subsequently sent to uBiome, Inc. for sample processing and V4 16S rDNA sequencing. Reads were clustered into operational taxonomic units (OTUs) and taxonomically identified against the GreenGenes 16S database. We find that an individuals alpha-diversity is not significantly altered during a Soylent-only diet. In addition, principal coordinate analysis using the unweighted UniFrac distance metric shows samples cluster strongly by individual and not by dietary phase. Among Soylent dieters, we find a significant increase in the ratio of Bacteroidetes to Firmicutes abundance, which is associated with several positive health outcomes, including reduced risks of obesity and intestinal inflammation.

clinical trials

Filling Gaps in Bacterial Amino Acid Biosynthesis Pathways with High-throughput Genetics

For many bacteria with sequenced genomes, we do not understand how they synthesize some amino acids. This makes it challenging to reconstruct their metabolism, and has led to speculation that bacteria might be cross-feeding amino acids. We studied heterotrophic bacteria from 10 different genera that grow without added amino acids even though an automated tool predicts that the bacteria have gaps in their amino acid synthesis pathways. Across these bacteria, there were 11 gaps in their amino acid biosynthesis pathways that we could not fill using current knowledge. Using genome-wide mutant fitness data, we identified novel enzymes that fill 9 of the 11 gaps and hence explain the biosynthesis of methionine, threonine, serine, or histidine by bacteria from six genera. We also found that the sulfate-reducing bacterium Desulfovibrio vulgaris synthesizes homocysteine (which is a precursor to methionine) by using DUF39, NIL/ferredoxin, and COG2122 proteins, and that homoserine is not an intermediate in this pathway. Our results suggest that most free-living bacteria can likely make all 20 amino acids and illustrate how high-throughput genetics can uncover previously-unknown amino acid biosynthesis genes.

microbiology

Magic pools: parallel assessment of transposon delivery vectors in bacteria

Transposon mutagenesis coupled to next-generation sequencing (TnSeq) is a powerful approach for discovering the functions of bacterial genes. However, the development of a suitable TnSeq strategy for a given bacterium can be costly and time-consuming. To meet this challenge, we describe a parts-based strategy for constructing libraries of hundreds of transposon delivery vectors, which we term \"magic pools\". Within a magic pool, each transposon vector has a different combination of promoters and antibiotic resistance markers as well as a random DNA barcode sequence, which allows the tracking of each vector during mutagenesis experiments. To identify an efficient vector for a given bacterium, we mutagenize it with a magic pool and sequence the resulting insertions; we then use the best vector to generate a large mutant library. We used the magic pool strategy to construct transposon mutant libraries in five genera of bacteria, including three genera of the phylum Bacteroidetes.

genomics

PaperBLAST: Text-mining papers for information about homologs

Large-scale genome sequencing has identified millions of protein-coding genes whose function is unknown. Many of these proteins are similar to characterized proteins from other organisms, but much of this information is missing from annotation databases and is hidden in the scientific literature. To make this information accessible, PaperBLAST uses EuropePMC to search the full text of scientific articles for references to genes. PaperBLAST also takes advantage of curated resources that link protein sequences to scientific articles (Swiss-Prot, GeneRIF, and EcoCyc). PaperBLASTs database includes over 700,000 scientific articles that mention over 400,000 different proteins. Given a protein of interest, PaperBLAST quickly finds similar proteins that are discussed in the literature and presents snippets of text from relevant articles or from the curators. PaperBLAST is available at http://papers.genomics.lbl.gov/.

bioinformatics

The DOE Systems Biology Knowledgebase (KBase)

The U.S. Department of Energy Systems Biology Knowledgebase (KBase) is an open-source software and data platform designed to meet the grand challenge of systems biology -- predicting and designing biological function from the biomolecular (small scale) to the ecological (large scale). KBase is available for anyone to use, and enables researchers to collaboratively generate, test, compare, and share hypotheses about biological functions; perform large-scale analyses on scalable computing infrastructure; and combine experimental evidence and conclusions that lead to accurate models of plant and microbial physiology and community dynamics. The KBase platform has (1) extensible analytical capabilities that currently include genome assembly, annotation, ontology assignment, comparative genomics, transcriptomics, and metabolic modeling; (2) a web-browser-based user interface that supports building, sharing, and publishing reproducible and well-annotated analyses with integrated data; (3) access to extensive computational resources; and (4) a software development kit allowing the community to add functionality to the system.

bioinformatics

Validating Regulatory Predictions from Diverse Bacteria with Mutant Fitness Data

Although transcriptional regulation is fundamental to understanding bacterial physiology, the targets of most bacterial transcription factors are not known. Comparative genomics has been used to identify likely targets of some of these transcription factors, but these predictions typically lack experimental support. Here, we used mutant fitness data, which measures the importance of each gene for a bacteriums growth across many conditions, to validate regulatory predictions from RegPrecise, a curated collection of comparative genomics predictions. Because characterized transcription factors often have correlated fitness with one of their targets (either positively or negatively), correlated fitness patterns provide support for the comparative genomics predictions. At a false discovery rate of 3%, we identified significant cofitness for at least one target of 158 TFs in 107 ortholog groups and from 24 bacteria. Thus, high-throughput genetics can be used to identify a high-confidence subset of the sequence-based regulatory predictions.

microbiology