bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Germline Cas9 Expression Yields Highly Efficient Genome Engineering in a Major Worldwide Disease Vector, Aedes aegypti

The development of CRISPR/Cas9 technologies has dramatically increased the accessibility and efficiency of genome editing in many organisms. In general, in vivo germline expression of Cas9 results in substantially higher activity than embryonic injection. However, no transgenic lines expressing Cas9 have been developed for the major mosquito disease vector Aedes aegypti. Here, we describe the generation of multiple stable, transgenic Ae. aegypti strains expressing Cas9 in the germline, resulting in dramatic improvements in both the consistency and efficiency of genome modifications using CRISPR. Using these strains, we disrupted numerous genes important for normal morphological development, and even generated triple mutants from a single injection. We have also managed to increase the rates of homology directed repair by more than an order of magnitude. Given the exceptional mutagenic efficiency and specificity of the Cas9 strains we built, they can be used for high-throughput reverse genetic screens to help functionally annotate the Ae. aegypti genome. Additionally, these strains represent a first step towards the development of novel population control technologies targeting Ae. aegypti that rely on Cas9-based gene drives.\n\nSignificance StatementAedes aegypti is the principal vector of multiple arboviruses that significantly affect human health including dengue, chikungunya, and zika. Development of tools for efficient genome engineering in this mosquito will not only lay the foundation for the application of novel genetic control strategies that do not rely on insecticides, but will also accelerate basic research on key biological processes involved in disease transmission. Here, we report the development of a transgenic CRISPR approach for rapid gene disruption in this organism. Given their high editing efficiencies, the Cas9 strains we developed can be used to quickly generate novel genome modifications allowing for high-throughput gene targeting, and can possibly facilitate the development of gene drives, thereby accelerating comprehensive functional annotation and development of innovative population control strategies for Ae. aegypti.

bioengineering

CRISPR/Cas9-APEX-mediated proximity labeling enables discovery of proteins associated with a predefined genomic locus in living cells

The activation or repression of a genes expression is primarily controlled by changes in the proteins that occupy its regulatory elements. The most common method to identify proteins associated with genomic loci is chromatin immunoprecipitation (ChIP). While having greatly advanced our understanding of gene expression regulation, ChIP requires specific, high quality, IP-competent antibodies against nominated proteins, which can limit its utility and scope for discovery. Thus, a method able to discover and identify proteins associated with a particular genomic locus within the native cellular context would be extremely valuable. Here, we present a novel technology combining recent advances in chemical biology, genome targeting, and quantitative mass spectrometry to develop genomic locus proteomics, a method able to identify proteins which occupy a specific genomic locus.

biochemistry

BEHST: genomic set enrichment analysis enhanced through integration of chromatin long-range interactions

Transforming data from genome-scale assays into knowledge of affected molecular functions and pathways is a key challenge in biomedical research. Using vocabularies of functional terms and databases annotating genes with these terms, pathway enrichment methods can identify terms enriched in a gene list. With data that can refer to intergenic regions, however, one must first connect the regions to the terms, which are usually annotated only to genes. To make these connections, existing pathway enrichment approaches apply unwarranted assumptions such as annotating non-coding regions with the terms from adjacent genes. We developed a computational method that instead links genomic regions to annotations using data on long-range chromatin interactions. Our method, Biological Enrichment of Hidden Sequence Targets (BEHST), finds Gene Ontology (GO) terms enriched in genomic regions more precisely and accurately than existing methods. We demonstrate BEHSTs ability to retrieve more pertinent and less ambiguous GO terms associated with results of in vivo mouse enhancer screens or enhancer RNA assays for multiple tissue types. BEHST will accelerate the discovery of affected pathways mediated through long-range interactions that explain non-coding hits in genome-wide association study (GWAS) or genome editing screens. BEHST is free software with a command-line interface for Linux or macOS and a web interface (http://behst.hoffmanlab.org/).

bioinformatics

Solagigasbacteria: Lone genomic giants among the uncultured bacterial phyla

Recent advances in single-cell genomic and metagenomic techniques have facilitated the discovery of numerous previously unknown, deep branches of the tree of life that lack cultured representatives. Many of these candidate phyla are composed of microorganisms with minimalistic, streamlined genomes lacking some core metabolic pathways, which may contribute to their resistance to growth in pure culture. Here we analyzed single-cell genomes and metagenome bins to show that the \"Candidate phylum SPAM\" represents an interesting exception, by having large genomes (6-8 Mbps), high GC content (66%-71%), and the potential for a versatile, mixotrophic metabolism. We also observed an unusually high genomic heterogeneity among individual SPAM cells in the studied samples. These features may have contributed to the limited recovery of sequences of this candidate phylum in prior metagenomic studies. Based on these observations, we propose renaming SPAM to \"Candidate phylum Solagigasbacteria\". Current evidence suggests that Solagigasbacteria are distributed globally in diverse terrestrial ecosystems, including soils, the rhizosphere, volcanic mud, oil wells, aquifers and the deep subsurface, with no reports from marine environments to date.

microbiology

Gaussian decomposition of high-resolution melt curve derivatives for measuring genome-editing efficiency

We describe a method for measuring genome editing efficiency from in silico analysis of high-resolution melt curve data. The melt curve data derived from amplicons of genome-edited or unmodified target sites were processed to remove the background fluorescent signal emanating from free fluorophore and then corrected for temperature-dependent quenching of fluorescence of double-stranded DNA-bound fluorophore. Corrected data were normalized and numerically differentiated to obtain the first derivatives of the melt curves. These were then mathematically modeled as a sum or superposition of minimal number of Gaussian components. Using Gaussian parameters determined by modeling of melt curve derivatives of unedited samples, we were able to model melt curve derivatives from genetically altered target sites where the mutant population could be accommodated using an additional Gaussian component. From this, the proportion contributed by the mutant component in the target region amplicon could be accurately determined. Mutant component computations compared well with the mutant frequency determination from next generation sequencing data. The results were also consistent with our earlier studies that used difference curve areas from high-resolution melt curves for determining the efficiency of genome-editing reagents. The advantage of the described method is that it does not require calibration curves to estimate proportion of mutants in amplicons of genome-edited target sites.\n\nSignificance StatementGenome editing has been revolutionized by the engineering of molecular scissors that cut DNA at a predetermined location on the chromosome. When these molecular scissors are expressed within cells, these scissors cut the genomic DNA at the designated target site, and the cells respond by repairing the cut sites. This repair process frequently introduces mutations at the target cut site. The more efficient the molecular scissors, the more number of cells in a treated culture dish exhibit these mutations at the cut site. Investigators therefore design several molecular scissors targeting the same region on the chromosome to identify the best ones. We describe a new recipe to measure scissors efficiency in target site cutting.

bioengineering

Prey range and genome evolution of Halobacteriovorax marinus predatory bacteria from an estuary

BackgroundHalobacteriovorax are saltwater-adapted predatory bacteria that attack Gram-negative bacteria and therefore may play an important role in shaping microbial communities. To understand the impact of Halobacteriovorax on ecosystems and develop them as biocontrol agents, it is important to characterize variation in predation phenotypes such as prey range and investigate the forces impacting Halobacteriovorax genome evolution across different phylogenetic distances.\n\nResultsWe isolated H. marinus BE01 from an estuary in Rhode Island using Vibrio from the same site as prey. Small, fast-moving attack phase BE01 cells attach to and invade prey cells, consistent with the intraperiplasmic predation strategy of H. marinus type strain SJ. BE01 is a prey generalist, forming plaques on Vibrio strains from the estuary as well as Pseudomonas from soil and E. coli. Genome analysis revealed that BE01 is very closely related to SJ, with extremely high conservation of gene order and amino acid sequences. Despite this similarity, we identified two regions of gene content difference that likely resulted from horizontal gene transfer. Analysis of modal codon usage frequencies supports the hypothesis that these regions were acquired from bacteria with different codon usage biases compared to Halobacteriovorax. In BE01, one of these regions includes genes associated with mobile genetic elements, such as a transposase not found in SJ and degraded remnants of an integrase occurring as a full-length gene in SJ. The corresponding region in SJ included unique mobile genetic element genes, such as a site-specific recombinase and bacteriophage-related genes not found in BE01. Acquired functions in BE01 include the dnd operon, which encodes a pathway for DNA modification that may protect DNA from nucleases, and a suite of genes involved in membrane synthesis and regulation of gene expression that was likely acquired from another Halobacteriovorax lineage.\n\nConclusionsOur results support previous observations that Halobacteriovorax prey on a broad range of Gram-negative bacteria. Genome analysis suggests strong selective pressure to maintain the genome in the H. marinus lineage represented by BE01 and SJ, although our results also provide further evidence that horizontal gene transfer plays an important role in genome evolution in predatory bacteria.

microbiology

Recovering genomic clusters of secondary metabolites from lakes: a Metagenomics 2.0 approach

BackgroundMetagenomic approaches became increasingly popular in the past decades due to decreasing costs of DNA sequencing and bioinformatics development. So far, however, the recovery of long genes coding for secondary metabolism still represents a big challenge. Often, the quality of metagenome assemblies is poor, especially in environments with a high microbial diversity where sequence coverage is low and complexity of natural communities high. Recently, new and improved algorithms for binning environmental reads and contigs have been developed to overcome such limitations. Some of these algorithms use a similarity detection approach to classify the obtained reads into taxonomical units and to assemble draft genomes. This approach, however, is quite limited since it can classify exclusively sequences similar to those available (and well classified) in the databases.\n\nIn this work, we used draft genomes from Lake Stechlin, north-eastern Germany, recovered by MetaBat, an efficient binning tool that integrates empirical probabilistic distances of genome abundance, and tetranucleotide frequency for accurate metagenome binning. These genomes were screened for secondary metabolism genes, such as polyketide synthases (PKS) and non-ribosomal peptide synthases (NRPS), using the Anti-SMASH and NAPDOS workflows.\n\nResultsWith this approach we were able to identify 243 secondary metabolite clusters from 121 genomes recovered from the lake samples. A total of 18 NRPS, 19 PKS and 3 hybrid PKS/NRPS clusters were found. In addition, it was possible to predict the partial structure of several secondary metabolite clusters allowing for taxonomical classifications and phylogenetic inferences.\n\nConclusionsOur approach revealed a great potential to recover and study secondary metabolites genes from any aquatic ecosystem.

bioinformatics

Correction of autoimmune IL2RA mutations in primary human T cells using non-viral genome targeting

Human T cells are central to physiological immune homeostasis, which protects us from pathogens without collateral autoimmune inflammation. They are also the main effectors in most current cancer immunotherapy strategies1. Several decades of work have aimed to genetically reprogram T cells for therapeutic purposes2-5, but as human T cells are resistant to most standard methods of large DNA insertion these approaches have relied on recombinant viral vectors, which do not target transgenes to specific genomic sites6, 7. In addition, the need for viral vectors has slowed down research and clinical use as their manufacturing and testing is lengthy and expensive. Genome editing brought the promise of specific and efficient insertion of large transgenes into target cells through homology-directed repair (HDR), but to date in human T cells this still requires viral transduction8, 9. Here, we developed a non-viral, CRISPR-Cas9 genome targeting system that permits the rapid and efficient insertion of individual or multiplexed large (>1 kilobase) DNA sequences at specific sites in the genomes of primary human T cells while preserving cell viability and function. We successfully tested the potential therapeutic use of this approach in two settings. First, we corrected a pathogenic IL2RA mutation in primary T cells from multiple family members with monogenic autoimmune disease and demonstrated enhanced signalling function. Second, we replaced the endogenous T cell receptor (TCR) locus with a new TCR redirecting T cells to a cancer antigen. The resulting TCR-engineered T cells specifically recognized the tumour antigen, with concomitant cytokine release and tumour cell killing. Taken together, these studies provide preclinical evidence that non-viral genome targeting will enable rapid and flexible experimental manipulation and therapeutic engineering of primary human immune cells.

genetics

Identifying Pleiotropic Effects: A Two-Stage Approach Using Genome-Wide Association Meta-Analysis Data

Pleiotropic effects occur when a single genetic variant independently influences multiple phenotypes. In genetic epidemiological studies, multiple endo-phenotypes or correlated traits are commonly tested separately in a univariate statistical framework to identify associations with genetic determinants. Subsequently, a simple look-up of overlapping univariate results is applied to identify pleiotropic genetic effects. However, this strategy offers limited power to detect pleiotropy. In contrast, combining correlated traits into a composite test provides a powerful approach for detecting pleiotropic genes. Here, we propose a two-stage approach to identify potential pleiotropic effects by utilizing aggregated results from large-scale genome-wide association (GWAS) meta-analyses. In the first stage, we developed two novel approaches (direct linear combining, dLC; and empirical combining, eLC) combining correlated univariate test statistics to screen potential pleiotropic variants on a genome-wide scale, using either individual-level or aggregated data. Our simulations indicated that dLC and eLC outperform other popular multivariate approaches (such as principal component analysis (PCA), multivariate analysis of variance (MANOVA), canonical correlation (CCA), generalized estimation equations (GEE), linear mixed effects models (LME) and OBrien combining approach). In particular, eLC provides a notable increase in power when the genetic variant exhibits both protective and deleterious effects. In the second stage, we developed a unique approach, conditional pleiotropy testing (cPLT), to examine pleiotropic effects using individual-level data for candidate variants identified in Stage 1. Simulation demonstrated reduced type 1 error for cPLT in identifying pleiotropic genetic variants compared to the typical conditional strategy. We validated our two-stage approach by performing a bivariate GWA study on two correlated quantitative traits, high-density lipoprotein (HDL) and triglycerides (TG), in the Genetic Analysis Workshop 16 (GAW16) simulation dataset. In summary, the proposed two-stage approach allows us to leverage aggregated summary statistics from univariate GWAS and improves the power to identify potential pleiotropy while maintaining valid false-positive rates.\n\nAuthor SummaryPleiotropy, occurring when a single genetic variant contributes to multiple phenotypes, remains difficult to identify in genome-wide association studies (GWAS). To leverage data for multiple phenotypes and incorporate univariate GWAS summary results, we propose a novel two-stage approach for discovering potential pleiotropic variants. In the first stage, two novel combining approaches were developed to screen potential pleiotropic variants on a genome-wide scale. Simulations demonstrated the superior statistical power of these approaches over other multivariate methods. In the second stage, our approach was used to identify potential pleiotropy in the candidate marker sets generated from the first stage. The proposed two-stage approach was applied to the GAW16 simulation dataset to discover pleiotropic variants associated with high-density lipoprotein and triglycerides. In summary, we demonstrate that the proposed two-stage approach can be applied as a viable and robust strategy to accommodate phenotypic and genetic heterogeneity for discovering potential pleiotropy on genome-wide scale.

genetics

Enabling rapid cloud-based analysis of thousands of human genomes via Butler

We present Butler, a computational framework developed in the context of the international Pan-cancer Analysis of Whole Genomes (PCAWG)1 project to overcome the challenges of orchestrating analyses of thousands of human genomes on the cloud. Butler operates equally well on public and academic clouds. This highly flexible framework facilitates management of virtual cloud infrastructure, software configuration, genomics workflow development, and provides unique capabilities in workflow execution management. By comprehensively collecting and analysing metrics and logs, performing anomaly detection as well as notification and cluster self-healing, Butler enables large-scale analytical processing of human genomes with 43% increased throughput compared to prior setups. Butler was key for delivering the germline genetic variant call-sets in 2,834 cancer genomes analysed by PCAWG1.

bioinformatics

Efficient management and analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr

MotivationGenome-wide datasets produced for association studies have dramatically increased in size over the past few years, with modern datasets commonly including millions of variants measured in dozens of thousands of individuals. This increase in data size is a major challenge severely slowing down genomic analyses. Specialized software for every part of the analysis pipeline have been developed to handle large genomic data. However, combining all these software into a single data analysis pipeline might be technically difficult.\n\nResultsHere we present two R packages, bigstatsr and bigsnpr, allowing for management and analysis of large scale genomic data to be performed within a single comprehensive framework. To address large data size, the packages use memory-mapping for accessing data matrices stored on disk instead of in RAM. To perform data pre-processing and data analysis, the packages integrate most of the tools that are commonly used, either through transparent system calls to existing software, or through updated or improved implementation of existing methods. In particular, the packages implement a fast derivation of Principal Component Analysis, functions to remove SNPs in Linkage Disequilibrium, and algorithms to learn Polygenic Risk Scores on millions of SNPs. We illustrate applications of the two R packages by analysing a case-control genomic dataset for the celiac disease, performing an association study and computing Polygenic Risk Scores. Finally, we demonstrate the scalability of the R packages by analyzing a simulated genome-wide dataset including 500,000 individuals and 1 million markers on a single desktop computer.\n\nAvailabilityhttps://privefl.github.io/bigstatsr/ & https://privefl.github.io/bigsnpr/\n\nContactflorian.prive@univ-grenoble-alpes.fr & michael.blum@univ-grenoble-alpes.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Drosophila larval brain neoplasms present tumour-type dependent genome instability

Single nucleotide polymorphisms (SNPs) and copy number variants (CNVs) are found at different rates in human cancer. To determine if these genetic lesions appear in Drosophila tumours we have sequenced the genomes of 17 malignant neoplasms caused by mutations in l(3)mbt, brat, aurA, or lgl. We have found CNVs and SNPs in all the tumours. Tumour-linked CNVs range between 11 and 80 per sample, affecting between 92 and 1546 coding sequences. CNVs are in average less frequent in l(3)mbt than in brat lines. Nearly half of the CNVs fall within the 10 to 100Kb range, all tumour samples contain CNVs larger that 100 Kb and some have CNVs larger than 1Mb. The rates of tumour-linked SNPs change more than 20-fold depending on the tumour type: late stage brat, l(3)mbt, and aurA and lgl lines present median values of SNPs/Mb of exome of 0.16, 0.48, and 3.6, respectively. Higher SNP rates are mostly accounted for by C>A transversions, which likely reflect enhanced oxidative stress conditions in the affected tumours. Both CNVs and SNPs turn over rapidly. We found no evidence for selection of a gene signature affected by CNVs or SNPs in the cohort. Altogether, our results show that the rates of CNVs and SNPs, as well as the distribution of CNV sizes in this cohort of Drosophila tumours are well within the range of those reported for human cancer. Genome instability is therefore inherent to Drosophila malignant neoplastic growth at a variable extent that is tumour type dependent.\n\nAUTHOR SUMMARYDrosophila models of malignant growth can help to understand the molecular mechanisms of malignancy. These models are known to exhibit some of the hallmarks of cancer like sustained growth, immortality, metabolic reprogramming, and others. However, it is currently unclear if these fly models are affected by genome instability, which is another hallmark of many human malignant tumours. To address this issue we have sequenced and analysed the genomes of a cohort of seventeen fly tumour samples. We have found that genome instability is a common trait of Drosophila malignant tumours, which occurs at an extent that is tumour-type dependent, at rates that are similar to those of different human cancers.

cancer biology

MeDEStrand: an improved method to infer genome-wide absolute methylation level from DNA enrichment experiment

BackgroundDNA methylation of dinucleotide CpG is an essential epigenetic modification that plays a key role in transcription. Bisulfite conversion method is a \"gold standard\" for DNA methylation profiling that provides single nucleotide resolution. However, whole-genome bisulfite conversion is very expensive. Alternatively, DNA enrichment-based methods offer high coverage of methylated CpG dinucleotides with the lowest cost per CpG covered genome-wide and have been used widely. They measure the DNA enrichment of methyl-CpG binding, therefore do not directly provide absolute methylation levels. Further, the enrichment is influenced by confounding factors besides the methylation status, e.g., CpG density. Computational models that can accurately derive the absolute methylation levels from the enrichment data are necessary.\n\nResultsWe present MeDEStrand, a method uses sigmoid function to estimate and correct the CpG bias from the numbers of reads that fell within bins that divide the genome. In addition, unlike the previous methods, which estimate CpG bias based on reads mapped at the same genomic loci, MeDEStrand processes the reads for the positive and negative DNA strands separately. We compare the performance of MeDEStrand with three other state-of-the-art methods MEDIPS, BayMeth and QSEA on four independent datasets generated using immortalized cell lines (GM12878 and K562) and human patient primary cells (foreskin fibroblast and mammary epithelial). Based on the comparison between the inferred absolute methylation levels from MeDIP-seq and the corresponding RRBS data, MeDEStrand shows the best performance at high resolution of 25, 50 and 100 base pairs.\n\nConclusions MeDEStrand benefits from the estimation of CpG bias with a sigmoid function and the procedure to process reads mapped to the positive and negative DNA strands separately. MeDEStrand is a tool to infer whole-genome absolute DNA methylation level at the cost of enrichment-based methods with adequate accuracy and resolution. R package MeDEStrand and its tutorial is freely available for download at https://github.com/jxu1234/MeDEStrand.git

bioinformatics

An alignment method for nucleic acid sequences against annotated genomes

MotivationBiological sequence alignment is fundamental to their further interpretation. Current alignment algorithms typically align either nucleic acid or amino acid sequences. Using only nucleic acid sequence similarity, divergent sequences cannot be aligned reliably because of the limited alphabet and genetic saturation. To align divergent coding nucleic acid sequences, one can align using the translated amino acid sequences. This requires the detection of the correct open reading frame, is prone to eventual frame shift errors, and typically requires the treatment of genes separately. It was our motivation to design a nucleic acid sequence alignment algorithm to align a nucleic acid sequence against a (reference) genome sequence, that works equally well for similar and divergent sequences, and produces an optimal alignment considering simultaneously the alignment of all annotated coding sequences.\n\nResultsWe define a genome alignment score for evaluating the quality of an alignment of a nucleic acid query sequence against a reference genome sequence, for which coding sequence features have been annotated (for example in a GenBank record). The genome alignment score combines the a ne gap score for the nucleic acid sequence with an a ne gap score for all amino acid alignments resulting from coding sequences in open reading frames contained within the query sequence. We present a Dynamic Programming algorithm to compute the optimal global or local alignment using this genomic alignment score and provide a formal proof of correctness. This algorithm allows the alignment of nucleic acid sequences from closely related and highly divergent sequences within the same software and using the same parameters, automatically correcting any eventual frame shift errors and produces at the same time the aligned translated amino acid sequences of all relevant coding sequence features.\n\nAvailabilityThe software is available as a web application at http://www.genomedetective.com/app/aga and as command-line application at https://github.com/emweb/aga

bioinformatics

Bayesian Reconstruction of Transmission within Outbreaks using Genomic Variants

Pathogen genome sequencing can reveal details of transmission histories and is a powerful tool in the fight against infectious disease. In particular, within-host pathogen genomic variants identified through heterozygous nucleotide base calls are a potential source of information to identify linked cases and infer direction and time of transmission. However, using such data effectively to model disease transmission presents a number of challenges, including differentiating genuine variants from those observed due to sequencing error, as well as the specification of a realistic model for within-host pathogen population dynamics.\n\nHere we propose a new Bayesian approach to transmission inference, BadTrIP (BAyesian epiDemiological TRansmission Inference from Polymorphisms), that explicitly models evolution of pathogen populations in an outbreak, transmission (including transmission bottlenecks), and sequencing error. BadTrIP enables the inference of host-to-host transmission from pathogen sequencing data and epidemiological data. By assuming that genomic variants are unlinked, our method does not require the computationally intensive and unreliable reconstruction of individual haplotypes. Using simulations we show that BadTrIP is robust in most scenarios and can accurately infer transmission events by efficiently combining information from genetic and epidemiological sources; thanks to its realistic model of pathogen evolution and the inclusion of epidemiological data, BadTrIP is also more accurate than existing approaches. BadTrIP is distributed as an open source package (https://bitbucket.org/nicofmay/badtrip) for the phylogenetic software BEAST2.\n\nWe apply our method to reconstruct transmission history at the early stages of the 2014 Ebola outbreak, showcasing the power of within-host genomic variants to reconstruct transmission events.\n\nAuthor SummaryWe present a new tool to reconstruct transmission events within outbreaks. Our approach makes use of pathogen genetic information, notably genetic variants at low frequency within host that are usually discarded, and combines it with epidemiological information of host exposure to infection. This leads to accurate reconstruction of transmission even in cases where abundant within-host pathogen genetic variation and weak transmission bottlenecks (multiple pathogen units colonising a new host at transmission) would otherwise make inference difficult due to the transmission history differing from the pathogen evolution history inferred from pathogen isolets. Also, the use of within-host pathogen genomic variants increases the resolution of the reconstruction of the transmission tree even in scenarios with limited within-outbreak pathogen genetic diversity: within-host pathogen populations that appear identical at the level of consensus sequences can be discriminated using within-host variants. Our Bayesian approach provides a measure of the confidence in different possible transmission histories, and is published as open source software. We show with simulations and with an analysis of the beginning of the 2014 Ebola outbreak that our approach is applicable in many scenarios, improves our understanding of transmission dynamics, and will contribute to finding and limiting sources and routes of transmission, and therefore preventing the spread of infectious disease.

genetics

MetQy: an R package to query metabolic functions of genes and genomes

SummaryWith the rapid accumulation of sequencing data from genomic and metagenomic studies, there is an acute need for better tools that facilitate their analyses against biological functions. To this end, we developed MetQy, an open-source R package designed for query-based analysis of functional units in [meta]genomes and/or sets of genes using the The Kyoto Encyclopedia of Genes and Genomes (KEGG) database. Furthermore, MetQy contains visualization and analysis tools and facilitates KEGGs flat file manipulation. Thus, MetQy enables better understanding of metabolic capabilities of known genomes or user-specified [meta]genomes by using the available information and can help guide studies in microbial ecology, metabolic engineering and synthetic biology.\n\nAvailability and ImplementationThe MetQy R package is freely available and can be downloaded from our groups website (http://osslab.lifesci.warwick.ac.uk) or GitHub (https://github.com/OSS-Lab/MetQy).\n\nContactO.Soyer@warwick.ac.uk

microbiology

Widespread ancient whole genome duplications in Malpighiales coincide with Eocene global climatic upheaval

Ancient whole genome duplications (WGDs) are important in eukaryotic genome evolution, and are especially prominent in plants. Recent genomic studies from large vascular plant clades, including ferns, gymnosperms, and angiosperms suggest that WGDs may represent a crucial mode of speciation. Moreover, numerous WGDs have been dated to events coinciding with major episodes of global and climatic upheaval, including the mass extinction at the KT boundary (~65 Ma) and during more recent intervals of global aridification in the Miocene (~10-5 Ma). These findings have led to the hypothesis that polyploidization may buffer lineages against the negative consequences of such disruptions. Here, we explore WGDs in the large, and diverse flowering plant clade Malpighiales using a combination of transcriptomes and complete genomes from 42 species. We conservatively identify 22 ancient WGDs, widely distributed across Malpighiales subclades. Our results provide strong support for the hypothesis that WGD is an important mode of speciation in plants. Importantly, we also identify that these events are clustered around the Eocene-Paleocene Transition (~54 Ma), during which time the planet was warmer and wetter than any period in the Cenozoic. These results establish that the Eocene Climate Optimum represents another, previously unrecognized, period of prolific WGDs in plants, and lends support to the hypothesis that polyploidization promotes adaptation and enhances plant survival during major episodes of global change. Malpighiales, in particular, may have been particularly influenced by these events given their predominance in the tropics where Eocene warming likely had profound impacts owing to the relatively tight thermal tolerances of tropical organisms.\n\nSignificance StatementWhole genome duplications (WGDs) are hypothesized to generate adaptive variations during episodes of climate change and global upheaval. Using large-scale phylogenomic assessments, we identify an impressive 22 ancient WGDs in the large, tropical flowering plant clade Malpighiales. This supports growing evidence that ancient WGDs are far more common than has been thought. Additionally, we identify that WGDs are clustered during a narrow window of time, ~54 Ma, when the climate was warmer and more humid than during any period in the last ~65 Ma. This lends support to the hypothesis that WGDs are associated with surviving climatic upheavals, especially for tropical organisms like Malpighiales, which have tight thermal tolerances.

evolutionary biology

GrapeTree: Visualization of core genomic relationships among 100,000 bacterial pathogens

O_LICurrent methods struggle to reconstruct and visualise the genomic relationships of [≥]100,000 bacterial genomes.\nC_LIO_LIGrapeTree facilitates the analyses of allelic profiles from 10,000s of core genomes within a web browser window.\nC_LIO_LIGrapeTree implements a novel minimum spanning tree algorithm to reconstruct genetic relationships despite missing data together with a static \"GrapeTree Layout\" algorithm to render interactive visualisations of large trees.\nC_LIO_LIGrapeTree is a stand-along package for investigating Newick trees plus associated metadata and is also integrated into EnteroBase to facilitate cutting edge navigation of genomic relationships among >160,000 genomes from bacterial pathogens.\nC_LIO_LIThe GrapeTree package was released under the GPL v3.0 Licence.\nC_LI

bioinformatics