bioRxiv ScienceSearch

Biology subjects

Heavens, D.

Publications and source records attributed to Heavens, D..

7 recordsLinked to original sources

City life: airborne DNA metagenomic biodiversity monitoring reveals dynamic changes across time and space

Airborne environmental DNA can capture biodiversity across the tree of life, but low sample biomass makes rapid, untargeted detection technically challenging. We combined 45-min air collection, nanopore sequencing and real-time taxonomic analysis in a shotgun metagenomic workflow capable of producing results within 3 hours. Across 77 samples from 13 London sites, including a year of weekly sampling at the Natural History Museum Wildlife Garden, we detected 1,916 species spanning bacteria, fungi, plants and animals. Communities varied spatially and seasonally, shifting from plant dominance in spring to ascomycete dominance in summer and basidiomycete dominance in late autumn and winter. Plant read abundance increased with upwind vegetation, linking airborne signals to surrounding habitat. Detection of catalogued garden plants depended on reference availability, dispersal biology, plant size and proximity to the collector. Together, these findings establish airborne shotgun metagenomics as a platform for rapid, repeated and scalable biodiversity assessment across space and time.

ecology

Arabidopsis thaliana populations support long-term maintenance and parallel expansions of related Pseudomonas pathogens

Crop disease outbreaks are often associated with clonal expansions of single pathogenic lineages. To determine whether similar boom-and-bust scenarios hold for wild plant pathogens, we carried out a multi-year multi-site survey of Pseudomonas in the natural host Arabidopsis thaliana. The most common Pseudomonas lineage corresponded to a pathogenic clade present in all sites. Sequencing of 1,524 Pseudomonas genomes revealed this lineage to have diversified approximately 300,000 years ago, containing dozens of genetically distinct pathogenic sublineages. These sublineages have expanded in parallel within the same populations and are differentiated both at the level of gene content and disease phenotype. Such coexistence of diverse sublineages indicates that in contrast to crop systems, no single strain has been able to overtake these A. thaliana populations in the recent past. Our results suggest that the selective pressures acting on a plant pathogen in wild hosts may be more complex than those in agricultural systems.

microbiology

Independent assessment and improvement of wheat genome assemblies using Fosill jumping libraries

BackgroundThe accurate sequencing and assembly of very large, often polyploid, genomes remain a challenging task, limiting long range sequence information and phased sequence variation for applications such as plant breeding. The 15 Gb hexaploid bread wheat genome has been particularly challenging to sequence, and several contending approaches recently generated accurate long range assemblies. Understanding errors in these assemblies is important for optimising future sequencing and assembly approaches and for comparative genomics.\n\nResultsHere we use a Fosill 38 Kb jumping library to assess medium and longer range order of different publicly available wheat genome assemblies. Modifications to the Fosill protocol generated longer Illumina sequences and enabled comprehensive genome coverage. Analyses of two independent BAC based chromosome-scale assemblies, two independent Illumina whole genome shotgun assemblies, and a hybrid long read (PacBio) and short read (Illumina) assembly were carried out. We revealed a variety of discrepancies using Fosill mate-pair mapping and validated several of each class. In addition, Fosill mate-pairs were used to scaffold a whole genome Illumina assembly, leading to a three-fold increase in N50 values.\n\nConclusionsOur analyses, using an independent means to validate different wheat genome assemblies, show that whole genome shotgun assemblies are significantly more accurate by all measures compared to BAC-based chromosome scale assemblies. Although current whole genome assemblies are reasonably accurate and useful, additional steps will be needed for the rapid, cost effective and complete sequencing and assembly of wheat genomes.

genomics

A critical comparison of technologies for a plant genome sequencing project

A high quality genome sequence of your model organism is an essential starting point for many studies. Old clone based methods are slow and expensive, whereas faster, cheaper short read only assemblies can be incomplete and highly fragmented, which minimises their usefulness. The last few years have seen the introduction of many new technologies for genome assembly. These new technologies and new algorithms are typically benchmarked on microbial genomes or, if they scale appropriately, human. However, plant genomes can be much more repetitive and larger than human, and plant biology makes obtaining high quality DNA free from contaminants difficult. Reflecting their challenging nature we observe that plant genome assembly statistics are typically poorer than for vertebrates. Here we compare Illumina short read, PacBio long read, 10x Genomics linked reads, Dovetail Hi-C and BioNano Genomics optical maps, singly and combined, in producing high quality long range genome assemblies of the potato species S. verrucosum. We benchmark the assemblies for completeness and accuracy, as well as DNA, compute requirements and sequencing costs. We expect our results will be helpful to other genome projects, and that these datasets will be used in benchmarking by assembly algorithm developers.

genomics

Rapid MinION metagenomic profiling of the preterm infant gut microbiota to aid in pathogen diagnostics

The Oxford Nanopore MinION sequencing platform offers near real time analysis of DNA reads as they are generated, which makes the device attractive for in-field or clinical deployment, e.g. rapid diagnostics. We used the MinION platform for shotgun metagenomic sequencing and analysis of gut-associated microbial communities; firstly, we used a 20-species human microbiota mock community to demonstrate how Nanopore metagenomic sequence data can be reliably and rapidly classified. Secondly, we profiled faecal microbiomes from preterm infants at increased risk of necrotising enterocolitis and sepsis. In single patient time course, we captured the diversity of the immature gut microbiota and observed how its complexity changes over time in response to interventions, i.e. probiotic, antibiotics and episodes of suspected sepsis. Finally, we performed real-time runs from sample to analysis using faecal samples of critically ill infants and of healthy infants receiving probiotic supplementation. Real-time analysis was facilitated by our new NanoOK RT software package which analysed sequences as they were generated. We reliably identified potentially pathogenic taxa (i.e. Klebsiella pneumoniae and Enterobacter cloacae) and their corresponding antimicrobial resistance (AMR) gene profiles within as little as one hour of sequencing. Antibiotic treatment decisions may be rapidly modified in response to these AMR profiles, which we validated using pathogen isolation, whole genome sequencing and antibiotic susceptibility testing. Our results demonstrate that our pipeline can process clinical samples to a rich dataset able to inform tailored patient antimicrobial treatment in less than 5 hours.

genomics

W2RAP: a pipeline for high quality, robust assemblies of large complex genomes from short read data

Producing high-quality whole-genome shotgun de novo assemblies from plant and animal species with large and complex genomes using low-cost short read sequencing technologies remains a challenge. But when the right sequencing data, with appropriate quality control, is assembled using approaches focused on robustness of the process rather than maximization of a single metric such as the usual contiguity estimators, good quality assemblies with informative value for comparative analyses can be produced. Here we present a complete method described from data generation and qc all the way up to scaffold of complex genomes using Illumina short reads and its application to data from plants and human datasets. We show how to use the w2rap pipeline following a metric-guided approach to produce cost-effective assemblies. The assemblies are highly accurate, provide good coverage of the genome and show good short range contiguity. Our pipeline has already enabled the rapid, cost-effective generation of de novo genome assemblies from large, polyploid crop species with a focus on comparative genomics.\n\nAvailabilityw2rap is available under MIT license, with some subcomponents under GPL-licenses. A ready-to-run docker with all software pre-requisites and example data is also available.\n\nhttp://github.com/bioinfologics/w2rap\n\nhttp://github.com/bioinfologics/w2rap-contigger

bioinformatics

An improved assembly and annotation of the allohexaploid wheat genome identifies complete families of agronomic genes and provides genomic evidence for chromosomal translocations.

Advances in genome sequencing and assembly technologies are generating many high quality genome sequences, but assemblies of large, repeat-rich polyploid genomes, such as that of bread wheat, remain fragmented and incomplete. We have generated a new wheat whole-genome shotgun sequence assembly using a combination of optimised data types and an assembly algorithm designed to deal with large and complex genomes. The new assembly represents more than 78% of the genome with a scaffold N50 of 88.8kbp that has a high fidelity to the input data. Our new annotation combines strand-specific Illumina RNAseq and PacBio full-length cDNAs to identify 104,091 high confidence protein-coding genes and 10,156 non-coding RNA genes. We confirmed three known and identified one novel genome rearrangements. Our approach enables the rapid and scalable assembly of wheat genomes, the identification of structural variants, and the definition of complete gene models, all powerful resources for trait analysis and breeding of this key global crop. [Supplemental material is available for this article.]

genomics