bioRxiv ScienceSearch

Biology subjects

Soares, A.

Publications and source records attributed to Soares, A..

3 recordsLinked to original sources

Computational haplotype recovery and long-read validation identifies novel isoforms of industrially relevant enzymes from natural microbial communities

Elucidation of population-level diversity of microbiomes is a significant step towards a complete understanding of the evolutionary, ecological and functional importance of microbial communities. Characterizing this diversity requires the recovery of the exact DNA sequence (haplotype) of each gene isoform from every individual present in the community. To address this, we present Hansel and Gretel: a freely-available data structure and algorithm, providing a software package that reconstructs the most likely haplotypes from metagenomes. We demonstrate recovery of haplotypes from short-read Illumina data for a bovine rumen microbiome, and verify our predictions are 100% accurate with long-read PacBio CCS sequencing. We show that Gretels haplotypes can be analyzed to determine a significant difference in mutation rates between core and accessory gene families in an ovine rumen microbiome. All tools, documentation and data for evaluation are open source and available via our repository: https://github.com/samstudio8/gretel

bioinformatics

Serial Crystallography with Multi-stage Merging of 1000’s of Images

KAMO and Blend provide particularly effective tools to manage automatically the merging of large numbers of datasets from serial crystallography. The requirement for manual intervention in the process can be reduced by extending Blend to support additional clustering options such as use of more accurate cell distance metrics and use of reflection-intensity correlation coefficients to infer "distances" among sets of reflec- tions. This increases the sensitivity to differences in unit cell parameters and allows for clustering to assemble nearly complete datasets on the basis of intensity or ampli- tude differences. If datasets are already sufficiently complete to permit it, one applies KAMO once and clusters the data using intensities only. If starting from incomplete datasets, one applies KAMO twice, first using cell parameters. In this step we use either the simple cell vector distance of the original Blend, or we use the more sensi- tive NCDist. This step tends to find clusters of sufficient size so that, when merged, each cluster is sufficiently complete to allow reflection intensities or amplitudes to be compared. One then uses KAMO again using the correlation between the reflections having a common hkl to merge clusters in a way sensitive to structural differences that may not have perturbed the cell parameters sufficiently to make meaningful clusters. Many groups have developed effective clustering algorithms that use a measurable physical parameter from each diffraction still or wedge to cluster the data into cate- gories which then can be merged, one hopes, to yield the electron density from a single protein form. Since these physical parameters are often largely independent from one another, it should be possible to greatly improve the efficacy of data clustering software by using a multi-stage partitioning strategy. Here, we have demonstrated one possible approach to multi-stage data clustering. Our strategy is to use unit-cell clustering until merged data is sufficiently complete then to use intensity-based clustering. We have demonstrated that, using this strategy, we are able to accurately cluster datasets from crystals that have subtle differences.

bioinformatics

Deep Sequencing: Intra-Terrestrial Metagenomics Illustrates The Potential Of Off-Grid Nanopore DNA Sequencing

Genetic and genomic analysis of nucleic acids from environmental samples has helped transform our perception of the Earths subsurface as a major reservoir of microbial novelty. Many of the microbial taxa living in the subsurface are under-represented in culture-dependent investigations. In this regard, metagenomic analyses of subsurface environments exemplify both the utility of metagenomics and its power to explore microbial life in some of the most extreme and inaccessible environments on Earth. Hitherto, the transfer of microbial samples to home laboratories for DNA sequencing and bioinformatics is the standard operating procedure for exploring microbial diversity. This approach incurs logistical challenges and delays the characterization of microbial biodiversity. For selected applications, increased portability and agility in metagenomic analysis is therefore desirable. Here, we describe the implementation of sample extraction, metagenomic library preparation, nanopore DNA sequencing and taxonomic classification using a portable, battery-powered, suite of off-the-shelf tools (the \"MetageNomad\") to sequence ochreous sediment microbiota while within the South Wales Coalfield. While our analyses were frustrated by short read lengths and a limited yield of DNA, within the assignable reads, Proteobacterial (-, {beta}-, {gamma}-Proteobacteria) taxa dominated, followed by members of Actinobacteria, Firmicutes and Bacteroidetes, all of which have previously been identified in coals. Further to this, the fungal genus Candida was detected, as well as a methanogenic archaeal taxon. To the best of our knowledge, this application of the MetageNomad represents an initial effort to conduct metagenomics within the subsurface, and stimulates further developments to take metagenomics off the beaten track.

microbiology