bioRxiv ScienceSearch

Biology subjects

Justin O Borevitz

Publications and source records attributed to Justin O Borevitz.

3 recordsLinked to original sources

kWIP: The k-mer Weighted Inner Product, a de novo Estimator of Genetic Similarity

Modern genomics techniques generate overwhelming quantities of data. Extracting population genetic variation demands computationally efficient methods to determine genetic relatedness between individuals or samples in an unbiased manner, preferably de novo. The rapid and unbiased estimation of genetic relatedness has the potential to overcome reference genome bias, to detect mix-ups early, and to verify that biological replicates belong to the same genetic lineage before conclusions are drawn using mislabelled, or misidentified samples.\n\nWe present the k-mer Weighted Inner Product (kWIP), an assembly-, and alignment-free estimator of genetic similarity. kWIP combines a probabilistic data structure with a novel metric, the weighted inner product (WIP), to efficiently calculate pairwise similarity between sequencing runs from their k-mer counts. It produces a distance matrix, which can then be further analysed and visualised. Our method does not require prior knowledge of the underlying genomes and applications include detecting sample identity and mix-up, non-obvious genomic variation, and population structure.\n\nWe show that kWIP can reconstruct the true relatedness between samples from simulated populations. By re-analysing several published datasets we show that our results are consistent with marker-based analyses. kWIP is written in C++, licensed under the GNU GPL, and is available from https://github.com/kdmurray91/kwip.\n\nAuthor SummaryCurrent analysis of the genetic similarity of samples is overly dependent on alignment to reference genomes, which are often unavailable and in any case can introduce bias. We address this limitation by implementing an efficient alignment free sequence comparison algorithm (kWIP). The fast, unbiased analysis kWIP performs should be conducted in preliminary stages of any analysis to verify experimental designs and sample metadata, catching catastrophic errors earlier.\n\nkWIP extends alignment-free sequence comparison methods by operating directly on sequencing reads. kWIP uses an entropy-weighted inner product over k-mers as a estimator of genetic relatedness. We validate kWIP using rigorous simulation experiments. We also demonstrate high sensitivity and accuracy even where there is modest divergence between genomes, and/or when sequencing coverage is low. We show high sensitivity in replicate detection, and faithfully reproduce published reports of population structure and stratification of microbiomes. We provide a reproducible workflow for replicating our validation experiments.\n\nkWIP is an efficient, open source software package. Our software is well documented and cross platform, and tutorial-style workflows are provided for new users.

Bioinformatics

DNA Methylation profiles of diverse Brachypodium distachyon aligns with underlying genetic diversity

DNA methylation, a common modification of genomic DNA, is known to influence the expression of transposable elements as well as some genes. Although commonly viewed as an epigenetic mark, evidence has shown that underlying genetic variation, such as transposable element polymorphisms, often associate with differential DNA methylation states. To investigate the role of DNA methylation variation, transposable element polymorphism, and genomic diversity, whole genome bisulfite sequencing was performed on genetically diverse lines of the model cereal Brachypodium distachyon. Although DNA methylation profiles are broadly similar, thousands of differentially methylated regions are observed between lines. An analysis of novel transposable element indel variation highlighted hundreds of new polymorphisms not seen in the reference sequence. DNA methylation and transposable element variation is correlated with the genome-wide amount of genetic variation present between samples. However, there was minimal evidence that novel transposon insertion or deletions are associated with nearby differential methylation. This study highlights the importance of genetic variation when assessing DNA methylation variation between samples and provides a valuable map of DNA methylation across diverse re-sequenced accessions of this model cereal species.\n\nReviewer Link to deposited dataAll data is publicly available in the NCBI short read archive under BioProject PRJNA281014. Data tables are available for download at https://drive.google.com/file/d/0BzBxfoxlBCneNkY1TFJDU29iSUU/view?usp=sharing

Genomics

Population scale mapping of novel transposable element diversity reveals links to gene regulation and epigenomic variation

Variation in the presence or absence of transposable elements (TEs) is a major source of genetic variation between individuals. Here, we identified 23,095 TE presence/absence variants between 216 Arabidopsis accessions. Most TE variants were rare, and we find a burden of rare variants associated with local extremes of gene expression and DNA methylation levels within the population. Of the common alleles identified, two thirds were not in linkage disequilibrium with nearby SNPs, implicating these variants as a source of novel genetic diversity. Nearly 200 common TE variants were associated with significantly altered expression of nearby genes, and a major fraction of inter-accession DNA methylation differences were associated with nearby TE insertions. Overall, this demonstrates that TE variants are a rich source of genetic diversity that likely plays an important role in facilitating epigenomic and transcriptional differences between individuals, and indicates a strong genetic basis for epigenetic variation.

Genomics