bioRxiv ScienceSearch

Biology subjects

Hazarika, R. R.

Publications and source records attributed to Hazarika, R. R..

2 recordsLinked to original sources

MethylStar: A fast and robust pipeline for high-throughput analysis of bulk or single-cell WGBS data

BackgroundWhole-Genome Bisulfite Sequencing (WGBS) is a Next Generation Sequencing (NGS) technique for measuring DNA methylation at base resolution. Continuing drops in sequencing costs are beginning to enable high-throughput surveys of DNA methylation in large samples of individuals and/or single cells. These surveys can easily generate hundreds or even thousands of WGBS datasets in a single study. The efficient pre-processing of these large amounts of data poses major computational challenges and creates unnecessary bottlenecks for downstream analysis and biological interpretation. ResultsTo offer an efficient analysis solution, we present MethylStar, a fast, stable and flexible pre-processing pipeline for WGBS data. MethylStar integrates well-established tools for read trimming, alignment and methylation state calling in a highly parallelized environment, manages computational resources and performs automatic error detection. MethylStar offers easy installation through a dockerized container with all preloaded dependencies and also features a user-friendly interface designed for experts/non-experts. Application of MethylStar to WGBS from human, maize and Arabidopsis shows that it outperforms existing pre-processing pipelines in terms of speed and memory requirements. ConclusionsMethylStar is a fast, stable and flexible pipeline for high-throughput pre-processing of bulk or single-cell WGBS data. Its easy installation and user-friendly interface should make it a useful resource for the wider epigenomics community. MethylStar is distributed under GPL-3.0 license and source code is publicly available for download from github https://github.com/jlab-code/MethylStar. Installation through a docker image is available from http://jlabdata.org/methylstar.tar.gz

bioinformatics

AlphaBeta: Computational inference of epimutation rates and spectra from high-throughput DNA methylation data in plants

IntroductionHeritable changes in cytosine methylation can arise stochastically in plant genomes independently of DNA sequence alterations. These so-called spontaneous epimutations appear to be a byproduct of imperfect DNA methylation maintenance during mitotic or meitotic cell divisions. Accurate estimates of the rate and spectrum of these stochastic events are necessary to be able to quantify how epimutational processes shape methylome diversity in the context of plant evolution, development and aging. MethodHere we describe AlphaBeta, a computational method for estimating epimutation rates and spectra from pedigree-based high-throughput DNA methylation data. The approach requires that the topology of the pedigree is known, which is typically the case in the experimental construction of mutation accumulation lines (MA-lines) in sexually or clonally reproducing species. However, this method also works for inferring somatic epimutation rates in long-lived perennials, such as trees, using leaf methylomes and coring data as input. In this case, we treat the tree branching structure as an intra-organismal phylogeny of somatic lineages and leverage information about the epimutational history of each branch. ResultsTo illustrate the method, we applied AlphaBeta to multi-generational data from selfing- and asexually-derived MA-lines in Arabidopsis and dandelion, as well as to intra-generational leaf methylome data of a single poplar tree. Our results show that the epimutation landscape in plants is deeply conserved across angiosperm species, and that heritable epimutations originate mainly during somatic development, rather than from DNA methylation reinforcement errors during sexual reproduction. Finally, we also provide the first evidence that DNA methylation data, in conjunction with statistical epimutation models, can be used as a molecular clock for age-dating trees. ConclusionAlphaBeta faciliates unprecedented quantitative insights into epimutational processes in a wide range of plant systems. Software implementing our method is available as a Bioconductor R package at http://bioconductor.org/packages/3.10/bioc/html/AlphaBeta.html

genomics