BoostMe accurately predicts DNA methylation values in whole-genome bisulfite sequencing of multiple human tissues
BackgroundBisulfite sequencing is widely employed to study the role of DNA methylation in disease; however, the data suffer from biases due to variability in depth of coverage. Imputation of methylation values at low-coverage sites may mitigate these biases while also identifying important genomic features and motifs associated with predictive power.\n\nResultsHere we describe BoostMe, a novel method for imputation of DNA methylation within whole-genome bisulfite sequencing (WGBS) data based on a gradient boosting algorithm. Importantly, we designed a new feature that leverages information from multiple samples in the same tissue and disease state, enabling BoostMe to outperform existing imputation methods in speed and accuracy. We show that imputation improves WGBS concordance with the Infinium MethylationEPIC array at low WGBS sequencing depth, suggesting improvement in WGBS accuracy after imputation. Furthermore, we compare the ability of BoostMe and DeepCpG - a deep neural network method - to identify interesting features and motifs associated with methylation in three human tissues implicated in type 2 diabetes (T2D) etiology. We find that while BoostMe only identifies features important to general methylation levels across tissues, DeepCpG is able to learn differences in methylation-associated sequence motifs among different tissues and identify tissue-specific regulators of differentiation such as EBF1 in adipose, ASCL2 in muscle, and FOXA1, TCF12, and NRF1 in islets. Neither algorithm readily identified T2D-associated features.\n\nConclusionsOur findings demonstrate the current power and limitations of machine and deep learning algorithms to both improve the quality of, and infer biological meaning from, genome-wide DNA methylation data.