DNAharvester: A Nextflow Pipeline for Analysing Highly Degraded DNA from Ancient and Historical Specimens
Ancient DNA (aDNA) research has advanced rapidly with the development of high-throughput sequencing, enabling genome-wide analyses of large collections of prehistoric specimens. However, analysing palaeontological and archaeological material with highly degraded DNA constitutes a major bioinformatic challenge. DNA from such samples is characterised by short fragment lengths, low endogenous content, post-mortem damage, and cross-species contamination, which can increase spurious mapping and reference bias, affecting downstream population genetic inferences. We present DNAharvester, a modular and reproducible pipeline designed specifically for processing highly degraded DNA from ancient and historical specimens. DNAharvester integrates metagenomic filtering, competitive mapping, adaptive aligner selection (incorporating BWA-aln, BWA-mem, and Bowtie2), and systematic evaluation of reference bias and spurious mapping. By incorporating flexible mapping and filtering strategies, the pipeline can be adapted to varying sample preservation, focusing on maximising authentic data recovery. DNAharvester features subworkflows for iterative assembly of mitogenomes, identification of genomic repeats and CpG sites, taxonomic classification, microbial/pathogen screening, genetic sex determination, and variant calling. To accommodate varying sequencing depths, the pipeline supports diploid variant calling, genotype likelihood estimation, and pseudo-haploid random allele calling. Implemented in Nextflow, DNAharvester provides a highly scalable, containerised framework that enhances reproducibility, portability, and robustness in aDNA analyses. We validated the pipeline using simulated and empirical datasets, demonstrating its ability to systematically mitigate complex background contamination while preserving authentic genomic signals. By streamlining complex bioinformatic tasks through simple configuration files, DNAharvester establishes a standardised approach for analysing aDNA datasets and makes genomic analyses of ancient remains accessible to the broader research community.