bioRxiv Science⌕ Search

Biology subjects

Guinet, B.

Publications and source records attributed to Guinet, B..

5 recordsLinked to original sources

metaJAM: a Nextflow integrated metagenomic workflow for sedimentary ancient DNA

The application of metagenomics in ancient DNA (aDNA) research is rapidly expanding, driven in particular by advances in sedimentary aDNA research and sequencing technologies. Although many ancient DNA studies rely on broadly similar bioinformatic strategies, there is still no single standardized, widely adopted workflow. These differences can directly affect how efficiently past biodiversity can be reconstructed and authenticated from the various archives analyzed using ancient metagenomic approaches. Although a few pipelines tackle the processing of ancient DNA data from shotgun sequencing, the ones applied to metagenomic datasets are scarce and often resource-intensive or challenging to install, update, or extend with new tools and parameters. metaJAM, a scalable and user-friendly pipeline, is presented here to specifically address the challenges of metagenomic aDNA analyses of eukaryotes. The pipeline has been designed in Nextflow to ensure continuous development and can be used on different high-performance computing (HPC) clusters. metaJAM integrates all key steps required for ancient DNA metagenomic analyses, from raw sequencing data pre-processing to microbial filtering, taxonomic assignment via competitive iterative mapping against Bowtie 2 reference indexes and reassignment using lowest common ancestor (LCA) inference. Validation and authentication are performed using the post-LCA toolkit bamdam together with alignment to an exhaustive reference database using MMseqs2. It allows users to choose among alternative tools and generates a series of plots to support data visualization and taxon authentication. metaJAM differs from existing pipelines through its implementation of rigorous filtering of microbial-like reads by Kraken 2 classification and masking microbial-like regions, iterative or parallel Bowtie 2 mapping, validation of the detected taxa and integration of up-to-date tools for ancient metagenomic analysis, along with diagnostic plots that help users assess the reliability of taxonomic assignments and visualize their data. It complies well with limited computational resources, customised databases for taxonomical groups, and provides an accessible workflow to support the investigation of metagenomic ancient DNA datasets. Its applications span a range of contexts, from ecosystem reconstructions in environmental aDNA archives such as sediments, to metagenomic studies on archaeological artefacts and even taxonomic identification of undiagnosed biological materials.

bioinformatics↗

DNAharvester: A Nextflow Pipeline for Analysing Highly Degraded DNA from Ancient and Historical Specimens

Ancient DNA (aDNA) research has advanced rapidly with the development of high-throughput sequencing, enabling genome-wide analyses of large collections of prehistoric specimens. However, analysing palaeontological and archaeological material with highly degraded DNA constitutes a major bioinformatic challenge. DNA from such samples is characterised by short fragment lengths, low endogenous content, post-mortem damage, and cross-species contamination, which can increase spurious mapping and reference bias, affecting downstream population genetic inferences. We present DNAharvester, a modular and reproducible pipeline designed specifically for processing highly degraded DNA from ancient and historical specimens. DNAharvester integrates metagenomic filtering, competitive mapping, adaptive aligner selection (incorporating BWA-aln, BWA-mem, and Bowtie2), and systematic evaluation of reference bias and spurious mapping. By incorporating flexible mapping and filtering strategies, the pipeline can be adapted to varying sample preservation, focusing on maximising authentic data recovery. DNAharvester features subworkflows for iterative assembly of mitogenomes, identification of genomic repeats and CpG sites, taxonomic classification, microbial/pathogen screening, genetic sex determination, and variant calling. To accommodate varying sequencing depths, the pipeline supports diploid variant calling, genotype likelihood estimation, and pseudo-haploid random allele calling. Implemented in Nextflow, DNAharvester provides a highly scalable, containerised framework that enhances reproducibility, portability, and robustness in aDNA analyses. We validated the pipeline using simulated and empirical datasets, demonstrating its ability to systematically mitigate complex background contamination while preserving authentic genomic signals. By streamlining complex bioinformatic tasks through simple configuration files, DNAharvester establishes a standardised approach for analysing aDNA datasets and makes genomic analyses of ancient remains accessible to the broader research community.

bioinformatics↗

AncientMetagenomeDir dating metadataset highlights need for standardised radiocarbon reporting in ancient DNA

Ancient DNA is a valuable data source for the understanding of our past. However, to effectively interpret this data, it is essential to know the age of the samples from which the DNA is obtained. Although the field of palaeogenomics has been recognised for its robust open data sharing practices, dating information associated with analysed samples is not reported consistently across palaeogenomic studies, nor is it included as metadata in most genetic data repositories. Here, we describe the addition of standardised precise dating information for ancient microbial genomes into the AncientMetagenomeDir metadata repository of published ancient metagenomic samples. This extension currently includes dating information for over 700 ancient microbial genomic datasets, of which 333 are dated using historical, contextual, or stratigraphic methods, and 405 are radiocarbon dated. We quantitatively assess the quality of radiocarbon date reporting and find that, despite established reporting conventions, radiocarbon dating information is often reported inconsistently across ancient metagenomic studies. This new resource provides ancient microbial researchers with standardised dating information that facilitates more accurate and consistent analysis of metagenomic sequencing data. The dataset also highlights the need for greater standardisation of radiocarbon date reporting in original publications in order to allow effective reuse of this and future ancient microbial data.

bioinformatics↗

The Impact of Reference Genome Divergence on Ancient DNA Damage Detection in Metagenomic Contexts

The reliability of ancient DNA (aDNA) authentication depends on detecting characteristic damage patterns, particularly cytosine deamination at fragment ends. However, in ancient metagenomic studies, sequence divergence between aDNA reads and available reference genomes may obscure such damage signals. We systematically evaluated how reference genome divergence, read count, read length, and damage levels affect aDNA damage profiles using both empirical datasets and controlled simulations. Using ancient Yersinia pestis and Hepatitis B virus data, we show that mapping to divergent reference genomes significantly reduces the detectability and intensity of characteristic damage patterns, particularly at low read counts. Simulations further revealed that reference genome identity is the strongest predictor of damage intensity, while read count primarily influences damage stochasticity. We introduce a correction matrix that adjusts C-to-T damage profiles for reference divergence, improving damage signal recovery. Our findings highlight methodological considerations for authenticating aDNA in metagenomic contexts, particularly when closely related reference genomes are unavailable.

bioinformatics↗

Disinfecting eukaryotic reference genomes to improve taxonomic inference from ancient environmental metagenomic data

Ancient environmental DNA is increasingly essential for reconstructing past ecosystems, particularly when palaeontological and archaeological tissue remains are absent. Detecting ancient plant and animal DNA in environmental samples often relies on using extensive eukaryotic reference genome databases for profiling shotgun metagenomics data. However, microbial contamination in these references can introduce substantial biases in taxonomic assignments, especially given the typical low abundance of plant and animal DNA in such samples. In this study, we present a method for identifying bacterial and archaeal-like sequences in eukaryotic genomes and apply it to nearly 3,000 reference genomes from NCBI RefSeq and GenBank (vertebrates, invertebrates, plants) as well as the 1,323 PhyloNorway plant genome assemblies from herbarium material from northern high-latitude regions. Our analysis reveals microbial-like sequences in many eukaryotic reference genomes, which are most pronounced in the PhyloNorway dataset. We provide a detailed map of the microbial-like regions, including genomic coordinates and taxonomic annotations. This resource enables the masking of microbial-like regions during profiling analyses, thereby improving the reliability of ancient environmental metagenomic datasets for downstream analyses.

genomics↗