bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.06.11.731724

The genome of the coffee bean weevil (Araecerus fasciculatus) reveals a cytochrome P450 repertoire as a convergent candidate mechanism of insect-origin caffeine detoxification

Abstract

The coffee bean weevil, Araecerus fasciculatus (Coleoptera, Curculionoidea, Anthribidae), is a cosmopolitan pest of over 100 stored agricultural commodities, with particular economic impact on coffee (Coffea arabica). Although two chromosome-level anthribid genomes have recently been released as part of the Darwin Tree of Life (DToL) project (Booth et al. 2024; Crowley et al. 2025), no functionally annotated genome has been available for the family. Here we present a draft genome assembly for A. fasciculatus, generated from PacBio HiFi long reads and processed through a three tiered metagenomic filtering pipeline to remove host plant (C. arabica) and microbial contamination. The final assembly spans 475 Mb across 3,617 scaffolds (N50 = 170 kb) with 88.5% BUSCO completeness (insecta_odb10) and only 3.1% duplication. Gene prediction with BRAKER2 identified 22,384 protein-coding genes, of which 11,783 received functional annotations through SwissProt similarity. Notably, we identified 92 cytochrome P450 (CYP) genes, including tandem gene clusters on two scaffolds (4 genes on ptg000464l, 5 genes on ptg001867l), suggestive of lineage-specific expansion through tandem duplication. Homology searches against Drosophila melanogaster caffeine-metabolizing P450s (CYP12D1, CYP6d5, CYP6a8) recovered strong matches (e-values 9.7 x 10-110 to 5.4 x 10-101, 33-38% identity). In stark contrast, comprehensive BLAST searches for bacterial caffeine N-demethylase genes (ndmA/B/C/D), which mediate caffeine degradation via horizontal gene transfer in the coffee berry borer Hypothenemus hampei (Scolytinae), returned zero hits across the A. fasciculatus genome, predicted proteome, and associated bacterial scaffolds. AlphaFold2 structure prediction of four top Araecerus P450 candidates produced high-confidence models (pLDDT 84.5-93.9, pTM 0.735-0.930) with conserved P450 catalytic motifs. Foldseek structural homology searches confirmed that all four candidates adopt cytochrome P450 folds (top hits: human CYP3A4, CYP3A7, CYP11A1; TM-scores 0.90-0.92; probability 1.000), with zero hits to bacterial Rieske-fold enzymes. Molecular docking of caffeine against these structures yielded binding affinities of -5.41 to -5.80 kcal/mol for the Araecerus candidates, comparable to or exceeding the -5.55 kcal/mol obtained for the experimentally validated Drosophila CYP6a8 and substantially stronger than the -3.70 kcal/mol for the bacterial NdmA structural outgroup (PDB: 6ICP). Phylogenetic analysis revealed that all four candidates have clear orthologs in two non-seed-feeding DToL anthribids (Pseudeuparius sepicola and Platystomos albinus), demonstrating that these P450 genes predate the dietary transition to caffeine-containing seeds. The Araecerus candidates predominantly belong to the CYP6 family (clan 3), whereas the primary Drosophila caffeine P450 CYP12D1 belongs to the mitochondrial clan, confirming convergent recruitment of different P450 subfamilies for caffeine metabolism. These results support the hypothesis that A. fasciculatus employs an insect-encoded, P450-mediated caffeine detoxification pathway fundamentally distinct from the bacterial horizontal gene transfer mechanism documented in Scolytinae. This represents convergent evolution of caffeine resistance via independent molecular strategies within Curculionoidea, and provides the first functionally annotated genomic resource for comparative studies across the Anthribidae.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Martinez Aponte, L. V., Rodriguez Ruiz, A., Locke, S. A., Colston, T. J., Van Dam, A. R.. 2026-06-15. The genome of the coffee bean weevil (Araecerus fasciculatus) reveals a cytochrome P450 repertoire as a convergent candidate mechanism of insect-origin caffeine detoxification. https://doi.org/10.64898/2026.06.11.731724

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Genomic correlates of metastatic competence and progression in human melanoma

Genomic events and their timing that grant a primary tumour the competence to disseminate remain poorly defined. We performed sequencing of 247 stage I/II primary cutaneous melanomas (CMs) and 60 matched metastases without intervening therapy from a prospectively followed registry cohort with a median followup of 92 months, integrating copy-number, mutational, protein and spatial-transcriptomic analyses. Relapse was not distinguished by oncogenic point mutations, which were largely shared between primaries and metastases, but by somatic copy-number alterations (SCNAs) and global chromosomal instability. We defined OncoCycle, a six-gene copy-number signature (amplification of CDK4, MCL1 and CD276; biallelic loss of CDKN2A, CDKN2B and TP53BP1) that predicted relapse independently of established clinicopathological features in melanoma, and a pan-cancer analysis. In matched pairs, metastatic progression was driven by continued copy-number evolution and reduction in intra-tumoural heterogeneity, rather than by acquired point mutations, and OncoCycle alterations from primary tumours were preserved in metastasis seeding clones. Clonal reconstruction revealed both monoclonal and polyclonal metastasis seeding, and spatial transcriptomics resolved copy-number-defined metastatic subclones occupying and programming distinct immune and stromal niches. Thus, metastatic competence was primed early by focal SCNAs on a background of chromosomal instability, elaborated by continued copy-number evolution during dissemination and spatio-temporal interactions with the tumour-microenvironment.

genomics↗

Identifying, phasing, and structurally annotating sex chromosomes for genome assemblies using CBS-tools

A complete reference genome for species with chromosomally-determined separate sexes should contain scaffolds for all sex chromosome homologs. However, sex chromosomes present distinct computational challenges compared to autosomes. Here we present a k-mer based analysis that utilizes whole-genome sequencing of a few sex-identified isolates: Cytogenetics-By-Sequencing (CBS) tools. Unlike other approaches that typically address one aspect of the sex chromosomes, CBS-tools strives to guide users from the discovery of the heterogametic sex through identifying the sex-determination region (SDR). The core of CBS-tools is automated quantification of sex-specific k-mers in order to predict the heterogametic sex. Using publicly-available datasets, CBS-tools correctly identified the known sex-system of the 31 species tested. Additionally, we used these k-mers to verify and correct phasing of sex chromosomes between haplotypes in species representing different sex-systems. Finally, we used these k-mers to delimit the SDR boundary using an interactive web platform. CBS-tools was developed with previously unexplored sex chromosome systems in mind, but is also suitable for well-examined sex chromosome pairs.

genomics↗

Evolutionary dynamics of the insertion sequence IS6110 in the Mycobacterium tuberculosis complex

Insertion sequences (IS) are the most common type of transposable element in prokaryotes and shape the structure of genomes through transposition and by providing a substrate for recombination. Despite the mutational impact of IS, the evolutionary dynamics of most elements in host species remain unknown. Here we study the dynamics of IS6110 in 10,000 strains of the Mycobacterium tuberculosis complex (MTBC). We developed a tool that allows the detection and comparison of IS insertions from short reads without using a reference genome. Using ancestral state reconstruction (ASR) on presence-absence patterns of IS6110, we describe the distribution of copy numbers (CNs) in the MTBC, infer birth rates of the element, and identify genomic regions with large numbers of parallel IS6110 insertions. Copy numbers in the MTBC range from 1 in some clades to more than 30 in strains of La3 (M. orygis). IS6110 birth rates scale approximately linearly with copy number and are elevated on terminal branches, consistent with the delayed action of purifying selection. A key characteristic of IS6110 is its occurrence in hotspots: the 5% most frequently targeted regions account for half of all independent insertion events. The motif 5'-TCTCAAAW-3' is enriched around target sites and in hotspots, suggesting that the accumulation of insertions in these regions results through a combination of non-random insertion and purifying selection in other regions. To conclude the study, we propose a niche constraints model according to which the distribution of IS6110 in the MTBC is governed by the rarity of regions that have both suitable DNA properties and little functional value for the host.

genomics↗