bioRxiv Science⌕ Search

Biology subjects

Gardette, A.

Publications and source records attributed to Gardette, A..

2 recordsLinked to original sources

Open-world fungal ITS embeddings improve higher-rank placement and cross-view retrieval, but percent identity is stronger on a matched ITS2 novelty benchmark

1. Learned sequence embeddings are increasingly proposed for fungal ITS classification and novelty detection, but are usually evaluated against weaker versions of themselves or default-configured alignment, by AUROC alone, and with queries treated as independent. We asked whether a purpose-trained encoder outperforms percent identity when both score identical queries under a firewalled open-world design, and which evaluation choices decide the answer. 2. From the UNITE release of 19 February 2025 we built ITS-core, ITS1 and ITS2 views and a genus-separated split with sealed calibration and test partitions. A convolutional encoder trained with genus-proxy, cross-view, hierarchical and episodic objectives was frozen and hash-sealed before test data were opened. We compared it with exhaustive, coverage-filtered VSEARCH identity on the same queries and references using genus-clustered paired intervals, and logged every post-opening correction before computing corrected metrics. 3. Three evaluation choices changed the comparison. Default VSEARCH heuristics returned a lower-identity hit than exhaustive search for 48.3% of benchmark queries; without a coverage filter, identity placed only 64.9% of novel ITS-core queries in the correct class, against 98.7% with it, because the conserved 5.8S let partial alignments win; and the encoder's embeddings depended on inference batching. Once corrected, identity exceeded the encoder in known-genus accuracy and novel-family placement at every view (family 80.0% to 84.6% against 53.8% to 64.8%) and detected more novel genera at a lower false-novelty rate. The encoder's novelty AUROC was within 0.023 of identity's, and no genus-clustered interval excluded zero. Identity's own development-to-test gap exceeded the encoder's, so that gap reflects partition composition rather than selection. Of the genera held out by a historical benchmark, 85% had been present in training; on a leakage-safe version, identity's AUROC advantage was 0.065, with a genus-clustered interval excluding zero. 4. Correctly configured alignment matched or exceeded the learned embedding throughout. The decisive results came from the evaluation, not the representation: baseline configuration, batch invariance, genus-level uncertainty, a parameter-free control for selection and a leakage audit each changed a conclusion, and each is inexpensive to apply to any learned barcode method.

bioinformatics↗

CurateMake: an auditable workflow for multi-source ITS reference database harmonisation and phylogenetic validation

O_LIReference databases shape the taxonomic resolution, uncertainty, and reproducibility of metabarcoding analyses. For ITS barcodes, public references are distributed across repositories with different taxonomic conventions, geographic coverage, and annotation practices, creating conflicts, missing ranks, and misannotations when databases are merged or compared. C_LIO_LIWe introduce CurateMake, a reproducible Snakemake workflow for ITS reference database construction, harmonisation, and validation. It integrates four public sources (UNITE, BOLD, PLANiTS, and CALeDNA) and user-supplied databases, combines Catalogue of Life name harmonisation with ITSx-based region standardisation, MSA/HMM-based alignment grouping, and SATIVA phylogenetic validation. Raw, CoL-harmonised, and SATIVA-validated annotation layers are retained throughout to compare curation effects while preserving flagged records for review. C_LIO_LIWe evaluated CurateMake on 3.58 million ingested sequences and controlled error-injection simulations. ITSx expanded the final harmonised database to 5.19 million barcode-resolved entries by recovering ITS1 and ITS2 sub-regions from full-length ITS records. Across the full dataset, normalised intra-cluster entropy decreased from Raw to CoL-harmonised to SATIVA-validated annotations, consistent with improved taxonomic coherence. In simulations, CurateMake achieved the highest correction rate across 1%-50% corruption and, at 15% corruption, corrected 42% {+/-} 1% of introduced errors, compared with 28% {+/-} 1% for CoL alone and 0% for SATIVA without the workflows alignment infrastructure. C_LIO_LIThese results show that nomenclatural harmonisation and phylogeny-informed validation address complementary error classes, with phylogenetic validation contributing measurably only within taxon-coherent alignments in this benchmark. CurateMake therefore provides a reproducible, provenance-tracked framework for auditable ITS reference database curation in metabarcoding workflows. C_LI

bioinformatics↗