Analyzing Genomic Foundation Models for Viral Sequence Identification
Rapid and accurate identification of viral sequences underpins clinical diagnostics, epidemiological surveillance, and the safety testing of biological products, yet the established alignment-based methods such as BLAST are inherently closed-set, with accuracy tied to how well a query is already represented in the reference database. Recent advances in genomic foundation models (GFMs) have enabled alignment-free approaches to sequence classification. However, their performance relative to traditional methods remains incompletely characterised. We evaluated GFMs, specifically the Nucleotide Transformer (NT) and DNABERT families, for viral sequence classification using a dataset derived from the Reference Viral Database (RVDB). Sequence embeddings were generated with mean, max, and CLS pooling, classified by nearest-neighbour search in an embedding index and compared with BLAST as a reference. Matching BLAST provided a particularly stringent benchmark here, because as a closed-set method it compares each query directly against the exact reference sequences it stores, whereas a foundation model must rely on representations shaped by broad, general-purpose pre-training. On sequences of standard length, the best model, NTv3-650M, reached an accuracy of 0.93, approaching the 0.97 achieved by BLAST. On long genomic sequences, evaluated on a much smaller test set, the difference narrowed further, with NTv3-650M reaching 0.95 against 0.98 for BLAST. Robustness analyses showed that all GFMs were markedly more sensitive to base masking than BLAST. Both methods indexed the same labelled reference sequences, and BLAST compared queries with them directly at the nucleotide level, so approaching it from frozen, general-purpose representations without any parameter update is a meaningful result. NTv3-650M emerged as the most promising candidate for further optimisation and deployment in bioinformatics applications.