bioRxiv Science⌕ Search

Biology subjects

Merritt, B.

Publications and source records attributed to Merritt, B..

3 recordsLinked to original sources

Evaluating the Effectiveness of Parameter-Efficient Fine-Tuning in Genomic Classification Tasks

Foundation models are increasingly being leveraged for biological tasks. To address the high memory requirements of fine-tuning large pre-trained language models, parameter efficient fine-tuning (PEFT) methods are also being increasingly utilized. Previous studies have shown minimal, if any, loss in performance when using PEFT on binary classification tasks. However, the impact of using PEFT on tasks with large classification spaces has not been systemically evaluated. In this work, we apply PEFT to the problem of taxonomic classification using pre-trained genomic language models as the classification backbone. We explore various training strategies--including PEFT, full fine-tuning, and partial fine-tuning--for classifying sequences at the superkingdom, phylum, and genus levels. We find that PEFT-trained models significantly underperform compared to those trained via full fine-tuning or partial fine-tuning. Additionally, we demonstrate increased performance of pretrained models over those randomly initialized.

bioinformatics↗

TaxTriage: An Open-Source Metagenomic Sequencing Data Analysis Pipeline Enabling Putative Pathogen Detection

MotivationTaxTriage is a comprehensive pathogen identification workflow designed for both short- and long-read untargeted DNA and RNA sequencing data. Combining read classification, mapping, and de novo assembly approaches, putative pathogens are identified through comparisons to curated pathogens and abundance expectations from healthy cohort data. Flexible installation options are enabled using Nextflow (NF), including cloud deployment via NF Tower (Seqera Platform) and local installation on a variety of systems, including standalone installations without external internet access. Final analysis summaries are compiled into an Organism Discovery Report, which lists likely pathogens and supporting data, including a custom confidence score. ResultsEvaluation of published in silico, clinical, and outbreak datasets identified performance comparable to alternative cloud-based processing pipelines for expected pathogen and co-infection detection with similar sensitivity and increased specificity. To support both public health and veterinary diagnostics communities, customization options have been incorporated to enable improved performance for host species of interest. Availability and ImplementationSource code for TaxTriage is freely available at https://github.com/jhuapl-bio/taxtriage.

bioinformatics↗

Improved Resolution of Highly Pathogenic Avian Influenza Virus Haemagglutinin Cleavage Site Using Oxford Nanopore R10 Sequencing Chemistry

Highly pathogenic avian influenza viruses continue to pose global risks to One Health, including agriculture, public, and animal health. Rapid and accurate genomic surveillance is critical for monitoring viral mutations, tracing transmission, and guiding interventions in near real-time. Oxford Nanopore sequencing holds promise for real-time influenza genotyping, but data quality from R9 chemistry has limited its adoption due to challenges resolving low-complexity regions such as the biologically critical hemagglutinin cleavage site, a homopolymer of basic amino acids that distinguish highly pathogenic strains. In this study, human and avian influenza isolates (n=45) from Cambodia were sequenced using both R9.4.1 and R10.4.1 flow cells and chemistries to evaluate performance between approaches. Overall, R10.4.1 yielded increased data output with higher average quality compared to R9.4.1, producing improved consensus sequences using a reference-based bioinformatics approach. R10.4.1 had significantly lower minor population insertion and deletion frequencies, driven by improved performance in low sequence complexity regions prone to insertion and deletion errors, such as homopolymers. Within the hemagglutinin cleavage site, R10.4.1 resolved the correct motif in 90% of genomes compared to only 60% with R9.4.1. Further examination showed reduced frameshift mutations in consensus sequences generated with R10.4.1 that could result in incorrectly classified virulence. Improved consensus genome quality from nanopore sequencing approaches, especially across biologically important low-complexity regions, is critical to reduce subjective hand-curation and will improve local and global genomic surveillance responses.

genomics↗