bioRxiv Science⌕ Search

Biology subjects

Gemler, B. T.

Publications and source records attributed to Gemler, B. T..

3 recordsLinked to original sources

Inter-tool analysis of a NIST dataset for assessing baseline nucleic acid sequence screening

Nucleic acid synthesis is a dual-use technology that can benefit fields such as biology, medicine, and information storage. However, synthetic nucleic acids could also potentially be used negligently and ultimately cause harm, or be used with malicious intent to cause harm. Thus, this technology needs to be appropriately safeguarded. Sequence screening is one component of a biosecurity protocol for preventing such harm and consists of differentiating Sequences of Concern (SOCs) from benign sequences that are not associated with pathogenicity or toxicity. There exist many fit-for-purpose tools that have been developed for DNA synthesis sequence screening. However, questions remain regarding their performance with respect to consistency of screening. To aid in determining if screening tools are harmonized in regard to baseline sequence screening, NIST constructed a test dataset based on current screening recommendations. NIST then sent blinded datasets to sequence screening tool developers for testing. Overall, there was a general agreement between the tools and NIST assignments of the sequences and all tools had a baseline performance of greater than 95% sensitivity and 97% accuracy. Disagreement on specific sequences largely arose from single tools and could be traced to differences in defining a SOC and/or methodological differences in screening algorithms.

bioinformatics↗

Toward AI-Resilient Screening of Nucleic Acid Synthesis Orders: Process, Results, and Recommendations

Fast-moving advances in AI-assisted protein engineering are enabling breakthroughs in the life sciences that promise numerous beneficial applications. At the same time, these new capabilities are creating potential biosecurity challenges by providing new pathways to intentional or accidental synthesis of genes that encode hazardous proteins. The synthesis of nucleic acids is a key choke point in the AI-assisted protein engineering pipeline as it is where digital designs are transformed into physical instructions that can produce potentially harmful proteins. Thus, one focus for efforts to enhance biosecurity in the face of new AI-enabled capabilities is on bolstering the screening of orders by nucleic acid synthesis providers. We describe a multistakeholder, cross-sector effort to address biosecurity challenges with uses of AI-powered biological design tools to reformulate naturally occurring proteins of concern to create synthetic homologs that have low sequence identity to the wild-type proteins. We evaluated the abilities of traditional nucleic acid biosecurity screening tools to detect these synthetic homologs and found that, of tools tested, not all could previously detect such AI-redesigned sequences reliably. However, as we report, patches were built and deployed to improve detection rates over the course of the project, resulting in a final mean detection rate over tools of 97% of the synthetic homologs that were determined, using in-silico metrics, to be more likely to retain wild-type-like function. Finally, we make recommendations on approaches for studying and addressing the rising risk of adversarial AI-assisted protein engineering attacks like the one we identified and worked to mitigate.

synthetic biology↗

UltraSEQ: a universal bioinformatic platform for information-based clinical metagenomics and beyond

Applied metagenomics is a powerful emerging capability enabling untargeted detection of pathogens, and its application in clinical diagnostics promises to alleviate the limitations of current targeted assays. While metagenomics offers a hypothesis-free approach to identify any pathogen, including unculturable and potentially novel pathogens, its application in clinical diagnostics has so far been limited by workflow-specific requirements, computational constraints, and lengthy expert review requirements. To address these challenges, we developed UltraSEQ, a first-of its kind metagenomics-based clinical diagnostics and biosurveillance tool that is accurate and scalable. Here we present results for evaluation of our novel UltraSEQ pipeline using an in silico synthesized metagenome, mock microbial community datasets, and publicly available clinical datasets from samples of different infection types, and both short-read and long-read sequencing data. Our results show that UltraSEQ successfully detected all expected species across the tree of life in the in silico sample and detected all 10 bacterial and fungal species in the mock microbial community dataset. For clinical datasets, even without requiring dataset-specific configuration settings changes, background sample subtraction, or prior sample information, UltraSEQ achieved an overall accuracy of 91%. Further, we demonstrated UltraSEQs ability to provide accurate antibiotic resistance and virulence factor genotypes that are consistent with phenotypic results. Taken together, the above results demonstrates that the UltraSEQ platform offers a transformative approach to microbial and metagenomic sample characterization, employing a biologically informed detection logic, deep metadata, and a flexible system architecture for classification and characterization of taxonomic origin, gene function, and user-defined functions, including disease-causing infection.

bioinformatics↗