bioRxiv Science⌕ Search

Biology subjects

Pershing, N. L.

Publications and source records attributed to Pershing, N. L..

2 recordsLinked to original sources

A prophage-encoded sRNA limits lytic phage infection of adherent-invasive E. coli

Prophages are prevalent features of bacterial genomes that can reduce susceptibility to lytic phage infection, yet the mechanisms involved are often elusive. Here, we identify a small RNA (svsR) encoded by the lambdoid prophage NC-SV in adherent-invasive Escherichia coli (AIEC) strain NC101 that confers resistance to lytic coliphages. Comparative genomic analyses revealed that NC-SV-like prophages and svsR homologs are conserved across diverse Enterobacteriaceae. Transcriptional analyses reveal that svsR represses maltodextrin transport genes, including lamB, which encodes the outer membrane maltoporin LamB--a known receptor for multiple phages. Nutrient supplementation experiments show that maltodextrin enhances phage adsorption, while glucose suppresses it, consistent with established effects of these sugars on lamB expression. In vivo, we compared wild-type NC101 and a prophage-deletion strain (NC101{Delta}NC-SV) in mice to assess the impact of NC-SV on lytic phage susceptibility. Although intestinal E. coli densities remained stable across groups, animals colonized with NC101 exhibited markedly reduced phage burdens in both the intestinal lumen and mucosa compared to mice colonized with NC101{Delta}NC-SV. This reduced phage pressure was associated with increased dissemination of NC101 to extraintestinal tissues, including the spleen and liver. Together, these findings highlight a nutrient-responsive, prophage-encoded mechanism that protects AIEC from phage predation and may promote bacterial persistence and dissemination in the inflamed gut.

microbiology↗

A Comparison of Tokenization Impact in AttentionBased and State Space Genomic Language Models

Genomic language models have recently emerged as a new method to decode, interpret, and generate genetic sequences. Existing genomic language models have utilized various tokenization methods, including character tokenization, overlapping and non-overlapping k-mer tokenization, and byte-pair encoding, a method widely used in natural language models. Genomic sequences differ from natural language because of their low character variability, complex and overlapping features, and inconsistent directionality. These features make sub-word tokenization in genomic language models significantly different from both traditional language models and protein language models. This study explores the impact of tokenization in genomic language models by evaluating their downstream performance on forty-four classification fine-tuning tasks. We also perform a direct comparison of byte pair encoding and character tokenization in Mamba, a state-space model. Our results indicate that character tokenization outperforms sub-word tokenization methods on tasks that rely on nucleotide level resolution, such as splice site prediction and promoter detection. While byte-pair tokenization had stronger performance on the SARS-CoV-2 variant classification task, we observed limited statistically significant differences between tokenization methods on the remaining downstream tasks.

bioinformatics↗