bioRxiv Science⌕ Search

Biology subjects

Feer, L.

Publications and source records attributed to Feer, L..

3 recordsLinked to original sources

Uncovering Cas9 PAM diversity through metagenomic mining and machine learning

Recognition of protospacer adjacent motifs (PAMs) is crucial for target site recognition by CRISPR-Cas systems. In genome editing applications, the requirement for specific PAM sequences at the target locus imposes substantial constraints, driving efforts to search for novel Cas9 orthologs with extended or alternative PAM compatibilities. Here, we present CRISPR-PAMdb, a comprehensive and publicly accessible database compiling Cas9 protein sequences from 3.8 million bacterial and archaeal genomes and PAM profiles from 7.4 million phage and plasmid sequences. Through spacer-protospacer alignment, we inferred consensus PAM preferences for 8,003 unique Cas9 clusters. To extend PAM discovery beyond traditional alignment-based approaches, we developed CICERO, a machine learning model predicting PAM preferences directly from Cas9 protein sequences. Built on the ESM2 protein language model and trained on the CRISPR-PAMdb database, CICERO achieved an average accuracy of 0.68 on test data and 0.75 on experimentally validated Cas9 orthologs. For Cas9 clusters where alignment-based predictions were infeasible, CICERO generated PAM profiles for an additional 50,308 Cas9 proteins, including 17,453 high-confidence predictions with accuracies above 0.86. CRISPR-PAMdb, alongside CICERO models, enables large-scale exploration of PAM diversity across Cas9 proteins, accelerating design of next-generation CRISPR-Cas9 tools for precise genome engineering.

synthetic biology↗

Monosaccharides Drive Salmonella Gut Colonization in a Context-Dependent Manner

The carbohydrates that fuel gut colonization by S. Typhimurium are not fully known. To investigate this, we designed a quality-controlled mutant pool to probe the metabolic capabilities of this enteric pathogen. Using WISH-barcoding, we tested 35 metabolic mutants across five different mouse models, allowing us to differentiate between context-dependent and context-independent nutrient sources. Results showed that S. Typhimurium uses D-glucose, D-mannose, D-fructose, and D-galactose as context-independent carbohydrates across all models. The utilization of N-acetylglucosamine and hexuronates, on the other hand, was context-dependent. Furthermore, we showed that D-fructose is important in strain-to-strain competition between Salmonella serovars. Complementary experiments confirmed that D-glucose, D-fructose, and D-galactose are excellent niches for S. Typhimurium to exploit during colonization. Quantitative measurements revealed sufficient amounts of D-glucose and D-galactose in the murine cecum to drive S. Typhimurium colonization. Understanding these key substrates and their context-dependent use by enteric pathogens will inform the future design of probiotics and therapeutics to prevent diarrheal infections such as non-typhoidal salmonellosis.

microbiology↗

mBARq: a versatile and user-friendly framework for the analysis of DNA barcodes from transposon insertion libraries, knockout mutants and isogenic strain populations

DNA barcoding has become a powerful tool for assessing the fitness of strains in a variety of studies, including random transposon mutagenesis screens, attenuation of site-directed mutants, and population dynamics of isogenic strain pools. However, the statistical analysis, visualization and contextualization of the data resulting from such experiments can be complex and require bioinformatic skills. Here, we developed mBARq, a user-friendly tool designed to simplify these steps for diverse experimental setups. The tool is seamlessly integrated with an intuitive web app for interactive data exploration via the STRING and KEGG databases to accelerate scientific discovery.

bioinformatics↗