bioRxiv Science⌕ Search

Biology subjects

Brenner, E. P.

Publications and source records attributed to Brenner, E. P..

3 recordsLinked to original sources

Genotype-phenotype modeling of light ecotypes in Prochlorococcusreveals genomic signatures of ecotypic divergence

Prochlorococcus species are the most abundant marine photosynthetic bacteria. Despite broadly shared phenotypic traits and marine habitats, they exhibit remarkable genomic diversity. We ask what genomic signatures underlie its ecotypic divergence into high- and low-light adapted lineages, and whether these signatures can still be recovered from incomplete assemblies. From [~]1,000 publicly available Prochlorococcus genomes, we focused on those with information on their light adaptation ecotype (high-light/low-light), phylogenetic clades, and depth of isolation. Across these divisions, we calculated average nucleotide identity and constructed pangenomes to assess cyanobacterial core genes vs. those that separate ecotypes. Despite scant conservation, we observe a sharp taxon separation by light ecotypes. Classical machine learning models trained to predict ecotype achieve near-perfect binary classification accuracy even when predicting on partial genomes (Matthews Correlation Coefficient = 0.86 - 1.00), while regression models trained to predict the depth of isolation performed poorly, with high root mean square error values (37.6 - 42.0m). For ecotype prediction, we analyzed top gene features across model runs and classes; these features included photosynthesis-associated genes and pathways, as well as many novel markers of unknown function. When separating ecotypes further by previously described phylogenetic clades, genomic content and composition show even clearer separation among clades, supporting the taxonomic breadth of the Prochlorococcus collective. These results emphasize the genomic specialization underlying ecotypic divergence and support the utility of ML approaches for cyanobacterial ecotype prediction from metagenomic data. Expanded sampling will yield novel clade-specific biology. All data, models, and results are available on GitHub: https://github.com/JRaviLab/cyano_adaptation. ImportanceProchlorococcus are common aquatic cyanobacteria that can derive energy from light. They can be classified into high-/low-light ecotypes depending on how they use light. Prochlorococcus have small genomes compared to other bacteria, but the gene sets they carry are also remarkably flexible, which may help them survive and adapt to their harsh oceanic environment. We studied hundreds of Prochlorococcus genomes from around the world in an effort to predict ecotypes from partial genome sequences. We used comparative genomics, machine learning, and other statistical methods to identify genomic features associated with ecotypes. These statistical approaches predicted ecotypes accurately, reliably, and according to large differences in gene content and genome structure. Our results support that Prochlorococcus can be divided into different species or genera based on clades, and provide many gene targets for further research to understand cyanobacterial circadian rhythms or improve their bioengineering potential as chassis organisms.

bioinformatics↗

Mapping genetic and phenotypic diversity of Pseudomonas aeruginosa across clinical and environmental isolation sites

Pseudomonas aeruginosa is a clinically significant opportunistic pathogen adept at thriving in both host-associated and environmental settings. To define the extent to which P. aeruginosa isolates specialize across niches and identify genotype-phenotype correlates, we performed whole genome sequencing and comprehensive phenotypic characterization of 125 P. aeruginosa isolates from diverse clinical and environmental sites, evaluating virulence-associated traits, including motility, cytotoxicity, biofilm formation, pyocyanin production, and antimicrobial susceptibility. We identify that genomic diversity does not correlate with isolation source or most virulence phenotypes. Instead, we find that the two major P. aeruginosa clades (Groups A and B) delineate phylogeny and cytotoxicity, with Group B strains showing significantly higher cytotoxicity than Group A. Sequence analysis revealed previously uncharacterized alleles of genes encoding type III secretion effector proteins. We observed high variability amongst strains and isolation sources in all four assayed virulence phenotypes. Antimicrobial resistance (AMR) is exclusively observed in clinical isolates, not environmental, reflecting antibiotic exposure-driven selection. Bacterial GWAS revealed a statistically significant association between cytotoxicity and exoU presence, and we identified a novel exoU allelic variant with decreased cytotoxicity, demonstrating that functional diversity within well-characterized virulence factors may still influence pathogenic outcomes. In summary, our analyses of 125 diverse isolates suggest that the ability of P. aeruginosa to thrive across diverse niches is driven by broadly conserved genetic repertoire rather than niche-specific accessory genes. ImportancePseudomonas aeruginosa is a clinically significant opportunistic pathogen adept at thriving in both host-associated and environmental niches. A major gap in our understanding of this difficult-to-treat pathogen is whether niche specialization occurs in the context of human disease. Addressing this question is critical for guiding effective infection control strategies. Previous large-scale studies have focused solely on genotypic or phenotypic analyses; when paired, they have been limited to a single phenotypic assay or to a small number of isolates from one source, or relied on PCR-based methods targeting a restricted set of genes. To comprehensively uncover niche specialization and pathogenic versatility, we performed whole genome sequencing and phenotypic characterization of 5 virulence-associated traits including AST of 125 clinical and environmental P. aeruginosa isolates. Our systems-level findings challenge reductionist models of bacterial niche specialization, instead supporting an integrated view where conserved genomic systems enable opportunistic pathogenesis across diverse environments.

microbiology↗

From sequence to signature: Uncovering multiscale AMR features across bacterial pathogens with supervised machine learning

Since the clinical introduction of antibiotics in the 1940s, antimicrobial resistance (AMR) has become an increasingly dire threat to global public health. Pathogens acquire AMR much faster than we discover new drugs (antibiotics), warranting innovative methods to better understand its molecular underpinnings. Traditional approaches for detecting AMR in novel bacterial strains are time-consuming and labor-intensive. However, advances in sequencing technology offer a plethora of bacterial genome data, and computational approaches like machine learning (ML) provide an optimistic scope for in silico AMR prediction. Here, we introduce a comprehensive multiscale ML approach to predict AMR phenotypes and identify AMR molecular features associated with a single drug or drug family, stratified by time and geographical locations. As a case study, we focus on a subset of the World Health Organizations Bacterial Priority Pathogens, the frequently drug-resistant and nosocomial ESKAPE pathogens: Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, and Enterobacter species. We started with sequenced genomes with lab-derived AMR phenotypes, constructed pangenomes, clustered gene and protein sequences, and extracted protein domains to generate pangenomic features across molecular scales. To uncover the molecular mechanisms behind drug-/drug class-specific resistance, we trained logistic regression ML models on our datasets. These yielded ranked lists of AMR-associated genes, proteins, and domains. In addition to recapitulating known AMR features, our models identified novel candidates for experimental validation. The models were performant across molecular scales, data types, and drugs while achieving a median normalized Matthews correlation coefficient of 0.89. Prediction performance showed resilience even when evaluated on geographical and temporal holdouts. We also evaluated model generalizability and cross-resistance across the drug-/drug class-specific models cross-tested on other available drug-/drug class genomes. Finally, we uncovered multiple drug class resistance features using multiclass and multilabel models. Our holistic approach promises reliable prediction of existing and developing resistance in newly sequenced pathogen genomes, while pinpointing the mechanistic molecular contributors of AMR. All our models and results are available at our interactive web app, https://jravilab.org/amr.

bioinformatics↗