bioRxiv Science⌕ Search

Biology subjects

Salamzade, R.

Publications and source records attributed to Salamzade, R..

5 recordsLinked to original sources

Comparative genomic and metagenomic investigations of the Corynebacterium tuberculostearicum species complex reveals potential mechanisms underlying associations to skin health and disease

Corynebacterium are a diverse genus and dominant member of the human skin microbiome. Recently, we reported that the most prevalent Corynebacterium species found on skin - including Corynebacterium tuberculostearicum and Corynebacterium kefirresidentii - comprise a narrow species complex despite the diversity of the genus. Here, we apply high-resolution phylogenomics and comparative genomics to describe the structure of the C. tuberculostearicum species complex. We find this species complex is missing a fatty acid biosynthesis gene family which is often found in multi-copy in approximately 99% of other Corynebacterium species. Conversely, this species complex is enriched for multiple genetic traits, including a gene encoding for a collagen-like peptide. Further, through metagenomic investigations, we find that one species within the complex, C. kefirresidentii, increases in relative abundance during atopic dermatitis flares and show that most members of this species possess a colocalized set of putative virulence genes.

microbiology↗

Deep surveys of transcriptional modules with Massive Associative Kbiclustering (MAK)

Biclustering can reveal functional patterns in common biological data such as gene expression. Biclusters are ordered submatrices of a larger matrix that represent coherent data patterns. A critical requirement for biclusters is high coherence across a subset of columns, where coherence is defined as a fit to a mathematical model of similarity or correlation. Biclustering, though powerful, is NP-hard, and existing biclustering methods implement a wide variety of approximations to achieve tractable solutions for real world datasets. High bicluster coherence becomes more computationally expensive to achieve with high dimensional data, due to the search space size and because the number, size, and overlap of biclusters tends to increase. This complicates an already difficult problem and leads existing methods to find smaller, less coherent biclusters. Our unsupervised Massive Associative K-biclustering (MAK) approach corrects this size bias while preserving high bicluster coherence both on simulated datasets with known ground truth and on real world data without, where we apply a new measure to evaluate biclustering. Moreover, MAK jointly maximizes bicluster coherence with biological enrichment and finds the most enriched biological functions. Another long-standing problem with these methods is the overwhelming data signal related to ribosomal functions and protein production, which can drown out signals for less common but therefore more interesting functions. MAK reports the second-most enriched non-protein production functions, with higher bicluster coherence and arrayed across a large number of biclusters, demonstrating its ability to alleviate this biological bias and thus reflect the mediation of multiple biological processes rather than recruitment of processes to a small number of major cell activities. Finally, compared to the union of results from 11 top biclustering methods, MAK finds 21 novel S. cerevisiae biclusters. MAK can generate high quality biclusters in large biological datasets, including simultaneous integration of up to four distinct biological data types. Author summaryBiclustering can reveal functional patterns in common biological data such as gene expression. A critical requirement for biclusters is high coherence across a subset of columns, where coherence is defined as a fit to a mathematical model of similarity or correlation. Biclustering, though powerful, is NP-hard, and existing biclustering methods implement a wide variety of approximations to achieve tractable solutions for real world datasets. This complicates an already difficult problem and leads existing biclustering methods to find smaller and less coherent biclusters. Using the MAK methodology we can correct the bicluster size bias while preserving high bicluster coherence on simulated datasets with known ground truth as well as real world datasets, where we apply a new data driven bicluster set score. MAK jointly maximizes bicluster coherence with biological enrichment and finds more enriched biological functions, including other than protein production. These functions are arrayed across a large number of MAK biclusters, demonstrating ability to alleviate this biological bias and reflect the mediation of multiple biological processes rather than recruitment of processes to a small number of major cell activities. MAK can generate high quality biclusters in large biological datasets, including simultaneous integration of up to four distinct biological data types.

bioinformatics↗

lsaBGC provides a comprehensive framework for evolutionary analysis of biosynthetic gene clusters within focal taxa

We developed lsaBGC, a bioinformatics suite that introduces several new methods to expand on the available infrastructure for genomic and metagenomic-based comparative and evolutionary investigation of biosynthetic gene clusters (BGCs). Through application of the suite to four genera commonly found in skin microbiomes, we uncover multiple novel findings on the evolution and diversity of their BGCs. We show that the virulence associated carotenoid staphyloxanthin in Staphylococcus aureus is ubiquitous across the Staphylococcus genus but has largely been lost in the skin-commensal species Staphylococcus epidermidis. We further identify thousands of novel single nucleotide variants (SNVs) within BGCs from the Corynebacterium tuberculostearicum sp. complex, which we describe here to be a narrow, multi-species clade that features the most prevalent Corynebacterium in healthy skin microbiomes. Although novel SNVs were approximately ten times as likely to correspond to synonymous changes when located in the top five percentile of conserved sites, lsaBGC identified SNVs which defied this trend and are predicted to underlie amino acid changes within functionally key enzymatic domains. Ultimately, beyond supporting evolutionary investigations, lsaBGC provides important functionalities to aid efforts for the discovery or synthesis of natural products.

microbiology↗

vRhyme enables binning of viral genomes from metagenomes

Genome binning has been essential for characterization of bacteria, archaea, and even eukaryotes from metagenomes. Yet, no approach exists for viruses. We developed vRhyme, a fast and precise software for construction of viral metagenome-assembled genomes (vMAGs). vRhyme utilizes single- or multi-sample coverage effect size comparisons between scaffolds and employs supervised machine learning to identity nucleotide feature similarities, which are compiled into iterations of weighted networks and refined bins. Using simulated viromes, we displayed superior performance of vRhyme compared to available binning tools in constructing more complete and uncontaminated vMAGs. When applied to 10,601 viral scaffolds from human skin, vRhyme advanced our understanding of resident viruses, highlighted by identification of a Herelleviridae vMAG comprised of 22 scaffolds, and another vMAG encoding a nitrate reductase metabolic gene, representing near-complete genomes post-binning. vRhyme will enable a convention of binning uncultivated viral genomes and has the potential to transform metagenome-based viral ecology.

bioinformatics↗

Inter-species geographic signatures for tracing horizontal gene transfer and long-term persistence of carbapenem resistance

BackgroundCarbapenem-resistant Enterobacterales (CRE) are an urgent global health threat. Inferring the dynamics of local CRE dissemination is currently limited by our inability to confidently trace the spread of resistance determinants to unrelated bacterial hosts. Whole genome sequence comparison is useful for identifying CRE clonal transmission and outbreaks, but high-frequency horizontal gene transfer (HGT) of carbapenem resistance genes and subsequent genome rearrangement complicate tracing the local persistence and mobilization of these genes across organisms. MethodsTo overcome this limitation, we developed a new approach to identify recent HGT of large, near-identical plasmid segments across species boundaries, which also allowed us to overcome technical challenges with genome assembly. We applied this to complete and near-complete genome assemblies to examine the local spread of CRE in a systematic, prospective collection of all CRE, as well as time- and species-matched carbapenem susceptible Enterobacterales, isolated from patients from four U.S. hospitals over nearly five years. ResultsOur CRE collection comprised a diverse range of species, lineages and carbapenem resistance mechanisms, many of which were encoded on a variety of promiscuous plasmid types. We found and quantified rearrangement, persistence, and repeated transfer of plasmid segments, including those harboring carbapenemases, between organisms over multiple years. Some plasmid segments were found to be strongly associated with specific locales, thus representing geographic signatures that make it possible to trace recent and localized HGT events. Functional analysis of these signatures revealed genes commonly found in plasmids of nosocomial pathogens, such as functions required for plasmid retention and spread, as well survival against a variety of antibiotic and antiseptics common to the hospital environment. ConclusionsCollectively, the framework we developed provides a clearer, high resolution picture of the epidemiology of antibiotic resistance importation, spread, and persistence in patients and healthcare networks.

genomics↗