bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Inferring biochemical reactions and metabolite structures to cope with metabolic pathway drift

Inferring genome-scale metabolic networks in emerging model organisms is challenging because of incomplete biochemical knowledge and incomplete conservation of biochemical pathways during evolution. This limits the possibility to automatically transfer knowledge from well-established model organisms. Therefore, specific bioinformatic tools are necessary to infer new biochemical reactions and new metabolic structures that can be checked experimentally. Using an integrative approach combining both genomic and metabolomic data in the red algal model Chondrus crispus, we show that, even metabolic pathways considered as conserved, like sterol or mycosporine-like amino acids (MAA) synthesis pathways, undergo substantial turnover. This phenomenon, which we formally define as \"metabolic pathway drift\", is consistent with findings from other areas of evolutionary biology, indicating that a given phenotype can be conserved even if the underlying molecular mechanisms are changing. We present a proof of concept with a new methodological approach to formalize the logical reasoning necessary to infer new reactions and new molecular structures, based on previous biochemical knowledge. We use this approach to infer previously unknown reactions in the sterol and MAA pathways.\n\nAuthor summaryGenome-scale metabolic models describe our current understanding of all metabolic pathways occuring in a given organism. For emerging model species, where few biochemical data are available about really occurring enzymatic activities, such metabolic models are mainly based on transferring knowledge from other more studied species, based on the assumption that the same genes have the same function in the compared species. However, integration of metabolomic data into genome-scale metabolic models leads to situations where gaps in pathways cannot be filled by known enzymatic reactions from existing databases. This is due to structural variation in metabolic pathways accross evolutionary time. In such cases, it is necessary to use complementary approaches to infer new reactions and new metabolic intermediates using logical reasoning, based on available partial biochemical knowledge. Here we present a proof of concept that this is feasible and leads to hypotheses that are precise enough to be a starting point for new experimental work.

systems biology

Using transcriptome sequencing and pooled exome capture to study local adaptation in the giga-genome of Pinus cembra

Despite decreasing sequencing costs, whole-genome sequencing for population-based genome scans for selection is still prohibitively expensive for organisms with large genomes. Moreover, the repetitive nature of large genomes often represents a challenge in bioinformatic and downstream analyses. Here we use in-depth transcriptome sequencing to design probes for exome capture in Swiss stone pine (Pinus cembra), a conifer with an estimated genome size of 29.3 Gbp and no reference genome available. We successfully applied around 55,000 self-designed probes, targeting 25,000 contigs, to DNA pools of seven populations from the Swiss Alps and identified > 140,000 SNPs in around 13,000 contigs. The probes performed equally well in pools of the closely related species Pinus sibirica; in both species, more than 70% of the targeted contigs were sequenced at a depth [≥] 40x, i.e. the number of haplotypes in the pool. However, a thorough analysis of individually sequenced P. cembra samples indicated that a majority of the contigs (63%) represented multi-copy genes. We therefore removed paralogous contigs based on heterozygote excess and deviation from allele balance. Without putatively paralogous contigs, allele frequencies of population pools represented accurate estimates of individually determined allele frequencies. Using population genetic and landscape genomic methods, we show that inferences of neutral and adaptive genetic variation may be biased when not accounting for such multi-copy genes. Future studies should therefore put more emphasis on identifying paralogous loci, which will be facilitated by the establishment of additional high-quality reference genomes.

genomics

Rational Discovery of Dual-Action Multi-Target Kinase Inhibitor for Precision Anti-Cancer Therapy Using Structural Systems Pharmacology

Although remarkable progresses have been made in the cancer treatment, existing anti-cancer drugs are associated with increasing risk of heart failure, variable drug response, and acquired drug resistance. To address these challenges, for the first time, we develop a novel genome-scale multi-target screening platform 3D-REMAP that integrates data from structural genomics and chemical genomics as well as synthesize methods from structural bioinformatics, biophysics, and machine learning. 3D-REMAP enables us to discover marked drugs for dual-action agents that can both reduce the risk of heart failure and present anti-cancer activity. 3D-REMAP predicts that levosimendan, a drug for heart failure, inhibits serine/threonine-protein kinase RIOK1 and other kinases. Subsequent experiments confirm this prediction, and suggest that levosimendan is active against multiple cancers, notably lymphoma, through the direct inhibition of RIOK1 and RNA processing pathway. We further develop machine learning models to identify cancer cell-lines and patients that may respond to levosimendan. Our findings suggest that levosimendan can be a promising novel lead compound for the development of safe and effective multi-targeted cancer therapy, and demonstrate the potential of genome-wide multi-target screening in designing polypharmacology and drug repurposing for precision medicine.\n\nAuthor SummaryMulti-target drug design (a.k.a targeted polypharmacology) has emerged as a new strategy for discovering novel therapeutics that can enhance therapeutic efficacy and overcome drug resistance in tackling multi-genic diseases such as cancer. However, it is extremely challenging for conventional computational tools that are either receptor-based or ligand-based to screen compounds for selectively targeting multiple receptors. Existing multi-target drug design mainly focuses on compound screening against receptors within the same gene family but not across different gene families. Here, we develop a new computational tool 3D-REMAP that enables us to identify chemical-protein interactions across fold space on a genome scale. The genome-scale chemical-protein interaction network allows us to discover dual-action drugs that can bind to two types of targets simultaneously, one for mitigating side effect and another for enhancing the therapeutic effect. Using 3D-REMAP, we predict and subsequently experiments validate that levosimendan, a drug for heart failure, is active against multiple cancers, notably, lymphoma. This study demonstrates the potential of genome-wide multi-target screening in designing polypharmacology and drug repurposing for precision medicine.

systems biology

The NDV-3A vaccine protects mice from multidrug resistant Candida auris infection

Candida auris is an emerging, multi-drug resistant, health care-associated fungal pathogen. Its predominant prevalence in hospitals and nursing homes indicates its ability to adhere to and colonize the skin, or persist in an environment outside the host - a trait unique from other Candida species. Besides being associated globally with life-threatening disseminated infections, C. auris also poses significant clinical challenges due to its ability to adhere to polymeric surfaces and form highly drug-resistant biofilms. Here, we performed bioinformatic studies to identify the presence of adhesin proteins in C. auris, with sequence as well as 3-D structural homologies to the major adhesin/invasin of C. albicans, Als3. Anti-Als3p antibodies generated by vaccinating mice with NDV-3A (a vaccine based on the N-terminus of Als3 protein formulated with alum) recognized C. auris in vitro, blocked its ability to form biofilms and enhanced macrophage-mediated killing of the fungus. Furthermore, NDV-3A vaccination induced significant levels of C. auris cross-reactive humoral and cellular immune responses, and protected immunosuppressed mice from lethal C. auris disseminated infection, compared to the control alum-vaccinated mice. Finally, NDV-3A potentiated the protective efficacy of the antifungal drug micafungin, against C. auris candidemia. Identification of Als3-like adhesins in C. auris makes it a target for immunotherapeutic strategies using NDV-3A, a vaccine with known efficacy against other Candida species and safety as well as efficacy in clinical trials. Considering that C. auris can be resistant to almost all classes of antifungal drugs, such an approach has profound clinical relevance.\n\nAuthor SummaryCandida auris has emerged as a major health concern to hospitalized patients and nursing home subjects. C. auris strains display multidrug resistance to current antifungal therapy and cause lethal infections. We have determined that C. auris harbors homologs of C. albicans Als cell surface proteins. The C. albicans NDV-3A vaccine, harboring the N-terminus of Als3p formulated with alum, generates cross-reactive antibodies against C. auris clinical isolates and protects neutropenic mice from hematogenously disseminated C. auris infection. Importantly, the NDV-3A vaccine displays an additive protective effect in neutropenic mice when combined with micafungin. Due to its proven safety and efficacy in humans against C. albicans infection, our studies support the expedited testing of the NDV-3A vaccine against C. auris in future clinical trials.

microbiology

Differential binding cell-SELEX: method for identification of cell specific aptamers using high throughput sequencing

Aptamers have evolved as a viable alternative to antibodies in recent years. High throughput sequencing (HTS) has revolutionized the aptamer research by increasing the number of reads from few using Sanger sequencing to millions of reads using HTS approach. Despite the availability and advantages of HTS compared to Sanger sequencing there are only 50 aptamer HTS sequencing samples available on public databases. HTS data for aptamer research are mostly used to compare sequence enrichment between subsequent selection cycles. This approach does not take full advantage of HTS because enrichment of sequences during selection can be due to inefficient negative selection when using live cells. Here we present differential binding cell-SELEX (systematic evolution of ligands by exponential enrichment) workflow that adapts FASTAptamer toolbox and bioinformatics tool edgeR that is mainly used for functional genomics to achieve more informative metrics about the selection process. We propose fast and practical high throughput aptamer identification method to be used with cell-SELEX technique to increase successful aptamer selection rate against live cells. The feasibility of our approach is demonstrated by performing aptamer selection against clear cell renal cell carcinoma (ccRCC) RCC-MF cell line using RC-124 cell line from healthy kidney tissue for negative selection.

molecular biology

Composition of the Survival Motor Neuron (SMN) complex in Drosophila melanogaster

Spinal Muscular Atrophy (SMA) is caused by homozygous mutations in the human survival motor neuron 1 (SMN1) gene. SMN protein has a well-characterized role in the biogenesis of small nuclear ribonucleoproteins (snRNPs), core components of the spliceosome. SMN is part of an oligomeric complex with core binding partners, collectively called Gemins. Biochemical and cell biological studies demonstrate that certain Gemins are required for proper snRNP assembly and transport. However, the precise functions of most Gemins are unknown. To gain a deeper understanding of the SMN complex in the context of metazoan evolution, we investigated the composition of the SMN complex in Drosophila melanogaster. Using a stable transgenic line that exclusively expresses Flag-tagged SMN from its native promoter, we previously found that Gemin2, Gemin3, Gemin5, and all nine classical Sm proteins, including Lsm10 and Lsm11, co-purify with SMN. Here, we show that CG2941 is also highly enriched in the pulldown. Reciprocal co-immunoprecipitation reveals that epitope-tagged CG2941 interacts with endogenous SMN in Schneider2 cells. Bioinformatic comparisons show that CG2941 shares sequence and structural similarity with metazoan Gemin4. Additional analysis shows that three other genes (CG14164, CG31950 and CG2371) are not orthologous to Gemins 6-7-8, respectively, as previously suggested. In D.melanogaster, CG2941 is located within an evolutionarily recent genomic triplication with two other nearly identical paralogous genes (CG32783 and CG32786). RNAi-mediated knockdown of CG2941 and its two close paralogs reveals that Gemin4 is essential for organismal viability.

genetics

A database of egg size and shape from more than 6,700 insect species

1Offspring size is a fundamental trait in disparate biological fields of study. This trait can be measured as the size of plant seeds, animal eggs, or live young, and it influences ecological interactions, organism fitness, maternal investment, and embryonic development. Although multiple evolutionary processes have been predicted to drive the evolution of offspring size, the phylogenetic distribution of this trait remains poorly understood, due to the difficulty of reliably collecting and comparing offspring size data from many species. Here we present a database of 10,449 morphological descriptions of insect eggs, with records for 6,706 unique insect species and representatives from every extant hexapod order. The dataset includes eggs whose volumes span more than eight orders of magnitude. We created this database by partially automating the extraction of egg traits from the primary literature. In the process, we overcame challenges associated with large-scale phenotyping by designing and employing custom bioinformatic solutions to common problems. We matched the taxa in this database to the currently accepted scientific names in taxonomic and genetic databases, which will facilitate the use of this data for testing pressing evolutionary hypotheses in offspring size evolution.

evolutionary biology

Detection of low-frequency mutations and removal of heat-induced artifactual mutations using Duplex Sequencing

We present a genome-wide comparative and comprehensive analysis of three different sequencing methods (conventional next generation sequencing (NGS), tag-based single strand sequencing (eg. SSCS), and Duplex Sequencing for investigating mitochondrial mutations in human breast epithelial cells. Duplex Sequencing produces a single strand consensus sequence (SSCS) and a duplex consensus sequence (DCS) analysis, respectively. Our study validates that although high-frequency mutations are detectable by all the three sequencing methods with the similar accuracy and reproducibility, rare (low-frequency) mutations are not accurately detectable by NGS and SSCS. Even with conservative bioinformatical modification to overcome the high error rate of NGS, the NGS frequency of rare mutations is 7.0x10-4. The frequency is reduced to 1.3x10-4 with SSCS and is further reduced to 1.0x10-5 using DCS. Rare mutation context spectra obtained from NGS significantly vary across independent experiments, and it is not possible to identify a dominant mutation context. In contrast, rare mutation context spectra are consistently similar in all independent DCS experiments. We have systematically identified heat-induced artifactual mutations and corrected the artifacts using Duplex Sequencing. All of these artifacts are stochastically occurring rare mutations. C>A/G>T, a signature of oxidative damage, is the most increased (170-fold) heat-induced artifactual mutation type. Our results strongly support the claim that Duplex Sequencing accurately detects low-frequency mutations and identifies and corrects artifactual mutations introduced by heating during DNA preparation.

genomics

Metagenomic Insights Unveil the Dominance of Undescribed Actinobacteria in Pond Ecosystem of an Indian Shrine

Metagenomic analysis holds immense potential for identifying rare and uncharacterized microorganisms from many ecological habitats. Actinobacteria have been proved to be an excellent source of novel antibiotics for several decades. The present study was designed to delineate and understand the bacterial diversity with special focus on Actinobacteria from pond sediment collected from Sanjeeviraya Hanuman Temple, Ayyangarkulam, Kanchipuram, Tamil Nadu, India. The sediment had an average temperature (25.32%), pH (7.13), salinity (0.960 mmhos/cm) and high organic content (10.7%) posing minimal stress on growth condition of the microbial community. Subsequent molecular manipulations, sequencing and bioinformatics analysis of V3 and V4 region of 16S rRNA metagenomics analysis confirmed the presence of 40 phyla, 100 classes, 223 orders, 319 families and 308 genera in the sediment sample dominated by Acidobacteria (18.14%), Proteobacteria (15.13%), Chloroflexi (12.34), Actinobacteria (10.84%), Cyanobacteria (5.58%), Verrucomicrobia (3.37%), Firmicutes (2.28%), and, Gemmatimonadetes (1.63%). Among the Actinobacteria phylum, Acidothermus (29.68%) was the predominant genus followed by Actinospica (17.65%), Streptomyces (14.64%), Nocardia (4.55%) and Sinomonas (2.9%). Culture-dependent isolation of Actinobacteria yielded all strains of similar morphology to that of Streptomyces genus which clearly indicating that the traditional based technique is incapable of isolating majority of the non-Streptomyces or the so called rare Actinobacteria. Although Actinobacteria were among the dominant phylum, a close look at the species level indicated that only 15.2% within the Actinobacterial phylum could be assigned to cultured species. This leaves a vast majority of the Actinobacterial species yet to be explored with possible novel metabolites have special pharmaceutical and industrial application. It also indicates that the microbial ecology of pond sediment is neglected fields which need attention.

microbiology

Identification of single nucleotide variants using position-specific error estimation in deep sequencing data

BackgroundTargeted deep sequencing is a highly effective technology to identify known and novel single nucleotide variants (SNVs) with many applications in translational medicine, disease monitoring and cancer profiling. However, identification of SNVs using deep sequencing data is a challenging computational problem as different sequencing artifacts limit the analytical sensitivity of SNV detection, especially at low variant allele frequencies (VAFs).\n\nMethodsTo address the problem of relatively high noise levels in amplicon-based deep sequencing data (e.g. with the Ion AmpliSeq technology) in the context of SNV calling, we have developed a new bioinformatics tool called AmpliSolve. AmpliSolve uses a set of normal samples to model position-specific, strand-specific and nucleotide-specific background artifacts (noise), and deploys a Poisson model-based statistical framework for SNV detection.\n\nResultsOur tests on both synthetic and real data indicate that AmpliSolve achieves a good trade-off between precision and sensitivity, even at VAF below 5% and as low as 1%. We further validate AmpliSolve by applying it to the detection of SNVs in 96 circulating tumor DNA samples at three clinically relevant genomic positions and compare the results to digital droplet PCR experiments.\n\nConclusionsAmpliSolve is a new tool for in-silico estimation of background noise and for detection of low frequency SNVs in targeted deep sequencing data. Although AmpliSolve has been specifically designed for and tested on amplicon-based libraries sequenced with the Ion Torrent platform it can, in principle, be applied to other sequencing platforms as well. AmpliSolve is freely available at https://github.com/dkleftogi/AmpliSolve.

genomics

Mutational analysis of N-ethyl-N-nitrosourea (ENU) in the fission yeast Schizosaccharomyces pombe.

Forward genetics has boosted our knowledge on genic function in a multitude of biological models and it has significantly contributed to the understanding of genetic bases of development, ageing and human diseases. With the advent of the next generation sequencing and use of powerful bioinformatic tools, this traditional genetic strategy has acquired a new impulse. At present, whole genome sequencing assisted by in silico analysis allows the rapid and efficient identification of gene variants that are responsible for a particular phenotype. In this experimental pipeline, it is crucial to start by provoking a large number of random changes in the genome of the model organisms to be screened. A range of chemical mutagens are used to this end because most of them display particular reactivity properties and act differently over DNA. Here we use N-ethyl-N-nitrosourea (ENU) as a mutagen for the first time to our knowledge in the fission yeast Schizosaccharomyces pombe. By comparison to the extensively used Ethyl methanesulfonate (EMS) in a phenotype-based study, we conclude that ENU is a very potent and easy-to-use mutagen. Judging from DNA sequence analysis of the identified mutants, ENU induces base changes rather than indels and the mutational spectrum in the fission yeast seems similar to that found in mice but different to that described in other single-celled organisms such as budding yeast and E. coli. Using ENU in S. pombe, we have gathered a collection of 49 auxotrophic mutants including two deleterious alleles of ATIC human ortholog. Defective alleles of this gene are causative of AICA-Ribosiduria, a severe genetic disease. We have also identified 5 aminoglycoside-resistance inactivating mutations in APH genes. All these mutations reported here may be of interest in the metabolism and antibiotic resistance research fields.

genetics

Genetics of adaptation of the ascomycetous fungus Podospora anserina to submerged cultivation

Podospora anserina is a model ascomycetous fungus which shows pronounced phenotypic senescence when grown on solid medium but possesses unlimited lifespan under submerged cultivation. In order to study the genetic aspects of adaptation of P. anserina to submerged cultivation, we initiated a long-term evolution experiment. In the course of the first four years of the experiment, 125 single-nucleotide substitutions and 23 short indels were fixed in eight independently evolving populations. Six proteins that affect fungal growth and development evolved in more than one population; in particular, the G-protein alpha subunit FadA evolved in seven out of eight experimental populations. Parallel evolution at the level of genes and pathways, an excess of nonsense and missense substitutions, and an elevated conservation of proteins and their sites where the changes occurred suggest that many of the observed allele replacements were adaptive and driven by positive selection.\n\nAuthor summaryLiving beings adapt to novel conditions that are far from their original environments in different ways. Studying mechanisms of adaptation is crucial for our understanding of evolution. The object of our interest is a multicellular fungus Podospora anserina. This fungus is known for its pronounced senescence and a definite lifespan, but it demonstrates an unlimited lifespan and no signs of senescence when grown under submerged conditions. Soon after transition to submerged cultivation, the rate of growth of P. anserina increases and its pigmentation changes. We wanted to find out whether there are any genetic changes that contribute to adaptation of P. anserina to these novel conditions and initiated a long-term evolutionary experiment on eight independent populations. Over the first four years of the experiment, 148 mutations were fixed in these populations. Many of these mutations lead to inactivation of the part of the developmental pathway in P. anserina, probably reallocating resources to vegetative proliferation in liquid medium. Our observations imply that strong positive selection drives changes in at least some of the affected protein-coding genes.\n\nData AvailabilityGenome sequence data have been deposited at DDBJ/ENA/GenBank under accessions QHKV00000000 (founder genotype A; version QHKV01000000) and QHKU00000000 (founder genotype B; version QHKU01000000), with the respective BioSample accessions SAMN09270751 and SAMN09270757, under BioProject PRJNA473312. Sequencing data have been deposited at the SRA with accession numbers SRR7233712-SRR7233727, under the same BioProject.\n\nFundingExperimental work and sequencing were supported by the Russian Foundation for Basic Research (grants no. 16-04-01845a and 18-04-01349a). Bioinformatic analysis was supported by the Russian Science Foundation (grant no. 16-14-10173). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

evolutionary biology

Molecular convergence and positive selection associated with the evolution of symbiont transmission mode in stony corals

Heritable symbioses are thought to be important for the maintenance of mutually beneficial relationships (1), and for facilitating major transitions in individuality, such as the evolution of the eukaryotic cell (2, 3). In stony corals, vertical transmission has evolved repeatedly (4), providing a unique opportunity to investigate the genomic basis of this complex trait. We conducted a comparative analysis of 25 coral transcriptomes to identify orthologous genes exhibiting both signatures of positive selection and convergent amino acid substitutions in vertically transmitting lineages. The frequency of convergence events tends to be higher among vertically transmitting lineages, consistent with the proposed role of natural selection in driving the evolution of convergent transmission mode phenotypes (5). Of the 10,774 total orthologous genes identified, 403 exhibited at least one molecular convergence event and evidence of positive selection in at least one vertically transmitting lineage. Functional enrichments among these top candidate genes include processes previously implicated in mediating the coral-Symbiodiniaceae symbiosis including endocytosis, immune response, cytoskeletal protein binding and cytoplasmic membrane-bounded vesicles (6). We also identified 100 genes showing evidence of positive selection at the particular convergence event. Among these, we identified several novel candidate genes, highlighting the value of our approach for generating new insight into the mechanistic basis of the coral symbiosis, in addition to uncovering host mechanisms associated with the evolution of heritable symbioses.\n\nDATA ARCHIVAL LOCATIONRaw sequencing data generated for this study have been uploaded to NCBIs SRA: PRJNA395352. All bioinformatic scripts and input files can be found at https://github.com/grovesdixon/convergent_evo_coral.

evolutionary biology

Comparative Transcriptomes Analysis of Taenia pisiformis at Different Development Stages

To understand the characteristics of the transcriptional group of Taenia pisiformis at different developmental stages, and to lay the foundation for the screening of vaccine antigens and drug target genes, the transcriptomes of adult and larva of T. pisiformis were assembled and analyzed using bioinformatic tools. A total of 36,951 unigenes with a mean length of 950bp were formed, among which 12,665, 8,188, 7,577, and 6,293 unigenes have been annotated respectively by sequence similarity analysis with four databases (NR, Swiss-Prot, KOG, and KEGG). It should be noted there are 5,662 unigenes that share good similarity with the four databases and get a relatively perfect functional annotation. Besides, a total of 10,247 differentially expressed genes were screened. To be specific, 6,910 unigenes were up-regulated in the larva stage while 3,337 were down-regulated in the adult stage. To sum up, this study sequenced and analyzed the transcriptomes of the larval and adult stages of T. pisiformis. The results of differentially expressed genes in these two stages could provide basis for functional genomics, immunology and gene expression profiles of T. pisiformis.

molecular biology

Real-time capture of horizontal gene transfers from gut microbiota by engineered CRISPR-Cas acquisition

Horizontal gene transfer (HGT) is central to the adaptation and evolution of bacteria. However, our knowledge about the flow of genetic material within complex microbiomes is lacking; most studies of HGT rely on bioinformatic analyses of genetic elements maintained on evolutionary timescales or experimental measurements of phenotypically trackable markers (e.g. antibiotic resistance). Consequently, our knowledge of the capacity and dynamics of HGT in complex communities is limited. Here, we utilize the CRISPR-Cas spacer acquisition process to detect HGT events from complex microbiota in real-time and at nucleotide resolution. In this system, a recording strain is exposed to a microbial sample, spacers are acquired from foreign transferred elements and permanently stored in genomic CRISPR arrays. Subsequently, sequencing and analysis of these spacers enables identification of the transferred elements. This approach allowed us to quantify transfer frequencies of individual mobile elements without the need for phenotypic markers or post-transfer replication. We show that HGT in human clinical fecal samples can be extensive and rapid, often involving multiple different plasmid types, with the IncX type being the most actively transferred. Importantly, the vast majority of transferred elements did not carry readily selectable phenotypic markers, highlighting the utility of our approach to reveal previously hidden real-time dynamics of mobile gene pools within complex microbiomes.

microbiology

METAGENOMIC SEQUENCING FOR COMBINED DETECTION OF RNA AND DNA VIRUSES IN RESPIRATORY SAMPLES FROM PAEDIATRIC PATIENTS

IntroductionViruses are the main cause of respiratory tract infections. Metagenomic next-generation sequencing (mNGS) enables the unbiased detection of all potential pathogens in a clinical sample, including variants and even unknown pathogens. To apply mNGS in viral diagnostics, there is a need for sensitive and simultaneous detection of RNA and DNA viruses. In this study, the performance of an in-house mNGS protocol for routine diagnostics of viral respiratory infections, with single tube DNA and RNA sample-pre-treatment and potential for automated pan-pathogen detection was studied.\n\nMaterials and MethodsThe sequencing protocol and bioinformatics analysis was designed and optimized including the optimal concentration of the spike-in internal controls equine arteritis virus (EAV) and phocine-herpes virus-1 (PhHV-1).The whole genome of PhHV-1 was sequenced and added to the NCBI database. Subsequently, the protocol was retrospectively validated using a selection of 25 respiratory samples with in total 29 positive and 346 negative PCR results, previously sent to the lab for routine diagnostics.\n\nResultsThe results demonstrated that our protocol using Illumina Nextseq 500 sequencing with 10 million reads showed high repeatability. The NCBI RefSeq database as opposed to the NCBI nucleotide database led to enhanced specificity of virus classification. A correlation was established between read counts and PCR cycle threshold value, demonstrating the semi-quantitative nature of viral detection by mNGS. The results as obtained by mNGS appeared condordant with PCR based diagnostics in 25 out of the 29 (86%) respiratory viruses positive by PCR and in 315 of 346 (91%) PCR-negative results. Viral pathogens only detected by mNGS, not present in the routine diagnostic workflow were influenza C, KI polyomavirus, and cytomegalovirus.\n\nConclusionsSensitivity and analytical specificity of this mNGS protocol was comparable with PCR and higher when considering off-PCR target viral pathogens. All potential viral pathogens were detected in one single test, while it simultaneously obtained detailed information on detected viruses.

microbiology

Uncovering the unexplored diversity of thioamidated ribosomal peptides in Actinobacteria using the RiPPER genome mining tool

The rational discovery of new specialized metabolites by genome mining represents a very promising strategy in the quest for new bioactive molecules. Ribosomally synthesized and post-translationally modified peptides (RiPPs) are a major class of natural product that derive from genetically encoded precursor peptides. However, RiPP gene clusters are particularly refractory to reliable bioinformatic predictions due to the absence of a common biosynthetic feature across all pathways. Here, we describe RiPPER, a new tool for the family-independent identification of RiPP precursor peptides and apply this methodology to search for novel thioamidated RiPPs in Actinobacteria. Until now, thioamidation was believed to be a rare post-translational modification, which is catalyzed by a pair of proteins (YcaO and TfuA) in Archaea. In Actinobacteria, the thioviridamide-like molecules are a family of cytotoxic RiPPs that feature multiple thioamides, and it has been proposed that a YcaO-TfuA pair of proteins also catalyzes their formation. Potential biosynthetic gene clusters encoding YcaO and TfuA protein pairs are common in Actinobacteria but the chemical diversity generated by these pathways is almost completely unexplored. A RiPPER analysis reveals a highly diverse landscape of precursor peptides encoded in previously undescribed gene clusters that are predicted to make thioamidated RiPPs. To illustrate this strategy, we describe the first rational discovery of a new family of thioamidated natural products, the thiovarsolins from Streptomyces varsoviensis.

microbiology

Uncultured marine cyanophages encode for active NblA, phycobilisome proteolysis adaptor protein

Phycobilisomes (PBS) are large water-soluble membrane-associated complexes in cyanobacteria and some chloroplasts that serve as a light-harvesting antennas for the photosynthetic apparatus. When short of nitrogen or sulfur, cyanobacteria readily degrade their phycobilisomes allowing the cell to replenish the vanishing nutrients. The key regulator in the degradation process is NblA, a small protein (~6 kDa) which recruits proteases to the PBS. It was discovered previously that not only do cyanobacteria possess nblA genes but also that they are encoded by genomes of some freshwater cyanophages. A recent study, using assemblies from oceanic metagenomes, revealed genomes of a novel uncultured marine cyanophage lineage which contain genes coding for the PBS degradation protein. Here, we examine the functionality of nblA-like genes from these marine cyanophages by testing them in a freshwater model cyanobacterial nblA knockout. One of the viral NblA variants could complement the non-bleaching phenotype and restore PBS degradation. Our findings reveal a functional NblA from a novel marine cyanophage lineage. Furthermore, we shed new light on the distribution of nblA genes in cyanobacteria and cyanophages.\n\nOriginality-Significance StatementThis is the first study to examine the distribution and function of nblA genes of marine cyanophage origin. We describe as well the distribution of nblA-like genes in marine cyanobacteria using bioinformatic methods.

microbiology