bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

The UCSC Repeat Browser allows discovery and visualization of evolutionary conflict across repeat families

BackgroundNearly half the human genome consists of repeat elements, most of which are retrotransposons, and many of these sequences play important biological roles. However repeat elements pose several unique challenges to current bioinformatic analyses and visualization tools, as short repeat sequences can map to multiple genomic loci resulting in their misclassification and misinterpretation. In fact, sequence data mapping to repeat elements are often discarded from analysis pipelines. Therefore, there is a continued need for standardized tools and techniques to interpret genomic data of repeats. ResultsWe present the UCSC Repeat Browser, which consists of a complete set of human repeat reference sequences derived from the gold standard repeat database RepeatMasker. The UCSC Repeat Browser contains mapped annotations from the human genome to these references, and presents all of them as a comprehensive interface to facilitate work with repetitive elements. Furthermore, it provides processed tracks of multiple publicly available datasets of biological interest to the repeat community, including ChIP-SEQ datasets for KRAB Zinc Finger Proteins (KZNFs) - a family of proteins known to bind and repress certain classes of repeats. Here we show how the UCSC Repeat Browser in combination with these datasets, as well as RepeatMasker annotations in several non-human primates, can be used to trace the independent trajectories of species-specific evolutionary conflicts. ConclusionsThe UCSC Repeat Browser allows easy and intuitive visualization of genomic data on consensus repeat elements, circumventing the problem of multi-mapping, in which sequencing reads of repeat elements map to multiple locations on the human genome. By developing a reference consensus, multiple datasets and annotation tracks can easily be overlaid to reveal complex evolutionary histories of repeats in a single interactive window. Specifically, we use this approach to retrace the history of several primate specific LINE-1 families across apes, and discover several species-specific routes of evolution that correlate with the emergence and binding of KZNFs.

genomics

Identification of the bacterial biosynthetic gene clusters of the oral microbiome illuminates the unexplored social language of bacteria during health and disease

Small molecules are the primary communication media of the microbial world. Recent bioinformatics studies, exploring the biosynthetic gene clusters (BGCs) which produce many small molecules, have highlighted the incredible biochemical potential of the signaling molecules encoded by the human microbiome. Thus far, most research efforts have focused on understanding the social language of the gut microbiome, leaving crucial signaling molecules produced by oral bacteria, and their connection to health versus disease, in need of investigation. In this study, a total of 4,915 BGCs were identified across 461 genomes representing a broad taxonomic diversity of oral bacteria. Sequence similarity networking provided a putative product class for over 100 unclassified novel BGCs. The newly identified BGCs were cross-referenced against 254 metagenomes and metatranscriptomes derived from individuals with either good oral health, dental caries, or periodontitis. This analysis revealed 2,473 BGCs, which were differentially represented across the oral microbiomes associated with health versus disease. Co-abundance network analysis identified numerous inverse correlations between BGCs and specific oral taxa. These correlations were present in health, but greatly reduced in dental caries, which may suggest a defect in colonization resistance. Finally, corroborating mass spectrometry identified several compounds with homology to products of the predicted BGC classes. Together, these findings greatly expand the number of known biosynthetic pathways present in the oral microbiome and provide an atlas for experimental characterization of these abundant, yet poorly understood, molecules and socio-chemical relationships, which impact the development of caries and periodontitis, two of the worlds most common chronic diseases.\n\nIMPORTANCEThe healthy oral microbiome is symbiotic with the human host, importantly providing colonization resistance against potential pathogens. Dental caries and periodontitis are two of the worlds most common and costly chronic infectious diseases, and are caused by a localized dysbiosis of the oral microbiome. Bacterially produced small molecules, often encoded by BGCs, are the primary communication media of bacterial communities, and play a crucial, yet largely unknown, role in the transition from health to dysbiosis. This study provides a comprehensive mapping of the BGC repertoire of the human oral microbiome and identifies major differences in health compared to disease. Furthermore, BGC representation and expression is linked to the abundance of particular oral bacterial taxa in health versus dental caries and periodontitis. Overall, this study provides a significant insight into the chemical communication network of the healthy oral microbiome, and how it devolves in the case of two prominent diseases.

microbiology

GRID - Genomics of Rare Immune Disorders: a highly sensitive and specific diagnostic gene panel for patients with primary immunodeficiencies

Primary Immune disorders affect 15,000 new patients every year in Europe. Genetic tests are usually performed on a single or very limited number of genes leaving the majority of patients without a genetic diagnosis. We designed, optimised and validated a new clinical diagnostic platform called GRID, Genomics of Rare Immune Disorders, to screen in parallel 279 genes, including 2015 IUIS genes, known to be causative of Primary Immune disorders (PID). Validation to clinical standard using more than 58,000 variants in 176 PID patients shows an excellent sensitivity, specificity. The customised and automated bioinformatics pipeline prioritises and reports pertinent Single Nucleotide Variants (SNVs), INsertions and DELetions (INDELs) as well as Copy Number Variants (CNVs). An example of the clinical utility of the GRID panel, is represented by a patient initially diagnosed with X-linked agammaglobulinemia due to a missense variant in the BTK gene with severe inflammatory bowel disease. GRID results identified two additional compound heterozygous variants in IL17RC, potentially driving the altered phenotype.

genomics

Quantitative proteomics of the 2016 WHO Neisseria gonorrhoeae reference strains surveys vaccine candidates and antimicrobial resistance determinants

The sexually transmitted disease gonorrhea (causative agent: Neisseria gonorrhoeae) remains an urgent public health threat globally due to the repercussions on reproductive health, high incidence, widespread antimicrobial resistance (AMR), and absence of a vaccine. To mine gonorrhea antigens and enhance our understanding of gonococcal AMR at the proteome level, we performed the first large-scale proteomic profiling of a diverse panel (n=15) of gonococcal strains, including the 2016 World Health Organization (WHO) reference strains. These strains show all existing AMR profiles, previously described in regard to phenotypic and reference genome characteristics, and are intended for quality assurance in laboratory investigations. Herein, these isolates were subjected to subcellular fractionation and labeling with tandem mass tags coupled to mass spectrometry and multi-combinatorial bioinformatics. Our analyses detected 901 and 723 common proteins in cell envelope and cytoplasmic subproteomes, respectively. We identified nine novel gonorrhea vaccine candidates. Expression and conservation of new and previously selected antigens were investigated. In addition, established gonococcal AMR determinants were evaluated for the first time using quantitative proteomics. Six new proteins, WHO_F_00238, WHO_F_00635, WHO_F_00745, WHO_F_01139, WHO_F_01144, and WHO_F_01226, were differentially expressed in all strains, suggesting that they represent global proteomic AMR markers, indicate a predisposition toward developing or compensating gonococcal AMR, and/or act as new antimicrobial targets. Finally, phenotypic clustering based on the isolates defined antibiograms and common differentially expressed proteins yielded seven matching clusters between established and proteome-derived AMR signatures. Together, our investigations provide a reference proteomics databank for gonococcal vaccine and AMR research endeavors, which enables microbiological, clinical, or epidemiological projects and enhances the utility of the WHO reference strains.

microbiology

Direct PCR amplification of 16S rRNA genes offers accelerated bacterial identification using the MinION™ nanopore sequencer

Rapid identification of bacterial pathogens is crucial for appropriate and adequate antibiotic treatment, which significantly improves patient outcomes. 16S ribosomal RNA (rRNA) gene amplicon sequencing has proven to be a powerful strategy for diagnosing bacterial infections. We have recently established a sequencing method and bioinformatics pipeline for 16S rRNA gene analysis utilizing the Oxford Nanopore Technologies MinION sequencer. In combination with our taxonomy annotation analysis pipeline, the system enabled the molecular detection of bacterial DNA in a reasonable timeframe for diagnostic purposes. However, purification of bacterial DNA from specimens remains a rate-limiting step in the workflow. To further accelerate the process of sample preparation, we adopted a direct PCR strategy that amplifies 16S rRNA genes from bacterial cell suspensions without DNA purification. Our results indicate that differences in cell wall morphology significantly affect direct PCR efficiency and sequencing data. Notably, mechanical cell disruption preceding direct PCR was indispensable for obtaining an accurate representation of the specimen bacterial composition. Furthermore, 16S rRNA gene analysis of mock polymicrobial samples indicated that primer sequence optimization is required to avoid preferential detection of particular taxa and to cover a broad range of bacterial species. This study establishes a relatively simple workflow for rapid bacterial identification via MinIONTM sequencing, which reduces the turnaround time from sample to result, and provides a reliable method that may be applicable to clinical settings.

microbiology

Defining the RNA Interactome by Total RNA-Associated Protein Purification

UV crosslinking can be used to identify precise RNA targets for individual proteins, transcriptome-wide. We sought to develop a technique to generate reciprocal data, identifying precise sites of RNA-binding proteome-wide. The resulting technique, total RNA-associated protein purification (TRAPP), was applied to yeast (S. cerevisiae) and bacteria (E. coli). In all analyses, SILAC labelling was used to quantify protein recovery in the presence and absence of irradiation. For S. cerevisiae, we also compared crosslinking using 254 nm (UVC) irradiation (TRAPP) with 4-thiouracil (4tU) labelling combined with ~350 nm (UVA) irradiation (PAR-TRAPP). Recovery of proteins not anticipated to show RNA-binding activity was substantially higher in TRAPP compared to PAR-TRAPP. As an example of preferential TRAPP-crosslinking, we tested enolase (Eno1) and demonstrated its binding to tRNA loops in vivo. We speculate that many protein-RNA interactions have biophysical effects on localization and/or accessibility, by opposing or promoting phase separation for highly abundant protein. Homologous metabolic enzymes showed RNA crosslinking in S. cerevisiae and E. coli, indicating conservation of this property. TRAPP allows alterations in RNA interactions to be followed and we initially analyzed the effects of weak acid stress. This revealed specific alterations in RNA-protein interactions; for example, during late 60S ribosome subunit maturation. Precise sites of crosslinking at the level of individual amino acids (iTRAPP) were identified in 395 peptides from 155 unique proteins, following phospho-peptide enrichment combined with a bioinformatics pipeline (Xi). TRAPP is quick, simple and scalable, allowing rapid characterization of the RNA-bound proteome in many systems.

cell biology

Illumina-based sequencing framework for accurate detection and mapping of influenza virus defective interfering particle-associated RNAs

The mechanisms and consequences of defective interfering particle (DIP) formation during influenza virus infection remain poorly understood. The development of next generation sequencing (NGS) technologies has made it possible to identify large numbers of DIP-associated sequences, providing a powerful tool to better understand their biological relevance. However, NGS approaches pose numerous technical challenges including the precise identification and mapping of deletion junctions in the presence of frequent mutation and base-calling errors, and the potential for numerous experimental and computational artifacts. Here we detail an Illumina-based sequencing framework and bioinformatics pipeline capable of generating highly accurate and reproducible profiles of DIP-associated junction sequences. We use a combination of simulated and experimental control datasets to optimize pipeline performance and demonstrate the absence of significant artifacts. Finally, we use this optimized pipeline to generate a high-resolution profile of DIP-associated junctions produced during influenza virus infection and demonstrate how this data can provide insight into mechanisms of DIP formation. This work highlights the specific challenges associated with NGS-based detection of DIP-associated sequences, and details the computational and experimental controls required for such studies.

microbiology

Galectin-1 promotes the invasion of bladder cancer urothelia through their matrix milieu

The progression of carcinoma of the urinary bladder involves migration of cancer epithelia through their surrounding tissue matrix microenvironment. This was experimentally confirmed when a gender- and grade-diverse set of bladder cancer cell lines were cultured in pathomimetic three-dimensional laminin-rich environments. The high-grade cells, particularly female, formed multicellular invasive morphologies in 3D. In comparison, low- and intermediate-grade counterparts showed growth-restricted phenotypes. A proteomic approach combining mass spectrometry and bioinformatics analysis identified the estrogen-driven lactose-binding lectin Galectin-1 (GAL-1) as a putative candidate that could drive this invasion. Expression of LGALS1, the gene encoding GAL-1 showed an association with tumor grade progression in bladder cell lines. Immunohisto- and cyto-chemical experiments suggested greater extracellular levels of GAL-1 in 3D cultures of high-grade bladder cells and cancer tissues. High levels of GAL-1 associated with increased proliferation- and adhesion- of bladder cancer cells when grown on laminin-rich matrices. Pharmacological inhibition and Gal-1 knockdown in high-grade female cells decreased their adhesion to, and viability on, laminin-rich substrata. Higher GAL-1 also correlated with reduced E-cadherin and increased N-cadherin levels in consonance with a mesenchymal-like phenotype that we observed in 3D culture. The inhibition of GAL-1 reversed the stellate invasive phenotype to a more growth-restricted one in high-grade cells embedded within both basement-membrane-like and stromal collagenous matrix scaffolds. Finally, inhibition of GAL-1 specifically altered cell surface sialic acids, suggesting the mechanism by which the levels of GAL-1 may underlie the aggression and poor prognosis of invasive bladder cancer, especially in women.

cancer biology

Robust Estimation of the Phylogenetic Origin of Plastids Using a tRNA-Based Phyloclassifier

The trait of oxygenic photosynthesis was acquired by the last common ancestor of Archaeplastida through endosymbiosis of the cyanobacterial progenitor of modern-day plastids. Although a single origin of plastids by endosymbiosis is broadly supported, recent phylogenomic studies report contradictory evidence that plastids branch either early or late within the cyanobacterial Tree of Life. Here we describe CYANO-MLP, a general-purpose phyloclassifier of cyanobacterial genomes implemented using a Multi-Layer Perceptron. CYANO-MLP exploits consistent phylogenetic signals in bioinformatically estimated structure-function maps of tRNAs. CYANO-MLP accurately classifies cyanobacterial genomes into one of eight well-supported cyanobacterial clades in a manner that is robust to missing data, unbalanced data and variation in model specification. CYANO-MLP supports a late-branching origin of plastids: we classify 99.32% of 440 plastid genomes into one of two late-branching cyanobacterial clades with strong statistical support, and confidently assign 98.41% of plastid genomes to one late-branching clade containing unicellular starch-producing marine/freshwater diazotrophic Cyanobacteria. CYANO-MLP correctly classifies the chromatophore of Paulinella chromatophora and rejects a sister relationship between plastids and the early-branching cyanobacterium Gloeomargarita lithophora. We show that recently applied phylogenetic models and character recoding strategies fit cyanobacterial/plastid phylogenomic datasets poorly, because of heterogeneity both in substitution processes over sites and compositions over lineages.

evolutionary biology

Metatranscriptome profiling of the dynamic transcription of mRNA and sRNA of a probiotic Lactobacillus strain in human gut

Metatranscriptomic sequencing has recently been applied to study how pathogens and probiotics affect human gastrointestinal (GI) tract microbiota, which provides new insights into their mechanisms of action. In this study, metatranscriptomic sequencing was applied to deduce the in vivo expression patterns of an ingested Lactobacillus casei strain, which was compared with its in vitro growth transcriptomes. Extraction of the strain-specific reads revealed that transcripts from the ingested L. casei were increased, while those from the resident L. paracasei strains remained unchanged. Mapping of all metatranscriptomic reads and transcriptomic reads to L. casei genome showed that gene expression in vitro and in vivo differed dramatically. About 39% (1163) mRNAs and 45% (93) sRNAs of L. casei well-expressed were repressed after ingested into human gut. Expression of ABC transporter genes and amino acid metabolism genes was induced at day-14 of ingestion; and genes for sugar and SCFA metabolisms were activated at day-28 of ingestion. Moreover, expression of sRNAs specific to the in vitro log phase was more likely to be activated in human gut. Expression of rli28c sRNA with peaked expression during the in vitro stationary phase was also activated in human gut; this sRNA repressed L. casei growth and lactic acid production in vitro. These findings implicate that the ingested L. casei might have to successfully change its transcription patterns to survive in human gut, and the time-dependent activation patterns indicate a highly dynamic cross-talk between the probiotic and human gut including its microbe community.\n\nImportanceProbiotic bacteria are important in food industry and as model microorganisms in understanding bacterial gene regulation. Although probiotic functions and mechanisms in human gastrointestinal tract are linked to the unique probiotic gene expression, it remains elusive how transcription of probiotic bacteria is dynamically regulated after being ingested. Previous study of probiotic gene expression in human fecal samples has been restricted due to its low abundance and the presence of of closely related species. In this study, we took the advantage of the good depth of metatranscriptomic sequencing reads and developed a strain-specific read analysis method to discriminate the transcription of the probiotic Lactobacillus casei and those of its resident relatives. This approach and additional bioinformatics analysis allowed the first study of the dynamic transcriptome profiles of probiotic L casei in vivo. The novel findings indicate a highly regulated repression and dynamic activation of probiotic genome in human GI tract.

microbiology

A novel signature derived from immunoregulatory and hypoxia genes predicts prognosis in liver and five other cancers

BackgroundDespite much progress in cancer research, its incidence and mortality continue to rise. A robust biomarker that would predict tumor behavior is highly desirable and could improve patient treatment and prognosis.\n\nMethodsIn a retrospective bioinformatics analysis involving patients with liver cancer (n=839), we developed a prognostic signature consisting of 45 genes associated with tumor-infiltrating lymphocytes and cellular responses to hypoxia. From this gene set, we were able to identify a second prognostic signature comprised of 8 genes. Its performance was further validated in five other cancers: head and neck (n=520), renal papillary cell (n=290), lung (n=515), pancreas (n=178) and endometrial (n=370).\n\nFindingsThe 45-gene signature predicted overall survival in three liver cancer cohorts: hazard ratio (HR)=1.82, P=0.006; HR=1.84, P=0.008 and HR=2.67, P=0.003. Additionally, the reduced 8-gene signature was sufficient and effective in predicting survival in liver and five other cancers: liver (HR=2.36, P=0.0003; HR=2.43, P=0.0002 and HR=3.45, P=0.0007), head and neck (HR=1.64, P=0.004), renal papillary cell (HR=2.31, P=0.04), lung (HR=1.45, P=0.03), pancreas (HR=1.96, P=0.006) and endometrial (HR=2.33, P=0.003). Receiver operating characteristic analyses demonstrated both signatures superior performance over current tumor staging parameters. Multivariate Cox regression analyses revealed that both 45-gene and 8-gene signatures were independent of other clinicopathological features in these cancers. Combining the gene signatures with somatic mutation profiles increased their prognostic ability.\n\nConclusionsThis study, to our knowledge, is the first to identify a gene signature uniting both tumor hypoxia and lymphocytic infiltration as a prognostic determinant in six cancer types (n=2,712). The 8-gene signature can be used for patient risk stratification by incorporating hypoxia information to aid clinical decision making.

cancer biology

Global analysis of the RpaB regulon based on the positional distribution of HLR1 sequences and comparative differential RNA-Seq data

The transcription factor RpaB regulates the expression of genes encoding photosynthesis-associated proteins during light acclimation. The binding site of RpaB is the HLR1 motif, a pair of imperfect octameric direct repeats, separated by two random nucleotides. Here, we used high-resolution mapping data of transcriptional start sites (TSSs) in the model Synechocystis sp. PCC 6803 in conjunction with the positional distribution of HLR1 sites for the global prediction of the RpaB regulon. The results demonstrate that RpaB regulates the expression of more than 150 promoters, driving the transcription of protein-coding and non-coding genes and antisense transcripts under low light and upon the shift to high light when DNA binding activity is lost. Transcriptional activation by RpaB is achieved when the HLR1 motif is located 66 to 45 nt upstream, repression occurs when it is close to or overlapping the TSS. Selected examples were validated by multiple experimental approaches, including chromatin affinity purification, reporter gene, northern hybridization and electrophoretic mobility shift assays. We found that RpaB controls ssr2016/pgr5, which is involved in cyclic electron flow and state transitions; six out of nine ferredoxins; three of four FtsH proteases; gcvP/slr0293, encoding a crucial photorespiratory protein; and nirA and isiA for which we suggest cross-regulation with the transcription factors NtcA or FurA, respectively. In addition to photosynthetic gene functions, RpaB contributes to the control of genes affiliated with nitrogen assimilation, cofactor biosyntheses, the CRISPR system and the circadian clock, making it one of the most versatile regulators in cyanobacteria.\n\nSignificance StatementRpaB is a transcription factor in cyanobacteria and in the chloroplasts of several lineages of eukaryotic algae. Like other important transcription factors, the gene encoding RpaB cannot be deleted, making the study of deletion mutants impossible. Based on a bioinformatic approach, we increased the number of known genes controlled by RpaB by a factor of 5. Depending on the distance to the TSS, RpaB mediates transcriptional activation or repression. The high number and functional diversity among its target genes and co-regulation with other transcriptional regulators characterize RpaB as a regulatory hub.

microbiology

PP4-dependent HDAC3 dephosphorylation discriminates between axonal regeneration and regenerative failure

The molecular mechanisms discriminating between regenerative failure and success remain elusive. While a regeneration-competent peripheral nerve injury mounts a regenerative gene expression response in bipolar dorsal root ganglia (DRG) sensory neurons, a regeneration-incompetent central spinal cord injury does not. This dichotomic response offers a unique opportunity to investigate the fundamental biological mechanisms underpinning regenerative ability. Following a pharmacological screen with small molecule inhibitors targeting key epigenetic enzymes in DRG neurons we identified HDAC3 signalling as a novel candidate brake to axonal regenerative growth. In vivo, we determined that only a regenerative peripheral but not a central spinal injury induces an increase in calcium, which activates protein phosphatase 4 that in turn dephosphorylates HDAC3 thus impairing its activity and enhancing histone acetylation. Bioinformatics analysis of ex vivo H3K9ac ChIPseq and RNAseq from DRG followed by promoter acetylation and protein expression studies implicated HDAC3 in the regulation of multiple regenerative pathways. Finally, genetic or pharmacological HDAC3 inhibition overcame regenerative failure of sensory axons following spinal cord injury. Together, these data indicate that PP4-dependent HDAC3 dephosphorylation discriminates between axonal regeneration and regenerative failure.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=122 SRC=\"FIGDIR/small/446963_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (40K):\norg.highwire.dtl.DTLVardef@1bd577corg.highwire.dtl.DTLVardef@1ba990aorg.highwire.dtl.DTLVardef@195813corg.highwire.dtl.DTLVardef@57abc7_HPS_FORMAT_FIGEXP M_FIG Graphical AbstractFollowing central nervous system (CNS) spinal injury, protein phosphatase 4/2 activity is not induced since calcium levels remain unchanged compared to uninjured conditions. HDAC3 remains phosphorylated and occupies deacetylated chromatin contributing to its compaction inhibiting gene expression. Following peripheral nervous system (PNS) sciatic injury, protein phosphatase 4/2 activity is induced by calcium. HDAC3 is dephosphorylated leading to its inhibition and release from chromatin sites contributing to increase in histone acetylation and in the expression of regeneration associated genes (RAGs).\n\nC_FIG

neuroscience

MicroRNA-200c suppresses epithelial-mesenchymal transition of ovarian cancer by targeting cofilin-2

This study investigated the effects of microRNA-200c (miR-200c) and cofilin-2 (CFL2) in regulating epithelial-mesenchymal transition (EMT) in ovarian cancer. The level of miR-200c was lower in invasive SKOV3 cells than that in non-invasive OVCAR3 cells, whereas CFL2 showed the opposite trend. Bioinformatics analysis and dual-luciferase reporter gene assays indicated that CFL2 was a direct target of miR-200c. Furthermore, SKOV3 and OVCAR3 cells were transfected with miR-200c mimic or inhibitor, pCDH-CFL2 (CFL2 overexpression), or CFL2 shRNA (CFL2 silencing). MiR-200c inhibition and CFL2 overexpression resulted in elevated levels of both CFL2 and vimentin while reducing E-cadherin expression. They also increased ovarian cancer cell invasion and migration in vitro and in vivo and increased the tumor volumes. Conversely, miR-200c mimic and CFL2 shRNA exerted the opposite effects as those aforementioned. In addition, the effects of pCDH-CFL2 and CFL2 shRNA were reversed by the miR-200c mimic and inhibitor, respectively. This finding suggested that miR-200c could be a potential tumor suppressor by targeting CFL2 in the EMT process.

cancer biology

RERconverge: an R package for associating evolutionary rates with convergent traits

Motivation: When different lineages of organisms independently adapt to similar environments, selection often acts repeatedly upon the same genes, leading to signatures of convergent evolutionary rate shifts at these genes. With the increasing availability of genome sequences for organisms displaying a variety of convergent traits, the ability to identify genes with such convergent rate signatures would enable new insights into the molecular basis of these traits.\n\nResults: Here we present the R package RERconverge, which tests for association between relative evolutionary rates of genes and the evolution of traits across a phylogeny. RERconverge can perform associations with binary and continuous traits, and it contains tools for visualization and enrichment analyses of association results.\n\nAvailability: RERconverge source code, documentation, and a detailed usage walk-through are freely available at https://github.com/nclark-lab/RERconverge. Datasets for mammals, Drosophila, and yeast are available at https://bit.ly/2J2QBnj.\n\nContact: mchikina@pitt.edu\n\nSupplementary information: Supplementary information, containing detailed vignettes for usage of RERconverge, are available at Bioinformatics online.

evolutionary biology

Metagenomic characterization of the viral community of the South Scotia Ridge

Viruses are the most abundant biological entities in aquatic ecosystems and harbor an enormous genetic diversity. While their great influence on the marine ecosystems is widely acknowledged, current information about their diversity remains scarce. Aviral metagenomic analysis of two surfaces and one bottom water sample was conducted from sites on the South Scotia Ridge (SSR) near the Antarctic Peninsula, during the austral summer 2016. The taxonomic composition and diversity of the viral communities were investigated and a functional assessment of the sequences was determined. Phylotypic analysis showed that most viruses belonging to the order Caudovirales, in particular, the family Podoviridae (41.92-48.7%), which is similar to the viral communities from the Pacific Ocean. Functional analysis revealed a relatively high frequency of phage-associated and metabolism genes. Phylogenetic analyses of phage TerL and Capsid_NCLDV (nucleocytoplasmic large DNA viruses) marker genes indicated that many of the sequences associated with Caudovirales and NCLDV were novel and distinct from known complete phage genomes. High Phaeocystis globosa virus virophage (Pgvv) signatures were found in SSR area and complete and partial Pgvv-like were obtained which may have an influence on host-virus interactions in the area during summer. Our study expands the existing knowledge of viral communities and their diversities from the Antarctic region and provides basic data for further exploring polar microbiomes.\n\nImportanceIn this study, we used high-throughput sequencing and bioinformatics analysis to analyze the viral community structure and biodiversity of SSR in the open sea near the Antarctic Peninsula. The results showed that the SSR viromes are novel, oceanic-related viromes and a high proportion of sequence reads was classified as unknown. Among known virus counterparts, members of the order Caudovirales were most abundant which is consistent with viromes from the Pacific Ocean. In addition, phylogenetic analyses based on the viral marker genes (TerL and MCP) illustrate the high diversity among Caudovirales and NCLDV. Combining deep sequencing and a random subsampling assembly approach, a new Pgvv-like group was also found in this region, which may a signification factor regulating virus-host interactions.

microbiology

Differential gene expression, including Sjfs800, in Schistosoma japonicum females before, during, and after male-female pairing

Schistosomiasis is a prevalent but neglected tropical disease caused by parasitic trematodes of the genus Schistosoma, with the primary disease-causing species being S. haematobium, S. mansoni, and S. japonicum. Male-female pairing of schistosomes is necessary for sexual maturity and the production of a large number of eggs, which are primarily responsible for schistosomiasis dissemination and pathology. Here, we used microarray hybridization, bioinformatics, quantitative PCR, in situ hybridization, and gene silencing assays to identify genes that play critical roles in S. japonicum reproduction biology, particularly in vitellarium development, a process that affects male-female pairing, sexual maturation, and subsequent egg production. Microarray hybridization analyses generated a comprehensive set of genes differentially transcribed before and after male-female pairing. Although the transcript profiles of females were similar 16 and 18 days after host infection, marked gene expression changes were observed at 24 days. The 30 most abundantly transcribed genes on day 24 included those associated with vitellarium development. Among these, genes for female-specific 800 (fs800), eggshell precursor protein, and superoxide dismutase (cu-zn-SOD) were substantially upregulated. Our in situ hybridization results in female S. japonicum indicated that cu-zn-SOD mRNA was highest in the ovary and vitellarium, eggshell precursor protein mRNA was expressed in the ovary, ootype, and vitellarium, and Sjfs800 mRNA was observed only in the vitellarium, localized in mature vitelline cells. Knocking down the Sjfs800 gene in female S. japonicum by approximately 60% reduced the number of mature vitelline cells, decreased rates of pairing and oviposition, and decreased the number of eggs produced in each male-female pairing by about 50%. These results indicate that Sjfs800 is essential for vitellarium development and egg production in S. japonicum and suggest that Sjfs800 regulation may provide a novel approach for the prevention or treatment of schistosomiasis.\n\nAuthor SummarySchistosomiasis is a common but largely unstudied tropical disease caused by parasitic trematodes of the genus Schistosoma. The eggs of schistosomes are responsible for schistosomiasis transmission and pathology, and the production of these eggs is dependent on the pairing of females and males. In this study, we determined which genes in Schistosoma japonicum females were differentially expressed before and after pairing with males, identifying the 30 most abundantly expressed of these genes. Among these 30 genes, we further characterized those in female S. japonicum that were upregulated after pairing and that were related to reproduction and vitellarium development, a process that affects male-female pairing, sexual maturation, and subsequent egg production. We identified three such genes, S. japonicum female-specific 800 (Sjfs800), eggshell precursor protein, and superoxide dismutase, and confirmed that the mRNAs for these genes were primarily localized in reproductive structures. By using gene silencing techniques to reduce the amount of Sjfs800 mRNA in females by about 60%, we determined that Sjfs800 plays a key role in development of the vitellarium and egg production. This finding suggests that regulation of Sjfs800 may provide a novel approach to reduce egg counts and thus aid in the prevention or treatment of schistosomiasis.

genomics

Reticulate evolution in eukaryotes: origin and evolution of the nitrate assimilation pathway

Genes and genomes can evolve through interchanging genetic material, this leading to reticular evolutionary patterns. However, the importance of reticulate evolution in eukaryotes, and in particular of horizontal gene transfer (HGT), remains controversial. Given that metabolic pathways with taxonomically-patchy distributions can be indicative of HGT events, the eukaryotic nitrate assimilation pathway is an ideal object of investigation, as previous results revealed a patchy distribution and suggested one crucial HGT event. We studied the evolution of this pathway through both multi-scale bioinformatic and experimental approaches. Our taxon-rich genomic screening shows this pathway to be present in more lineages than previously proposed and that nitrate assimilation is restricted to autotrophs and to distinct osmotrophic groups. Our phylogenies show a pervasive role of HGT, with three bacterial transfers contributing to the pathway origin, and at least seven well-supported transfers between eukaryotes. Our results, based on a larger dataset, differ from the previously proposed transfer of a nitrate assimilation cluster from Oomycota (Stramenopiles) to Dikarya (Fungi, Opisthokonta). We propose a complex HGT path involving at least two cluster transfers between Stramenopiles and Opisthokonta. We also found that gene fusion played an essential role in this evolutionary history, underlying the origin of the canonical eukaryotic nitrate reductase, and of a novel nitrate reductase in Ichthyosporea (Opisthokonta). We show that the ichthyosporean pathway, including this novel nitrate reductase, is physiologically active and transcriptionally co-regulated, responding to different nitrogen sources; similarly to distant eukaryotes with independent HGT-acquisitions of the pathway. This indicates that this pattern of transcriptional control evolved convergently in eukaryotes, favoring the proper integration of the pathway in the metabolic landscape. Our results highlight the importance of reticulate evolution in eukaryotes, by showing the crucial contribution of HGT and gene fusion in the evolutionary history of the nitrate assimilation pathway.

evolutionary biology