bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

A predicted deleterious allele of the essential meiosis gene MND1, present in ~3% of East Asians, does not disrupt reproduction in mice.

Infertility is a major health problem affecting ~15% of couples worldwide. Except for cases involving readily-detectable chromosome aberrations, confident identification of a causative genetic defect is problematic. Despite the advent of genome sequencing for diagnostic purposes, the preponderance of segregating genetic variants complicates identification of culprit genetic alleles or mutations. Many algorithms have been developed to predict the effects of \"variants of unknown significance\" (VUS), typically SNPs (single nucleotide polymorphisms), but these predictions are not sufficiently accurate for clinical action. As part of a project to identify population variants that impact fertility, we have been generating CRISPR-Cas9 edited mouse models of suspect SNPs in genes that are known to be required for fertility in mice. Here, we present data on a non-synonymous (amino acid altering) SNP (rs140107488) in the meiosis gene Mnd1, which is predicted bioinformatically to be deleterious to protein function. We report that when modeled in mice, this allele (MND1K85M), which is present allele frequency of ~3% in East Asians, has no discernable effect upon fertility, fecundity, or gametogenesis, although it may cause sex skewing of progeny from homozygous males. In sum, this appears to be a benign allele that can be eliminated or de-prioritized in clinical genomic analyses of infertility patients.

genetics

Characterization of consensus operator site for Streptococcus pneumoniae copper repressor, CopY

Copper is broadly toxic to bacteria. As such, bacteria have evolved specialized copper export systems (cop operons) often consisting of a DNA-binding/copper-responsive regulator (which can be a repressor or activator), a copper chaperone, and a copper exporter. For those bacteria using DNA-binding copper repressors, few studies have examined the regulation of this operon regarding the operator DNA sequence needed for repression. In Streptococcus pneumoniae (the pneumococcus), CopY is the copper repressor for the cop operon. Previously, these homologs have been characterized to bind a 10-base consensus sequence T/GACAnnTGTA. Here, we bioinformatically and empirically characterize these operator sites across species using S. pneumoniae CopY as a guide for binding. By examining the 21-base repeat operators for the pneumococcal cop operon and comparing binding of recombinant CopY to this, and the operator sites found in Enterococcus hirae, we show using biolayer interferometry that the T/GACAnnTGTA sequence is essential to binding, but it is not sufficient. We determine a more comprehensive S. pneumoniae CopY operator sequence to be RnYKACAAATGTARnY (where \"R\" is purine, \"Y\" is pyrimidine, and \"K\" is either G or T) binding with an affinity of 28 nM. We further propose that the cop operon operator consensus site of pneumococcal homologs be RnYKACAnnYGTARnY. This study illustrates the necessity to explore bacterial operator sites further to better understand bacterial gene regulation.

microbiology

Stress-responsive Entamoeba topoisomerase II: a potential anti-amoebic target

Topoisomerases are ubiquitous enzymes, involved in all DNA processes across the biological world. These enzymes are also targets for various anticancer and antimicrobial agents. The causative organism of amoebiasis, Entamoeba histolytica (Eh), has seven unexplored genes annotated as putative topoisomerases. One of the seven topoisomerases in this parasite was found to be highly up-regulated during heat shock and oxidative stress. The bioinformatic analysis shows that it is a eukaryotic type IIA topoisomerase. Its ortholog was also highly up-regulated during the late hours of encystation in E. invadens (Ei), the encystation model of Eh. Immunoprecipitated endogenous EhTopoII showed topoisomerase II activity in vitro. Immunolocalization studies show that this enzyme colocalized with newly forming nuclei during encystation, which is a significant event in maturing cysts. Double-stranded RNA mediated down-regulation of the TopoII both in Eh and Ei reduced the viability of actively growing trophozoites and also reduced the encystation efficiency in Ei. Drugs, targeting eukaryotic topoisomerase II, e.g., etoposide, ICRF193, and amsacrine, show 3-5 times higher EC50 in Eh than that of mammalian cells. Sequence comparison with human TopoII showed that key amino acid residues involved in the interactions with etoposide and ICRF193 are different in Entamoeba TopoII. Interestingly, ciprofloxacin an inhibitor of prokaryotic DNA gyrase showed about six times less EC50 value in Eh than that of human cells. The parasites notable susceptibility to prokaryotic topoisomerase drugs in comparison to human cells opens up the scope to study this invaluable enzyme in the light of an antiamoebic target.

biochemistry

Show me your secret(ed) weapons: a multifaceted approach reveals novel type III-secreted effectors of a plant pathogenic bacterium

Many Gram-negative plant and animal pathogenic bacteria employ a type III secretion system (T3SS) to secrete protein effectors into the cells of their hosts and promote disease. The plant pathogen Acidovorax citrulli requires a functional T3SS for pathogenicity. As with Xanthomonas and Ralstonia spp., an AraC-type transcriptional regulator, HrpX, regulates expression of genes encoding T3SS components and type III-secreted effectors (T3Es) in A. citrulli. A previous study reported eleven T3E genes in this pathogen, based on the annotation of a sequenced strain. We hypothesized that this was an underestimation. Guided by this hypothesis, we aimed at uncovering the T3E arsenal of the A. citrulli model strain, M6. We carried out a thorough sequence analysis searching for similarity to known T3Es from other bacteria. This analysis revealed 51 A. citrulli genes whose products are similar to known T3Es. Further, we combined machine learning and transcriptomics to identify novel T3Es. The machine learning approach ranked all A. citrulli M6 genes according to their propensity to encode T3Es. RNA-Seq revealed differential gene expression between wild-type M6 and a mutant defective in HrpX. Data combined from these approaches led to the identification of seven novel T3E candidates, that were further validated using a T3SS-dependent translocation assay. These T3E genes encode hypothetical proteins, do not show any similarity to known effectors from other bacteria, and seem to be restricted to plant pathogenic Acidovorax species. Transient expression in Nicotiana benthamiana revealed that two of these T3Es localize to the cell nucleus and one interacts with the endoplasmic reticulum. This study not only uncovered the arsenal of T3Es of an important pathogen, but it also places A. citrulli among the \"richest\" bacterial pathogens in terms of T3E cargo. It also revealed novel T3Es that appear to be involved in the pathoadaptive evolution of plant pathogenic Acidovorax species.\n\nAuthor summaryAcidovorax citrulli is a Gram-negative bacterium that causes bacterial fruit blotch (BFB) disease of cucurbits. This disease represents a serious threat to cucurbit crop production worldwide. Despite the agricultural importance of BFB, the knowledge about basic aspects of A. citrulli-plant interactions is rather limited. As many Gram-negative plant and animal pathogenic bacteria, A. citrulli employs a complex secretion system, named type III secretion system, to deliver protein virulence effectors into the host cells. In this work we aimed at uncovering the arsenal of type III-secreted effectors (T3Es) of this pathogen by combination of bioinformatics and experimental approaches. We found that this bacterium possesses at least 51 genes that are similar to T3E genes from other pathogenic bacteria. In addition, our study revealed seven novel T3Es that seem to occur only in A. citrulli strains and in other plant pathogenic Acidovorax species. We found that two of these T3Es localize to the plant cell nucleus while one partially interacts with the endoplasmic reticulum. Further characterization of the novel T3Es identified in this study may uncover new host targets of pathogen effectors and new mechanisms by which pathogenic bacteria manipulate their hosts.

microbiology

Effector prediction and characterization in the oomycete pathogen Bremia lactucae reveal host-recognized WY domain proteins that lack the canonical RXLR motif

Pathogens infecting plants and animals use a diverse arsenal of effector proteins to suppress the host immune system and promote infection. Identification of effectors in pathogen genomes is foundational to understanding mechanisms of pathogenesis, for monitoring field pathogen populations, and for breeding disease resistance. We identified candidate effectors from the lettuce downy mildew pathogen, Bremia lactucae, using comparative genomics and bioinformatics to search for the WY domain. This conserved structural element is found in Phytophthora effectors and some other oomycete pathogens; it has been implicated in the immune-suppressing function of these effectors as well as their recognition by host resistance proteins. We predicted 54 WY domain containing proteins in isolate SF5 of B. lactucae that have substantial variation in both sequence and domain architecture. These candidate effectors exhibit several characteristics of pathogen effectors, including an N-terminal signal peptide, lineage specificity, and expression during infection. Unexpectedly, only a minority of B. lactucae WY effectors contain the canonical N-terminal RXLR motif, which is a conserved feature in the majority of cytoplasmic effectors reported in Phytophthora spp. Functional analysis effectors containing WY domains revealed eleven out of 21 that triggered necrosis, which is characteristic of the immune response on wild accessions and domesticated lettuce lines containing resistance genes. Only two of the eleven recognized effectors contained a canonical RXLR motif, suggesting that there has been an evolutionary divergence in sequence motifs between genera; this has major consequences for robust effector prediction in oomycete pathogens.\n\nAuthor SummaryThere is a microscopic battle that takes place at the molecular level during infection of plants and animals by pathogens. Some of the weapons that pathogens battle with are known as \"effectors,\" which are secreted proteins that enter host cells to alter physiology and suppress the immune system. Effectors can also be a liability for plant pathogens because plants have evolved ways to recognize these effectors, triggering a defense response leading to localized cell death, which prevents the spread of the pathogen. Here we used computer models to predict effectors from the genome of Bremia lactucae, the causal agent of lettuce downy mildew. Three effectors were demonstrated to suppress the basal immune system of lettuce. Eleven effectors were recognized by one or more resistant lines of lettuce. In addition to contributing to our understanding of the mechanisms of pathogenesis, this study of effectors is useful for breeding disease resistant lettuce, decreasing agricultural reliance on fungicides.

plant biology

Estimation of Speciation Times Under the Multispecies Coalescent

MotivationThe multispecies coalescent model is now widely accepted as an effective model for incorporating variation in the evolutionary histories of individual genes into methods for phylogenetic inference from genome-scale data. However, because model-based analysis under the coalescent can be computationally expensive for large data sets, a variety of inferential frameworks and corresponding algorithms have been proposed for estimation of species-level phylogenies and associated parameters, including speciation times and effective population sizes. ResultsWe consider the problem of estimating the timing of speciation events along a phylogeny in a coalescent framework. We propose a maximum a posteriori estimator based on composite likelihood (MAPCL) for inferring these speciation times under a model of DNA sequence evolution for which exact site pattern probabilities can be computed under the assumption of a constant{theta} throughout the species tree. We demonstrate that the MAPCL estimates are statistically consistent and asymptotically normally distributed, and we show how this result can be used to estimate their asymptotic variance. We also provide a more computationally efficient estimator of the asymptotic variance based on the nonparametric bootstrap. We evaluate the performance of our method using simulation and by application to an empirical dataset for gibbons. Availability and implementationThe method has been implemented in the PAUP* program, freely available at https://paup.phylosolutions.com for Macintosh, Windows, and Linux operating systems. Contactpeng.650@osu.edu Supplementary informationSupplementary data are available at Bioinformatics online.

evolutionary biology

Identification and characterization of cis-regulatory elements for photoreceptor type-specific transcription in zebrafish

Tissue-specific or cell type-specific transcription of protein-coding genes is controlled by both trans-regulatory elements (TREs) and cis-regulatory elements (CREs). However, it is challenging to identify TREs and CREs, which are unknown for most genes. Here, we describe a protocol for identifying two types of transcription-activating CREs--core promoters and enhancers--of zebrafish photoreceptor type-specific genes. This protocol is composed of three phases: bioinformatic prediction, experimental validation, and characterization of the CREs. To better illustrate the principles and logic of this protocol, we exemplify it with the discovery of the core promoter and enhancer of the mpp5b apical polarity gene (also known as ponli), whose red, green, and blue (RGB) cone-specific transcription requires its enhancer, a member of the rainbow enhancer family. While exemplified with an RGB cone-specific gene, this protocol is general and can be used to identify the core promoters and enhancers of other protein-coding genes.

developmental biology

QTL analysis of macrophages from an AKR/JxDBA/2J intercross identified the Gpnmb gene as a modifier of lysosome function

Our prior studies found differences in the AKR/J and DBA/2J strains in regard to atherosclerosis and macrophage phenotypes including cholesterol ester loading, cholesterol efflux, and autolysosome formation. The goal of this study was to determine if there were differences in macrophage lysosome function, and if so to use quantitative trait locus (QTL) analysis to identify the causal gene. Lysosome function was measured by incubation with an exogenous double-labeled ovalbumin indicator sensitive to proteolysis. DBA/2J vs. AKR/J bone marrow macrophages had significantly decreased lysosome function. Macrophages were cultured from 120 mice derived from an AKR/JxDBA/2J F4 intercross. We measured lysosome function and performed a high density genome scan. QTL analysis yielded two genome wide significant loci on chromosomes 6 and 17, called macrophage lysosome function modifier (Mlfm) loci Mlfm1 and Mlfm2. After adjusting for Mlfm1, two additional loci were identified. Based on proximity to the Mlfm1 peak, macrophage mRNA expression differences with AKR/J >> DBA/2J, and a protein coding nonsense variant in DBA/2J, the Gpnmb gene, encoding a lysosomal membrane protein, was our top candidate. To test this candidate, Gpnmb expression was knocked down with siRNA in AKR/J macrophages; and, to express the wildtype Gpnmb in DBA/2J macrophages, we obtained a DBA/2 substrain, DBA/2J-Gpnmb+/SjJ, which was isolated from the parental strain prior to its acquiring the nonsense mutation, and subsequently back crossed to the modern DBA/2J background. Knockdown of Gpnmb in AKR/J macrophages decreased lysosome function, while restoration of the wildtype Gpnmb allele in DBA/2J macrophages increased lysosome function. However, this modifier of lysosome function was not responsible for the strain differences in macrophage cholesterol ester loading or cholesterol efflux. In conclusion, we identified the Gpnmb gene as the major modifier of lysosome function and we showed that the QTL in a dish strategy is efficient in identifying modifier genes.\n\nAuthor SummaryInbred strains of mice differ in both their genetic backgrounds as well as in many traits; and, classical mouse genetics allows the mapping of genes responsible for these traits. We identified many traits that differ between the inbred strains AKR/J and DBA/2J, including atherosclerosis susceptibility, macrophage cholesterol metabolism, and in the current study, macrophage protein degradation via an organelle called the lysosome. Using mouse genetic mapping and bioinformatics we identified a candidate gene, called Gpnmb, responsible for modifying lysosome function; and, the DBA/2J strain carries a mutation in this gene. Here we demonstrate that the Gpnmb gene is a modifier of lysosome function by either correcting this Gpnmb mutation in DBA/2J macrophages, or by knocking down Gpnmb expression in AKR/J macrophages. This study is noteworthy as the human GPNMB gene has been implicated in many diseases including cancer, kidney injury, obesity, non-alcoholic steatohepatitis, Parkinson disease, osteoarthritis, and lysosome storage disorders.

genetics

MitoFinder: efficient automated large-scale extraction of mitogenomic data in target enrichment phylogenomics

Thanks to the development of high-throughput sequencing technologies, target enrichment sequencing of nuclear ultraconserved DNA elements (UCEs) now allows routinely inferring phylogenetic relationships from thousands of genomic markers. Recently, it has been shown that mitochondrial DNA (mtDNA) is frequently sequenced alongside the targeted loci in such capture experiments. Despite its broad evolutionary interest, mtDNA is rarely assembled and used in conjunction with nuclear markers in capture-based studies. Here, we developed MitoFinder, a user-friendly bioinformatic pipeline, to efficiently assemble and annotate mitogenomic data from hundreds of UCE libraries. As a case study, we used ants (Formicidae) for which 501 UCE libraries have been sequenced whereas only 29 mitogenomes are available. We compared the efficiency of four different assemblers (IDBA-UD, MEGAHIT, MetaSPAdes, and Trinity) for assembling both UCE and mtDNA loci. Using MitoFinder, we show that metagenomic assemblers, in particular MetaSPAdes, are well suited to assemble both UCEs and mtDNA. Mitogenomic signal was successfully extracted from all 501 UCE libraries allowing confirming species identification using COI barcoding. Moreover, our automated procedure retrieved 296 cases in which the mitochondrial genome was assembled in a single contig, thus increasing the number of available ant mitogenomes by an order of magnitude. By leveraging the power of metagenomic assemblers, MitoFinder provides an efficient tool to extract complementary mitogenomic data from UCE libraries, allowing testing for potential mito-nuclear discordance. Our approach is potentially applicable to other sequence capture methods, transcriptomic data, and whole genome shotgun sequencing in diverse taxa.

evolutionary biology

Developing critical thinking in STEM education through inquiry-based writing in the laboratory classroom

Laboratory pedagogy is moving away from step-by-step instructions and toward inquiry-based learning (IBL), but only now developing methods for integrating inquiry-based writing (IBW) practices into the laboratory course. Based on an earlier proposal (Science 332:919 (2011)), we designed and implemented an IBW sequence in a university bioinformatics course.\n\nWe automatically generated unique, double-blinded, biologically plausible DNA sequences for each student. After guided instruction, students investigated sequences independently and responded through IBW writing assignments. IBW assignments were structured as condensed versions of a scientific research paper, and because the sequences were double blinded, they were also assessed as authentic science and evaluated on clarity and persuasiveness.\n\nWe piloted the approach in a seven-day workshop (35 students) at Perdana University Graduate School of Medicine in Kuala Lumpur. We observed dramatically improved student engagement and indirect evidence of improved learning outcomes over a similar workshop without IBW. Based on student feedback, initial discomfort with the writing component abated in favor of an overall positive response and increasing comfort with the high demands of student writing. Similarly encouraging results were found in a semester length undergraduate module at the National University of Singapore (155 students).

scientific communication and education

Brain Regional Gene Expression Network Analysis Identifies Unique Interactions Between Chronic Ethanol Exposure and Consumption

Progressive increases in ethanol consumption is a hallmark of alcohol use disorder (AUD). Persistent changes in brain gene expression are hypothesized to underlie the altered neural signaling producing abusive consumption in AUD. To identify brain regional gene expression networks contributing to progressive ethanol consumption, we performed microarray and scale-free network analysis of expression responses in a C57BL/6J mouse model utilizing chronic intermittent ethanol by vapor chamber (CIE) in combination with limited access oral ethanol consumption. This model has previously been shown to produce long-lasting increased ethanol consumption, particularly when combining oral ethanol access with repeated cycles of intermittent vapor exposure. The interaction of CIE and oral consumption was studied by expression profiling and network analysis in medial prefrontal cortex, nucleus accumbens, hippocampus, bed nucleus of the stria terminalis, and central nucleus of the amygdala. Brain region expression networks were analyzed for ethanol-responsive gene expression, correlation with ethanol consumption and functional content using extensive bioinformatics studies. In all brain-regions studied the largest number of changes in gene expression were seen when comparing ethanol naive mice to those exposed to CIE and drinking. In the prefrontal cortex, however, unique patterns of gene expression were seen compared to other brain-regions. Network analysis identified modules of co-expressed genes in all brain regions. The prefrontal cortex and nucleus accumbens showed the greatest number of modules with significant correlation to drinking behavior. Across brain-regions, however, many modules with strong correlations to drinking, both baseline intake and amount consumed after CIE, showed functional enrichment for synaptic transmission and synaptic plasticity.

genomics

The Drosophila fertility factor kl-3 is linked to the Y-chromosome of the vector of Chagas’ disease Triatoma infestans (Hemiptera: Reduviidae) and is essential for male fertility

In many insects, the Y chromosome plays a key role in sexual determination and male fertility. The Chagas disease vector Triatoma infestans has 22 autosomal chromosomes and a pair of XY sex chromosomes. However, the knowledge on the Y chromosome of this species, its genetic content or its biological function, is very poor. Due to repetitive DNA, Y chromosome sequences are poorly assembled in genome projects, hindering structural and functional studies on Y-linked genes. Our group has developed many of the bioinformatic tools to identify Y-linked sequences in assembled genomes. Here, we describe the identification of a {gamma}-dynein heavy chain linked to the Y-chromosome of T. infestans. This protein is orthologous to the Drosophila melanogaster Y-linked gene kl-3. In D. melanogaster, dyneins of the Y chromosome are known as male fertility factors and their deletion causes male infertility. We performed knockdown of the kl-3 expression to ascertain its function in T. infestans. Our results showed that injection of dsKL3 reduced, significantly, the fertility of T. infestans males (p<0.01). The mean number of eggs laid by the control group was 35.64 eggs/couple while the kl-3 knockdown group was of 11.82 eggs/couple (five couples did not lay any eggs). Differences in eclosion rate was even more significant, with a hatching mean rate of 16.85{+/-}10.03 and 1.69{+/-}3.58 (p<0.001) for the control and the silenced groups respectively. Our results suggest that kl-3 maintains its functional role as essential for male fertility in T. infestans. Hence, it seems that the Y-chromosome of T. infestans has a key role in male fertility. This is the first report of a kl-3 orthologue linked to the Y chromosome of an insect species outside the diptera clade. In addition to the first report of a Y-linked gene in T. infestans with a role for male fertility, this finding is of great relevance for the study of the evolution of Y chromosomes and further studies that could lead to novel approaches in insect control.

genomics

Druggable genome screen identifies new regulators of the abundance and toxicity of ATXN3, the Spinocerebellar Ataxia Type 3 disease protein

BackgroundSpinocerebellar Ataxia type 3 (SCA3, also known as Machado-Joseph disease) is a neurodegenerative disorder caused by a CAG repeat expansion encoding an abnormally long polyglutamine (polyQ) tract in the disease protein, ataxin-3 (ATXN3). No preventive treatment is yet available for SCA3. Because SCA3 is likely caused by a toxic gain of ATXN3 function, a rational therapeutic strategy is to reduce mutant ATXN3 levels by targeting pathways that control its production or stability. Here, we sought to identify genes that modulate ATXN3 levels as potential therapeutic targets in this fatal disorder.\n\nMethodsWe screened a collection of siRNAs targeting 2742 druggable human genes using a cell-based assay based on luminescence readout of polyQ-expanded ATXN3. From 317 candidate genes identified in the primary screen, 100 genes were selected for validation. Among the 33 genes confirmed in secondary assays, 15 were validated in an independent cell model as modulators of pathogenic ATXN3 protein levels. Ten of these genes were then assessed in a Drosophila model of SCA3, and one was confirmed as a key modulator of physiological ATXN3 abundance in SCA3 neuronal progenitor cells.\n\nResultsAmong the 15 genes shown to modulate ATXN3 in mammalian cells, orthologs of CHD4, FBXL3, HR and MC3R regulate mutant ATXN3-mediated toxicity in fly eyes. Further mechanistic studies of one of these genes, FBXL3, encoding a F-box protein that is a component of the SKP1-Cullin-F-box (SCF) ubiquitin ligase complex, showed that it reduces levels of normal and pathogenic ATXN3 in SCA3 neuronal progenitor cells, primarily via a SCF complex-dependent manner. Bioinformatic analysis of the 15 genes revealed a potential molecular network with connections to tumor necrosis factor-/nuclear factor-kappa B (TNF/NF-kB) and extracellular signal-regulated kinases 1 and 2 (ERK1/2) pathways.\n\nConclusionsWe identified 15 druggable genes with diverse functions to be suppressors or enhancers of pathogenic ATXN3 abundance. Among identified pathways highlighted by this screen, the FBXL3/SCF axis represents a novel molecular pathway that regulates physiological levels of ATXN3 protein.

neuroscience

miR-1/206 down-regulates splicing factor Srsf9 to promote myogenesis

BackgroundMyogenesis is driven by specific changes in the transcriptome that occur during the different stages of muscle differentiation. In addition to controlled transcriptional transitions, several other post-transcriptional mechanisms direct muscle differentiation. Both alternative splicing and miRNA activity regulate gene expression and production of specialized protein isoforms. Importantly, disruption of either process often results in severe phenotypes as reported for several muscle diseases. Thus, broadening our understanding of the post-transcriptional pathways that operate in muscles will lay the foundation for future therapeutic interventions.\n\nMethodsWe employed bioinformatics analysis in concert with the well-established C2C12 cell system for predicting and validating novel miR-1 and miR-206 targets engaged in muscle differentiation. We used reporter gene assays to test direct miRNA targeting and studied C2C12 cells stably expressing one of the cDNA candidates fused to a heterologous, miRNA-resistant 3 UTR. We monitored effects on differentiation by measuring fusion index, myotube area, and myogenic gene expression during time course differentiation experiments.\n\nResultsGene ontology analysis revealed a strongly enriched set of putative miR-1 and miR-206 targets associated with RNA metabolism. Notably, the expression levels of several candidates decreased during C2C12 differentiation. We discovered that the splicing factor Srsf9 is a direct target of both miRNAs during myogenesis. Persistent Srsf9 expression during differentiation impaired myotube formation and blunted induction of the early pro-differentiation factor myogenin as well as the late differentiation marker sarcomeric myosin, Myh8.\n\nConclusionsOur data uncover novel miR-1 and miR-206 cellular targets and establish a functional link between the splicing factor Srsf9 and myoblast differentiation. The finding that miRNA-mediated clearance of Srsf9 is a key myogenic event illustrates the coordinated and sophisticated interplay between the diverse components of the gene regulatory network.

cell biology

Insights into the bacterial profiles and resistome structures following severe 2018 flood in Kerala, South India

Extreme flooding is one of the major risk factors for human health, and it can significantly influence the microbial communities and enhance the mobility of infectious disease agents within its affected areas. The flood crisis in 2018 was one of the severe natural calamities recorded in the southern state of India (Kerala) that significantly affected its economy and ecological habitat. We utilized a combination of shotgun metagenomics and bioinformatics approaches for understanding microbiome disruption and the dissemination of pathogenic and antibiotic-resistant bacteria on flooded sites. Here we report, altered bacterial profiles at the flooded sites having 77 significantly different bacterial genera in comparison with non-flooded mangrove settings. The flooded regions were heavily contaminated with faecal contamination indicators such as Escherichia coli and Enterococcus faecalis and resistant strains of Pseudomonas aeruginosa, Salmonella Typhi/Typhimurium, Klebsiella pneumoniae, Vibrio cholerae and Staphylococcus aureus. The resistome of the flooded sites contains 103 resistant genes, of which 38% are encoded in plasmids, where most of them are associated with pathogens. The presence of 6 pathogenic bacteria and its susceptibility to multiple antibiotics including ampicillin, chloramphenicol, kanamycin and tetracycline hydrochloride were confirmed in flooded and post-flooded sites using traditional culture-based analysis followed by 16S rRNA sequencing. Our results reveal altered bacterial profile following a devastating flood event with elevated levels of both faecal contamination indicators and resistant strains of pathogenic bacteria. The circulation of raw sewage from waste treatment settings and urban area might facilitate the spreading of pathogenic bacteria and resistant genes.

microbiology

A restriction enzyme reduced representation sequencing approach for low-cost, high-throughput metagenome profiling

Microbial community profiles have been associated with a variety of traits, including methane emissions in livestock, however, these profiles can be difficult and expensive to obtain for thousands of samples. The objective of this work was to develop a low-cost, high-throughput approach to capture the diversity of the rumen microbiome. Restriction enzyme reduced representation sequencing (RE-RRS) using ApeKI or PstI, and two bioinformatic pipelines (reference-based and reference-free) were compared to 16S rRNA gene sequencing using repeated samples collected two weeks apart from 118 sheep that were phenotypically extreme (60 high and 58 low) for methane emitted per kg dry matter intake (n=236). DNA was extracted from freeze-dried rumen samples using a phenol chloroform and bead-beating protocol prior to sequencing. The resulting sequences were used to investigate the repeatability of the rumen microbial community profiles, the effect of host genetics, laboratory and analytical method, and the genetic and phenotypic correlations with methane production. The results suggested that the best method was PstI RE-RRS analyzed with the reference-free approach via a correspondence analysis, with estimates for repeatability of 0.62{+/-}0.06, heritability 0.31{+/-}0.29, and genetic and phenotypic correlation with methane emissions of 0.88{+/-}0.25 and 0.64{+/-}0.05 respectively for the first component of correspondence analysis. The reference-free approach assigned 62.0{+/-}5.7% of reads to common 65 bp tags, much higher than the reference-based approach of 6.8{+/-}1.8% of reads assigned. Sensitivity studies suggested approximately 2000 samples could be sequenced in a single lane on an Illumina HiSeq 2500, therefore the current work of 118 samples/lane and future proposed 384 samples/lane are well within that threshold. Our approach is now being used to investigate host factors affecting the rumen and its association with a variety of production and environmental traits. With minor adaptations, our approach could be used to obtain microbial profiles from other metagenomic samples.

microbiology

Conservation of gene architecture and domains amidst sequence divergence in the hsrω lncRNA gene across the Drosophila genus

The developmentally active and cell-stress responsive hsr{omega} locus in Drosophila melanogaster carries two exons, one omega intron, one short translatable open reading frame ORF{omega}, long stretch of unique tandem repeats and an overlapping mir-4951 near its 3 end. It produces multiple lncRNAs using two transcription start and four termination sites. Earlier studies revealed functional conservation in several Drosophila species but with little sequence conservation, in three experimentally examined species, of ORF{omega}, tandem repeat and other regions but ultra-conservation of 16nt at 5 and 60nt at 3 splice-junctions of the omega intron. Present bioinformatic study, using the splice-junction landmarks in Drosophila melanogaster hsr{omega}, identified orthologues in publicly available 34 Drosophila species genomes. Each orthologue carries the short ORF{omega}, ultra-conserved splice junctions of omega intron, repeat region, conserved 3-end located mir-4951, and syntenic neighbours. Multiple copies of conserved nonamer motifs are seen in the tandem repeat region, despite a high variability in repeat sequences. Intriguingly, only the intron sequences in different species show evolutionary relationships matching the general phylogenetic history in the genus. Search in other known insect genomes did not reveal sequence homology although a locus with similar functional properties is suggested in Chironomus and Ceratitis species. Amidst the high sequence divergence, the conserved organization of exons, ORF{omega} and omega intron in this genes proximal part and tandem repeats in distal part across the Drosophila genus is remarkable and possibly reflects functional importance of higher order structure of hsr{omega} lncRNAs and the small Omega peptide.

genomics

A comprehensive dataset of TLX1 positive ALL-SIL lymphoblasts and primary T-cell acute lymphoblastic leukemias

Most currently available transcriptome data of T-cell acute lymphoblastic leukemia (T-ALL) are based on polyA[+] RNA sequencing methods thus lacking non-polyadenylated transcripts. Here, we present the data of polyA[+] and total RNA sequencing in the context of in vitro TLX1 knockdown in ALL-SIL cells and a primary T-ALL cohort. We extended this dataset with ATAC sequencing and H3K4me1 and H3K4me3 ChIP sequencing data to map putative gene regulatory regions. In this data descriptor, we present a detailed report of how the data were generated and which bioinformatics analyses were performed. Through several technical validations, we showed that our sequencing data are of high quality and that our in vitro TLX1 knockdown was successful. We also validated the quality of the ATAC and ChIP sequencing data and showed that ATAC and H3K4me3 ChIP peaks are enriched at transcription start sites. We believe that this comprehensive set of sequencing data can be reused by others to further unravel the complex biology of T-ALL in general and TLX1 in particular.

molecular biology