bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

A Network of Networks Approach for Modeling Interconnected Brain Tissue-Specific Networks

MotivationRecent sequence-based analyses have identified a lot of gene variants that may contribute to neurogenetic disorders such as autism spectrum disorder and schizophrenia. Several state-of-the-art network-based analyses have been proposed for mechanical understanding of genetic variants in neurogenetic disorders. However, these methods were mainly designed for modeling and analyzing single networks that do not interact with or depend on other networks, and thus cannot capture the properties between interdependent systems in brain-specific tissues, circuits, and regions which are connected each other and affect behavior and cognitive processes.\n\nResultsWe introduce a novel and efficient framework, called a \"Network of Networks\" (NoN) approach, to infer the interconnectivity structure between multiple networks where the response and the predictor variables are topological information matrices of given networks. We also propose Graph-Oriented SParsE Learning (GOSPEL), a new sparse structural learning algorithm for network graph data to identify a subset of the topological information matrices of the predictors related to the response. We demonstrate on simulated data that GOSPEL outperforms existing kernel-based algorithms in terms of F-measure. On real data from human brain region-specific functional networks associated with the autism risk genes, we show that the NoN model provides insights on the autism-associated interconnectivity structure between functional interaction networks and a comprehensive understanding of the genetic basis of autism across diverse regions of the brain.\n\nAvailabilityOur software is available from https://github.com/infinite-point/GOSPEL.\n\nContactkawakubo@med.nagoya-u.ac.jp, shimamura@med.nagoya-u.ac.jp\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Potent Cas9 inhibition in bacterial and human cells by new anti-CRISPR protein families

CRISPR-Cas systems are widely used for genome engineering technologies, and in their natural setting, they play crucial roles in bacterial and archaeal adaptive immunity, protecting against phages and other mobile genetic elements. Previously we discovered bacteriophage-encoded Cas9-specific anti-CRISPR (Acr) proteins that serve as countermeasures against host bacterial immunity by inactivating their CRISPR-Cas systems1. We hypothesized that the evolutionary advantages conferred by anti-CRISPRs would drive the widespread occurrence of these proteins in nature2-4. We have identified new anti-CRISPRs using the bioinformatic approach that successfully identified previous Acr proteins1 against Neisseria meningitidis Cas9 (NmeCas9). In this work we report two novel anti-CRISPR families in strains of Haemophilus parainfluenzae and Simonsiella muelleri, both of which harbor type II-C CRISPR-Cas systems5. We characterize the type II-C Cas9 orthologs from H. parainfluenzae and S. muelleri, show that the newly identified Acrs are able to inhibit these systems, and define important features of their inhibitory mechanisms. The S. muelleri Acr is the most potent NmeCas9 inhibitor identified to date. Although inhibition of NmeCas9 by anti-CRISPRs from H. parainfluenzae and S. muelleri reveals cross-species inhibitory activity, more distantly related type II-C Cas9s are not inhibited by these proteins. The specificities of anti-CRISPRs and divergent Cas9s appear to reflect co-evolution of their strategies to combat or evade each other. Finally, we validate these new anti-CRISPR proteins as potent off-switches for Cas9 genome engineering applications.

molecular biology

Single-cell level transcriptome of the maize pathogenic fungi cochliobolus heterostrophus race O in infection reveal the virulence related genes, and potential circRNA effector

Cochliobolus heterostrophus is a crucial pathogenic fungus that causes southern corn leaf blight (SCLB) in maize worldwide, however, the virulence mechanism of the dominant race O remains unclear. In this report, the single-cell level of pathogen tissue at three infection stages were collected from the host interaction-situ, and were performed next-generation sequencing from the perspectives of mRNA, circular RNA(circRNA) and long noncoding RNA(lncRNA). In the mRNA section, signal transduction, kinase, oxidoreductase, and hydrolase, et al. were significantly related in both differential expression and co-expression between virulence differential race O strains. The expression pattern of the traditional virulence factors nonribosomal peptide synthetases (NPSs), polyketide synthases (PKSs) and small secreted proteins (SSPs) were multifarious. In the noncoding RNA section, a total of 2279 circRNAs and 169 lncRNAs were acquired. Noncoding RNAs exhibited differential expression at three stages. The high virulence strain DY transcribed 450 more circRNAs than low virulence strain WF. Informatics analysis revealed numbers of circRNAs which positively correlate with race O virulence, and a cross-kingdom interaction between the pathogenic circRNA and host miRNA was predicted. An important exon-intron circRNA Che-cirC2410 combines informatics characteristics above, and highly expressed in the DY strain. Che-cirC2410 initiate from the pseudogene chhtt, which doesnt translate genetic code into protein. In-situ hybridization tells the sub-cellular localization of Che-cirC2410 include pathogen`s mycelium, periplasm, and the diseased host tissues. The target of Che-cirC2410 was predicted to be zma-miR399e-5P, and the interaction between noncoding RNAs was proved. More, the expression of zma-miR399e-5P exhibited a negative correlation to Che-cirC2410 in vivo. The deficiency of Che-circ2410 decreased the race O virulence. The host resistance to SCLB was weakened when zma-miR399e-5P was silenced. Thus, a novel circRNA-type effector and its resistance related miRNA target are proposed cautiously in this report. These findings enriched the pathogen-host dialogue by using noncoding RNAs as language, and revealed a new perspective for understanding the virulence of race O, which may provide valuable strategy of maize breeding for disease resistance.\n\nAuthor SummaryThe southern corn leaf blight (caused by Cochliobolus heterostrophus) is not optimistic in Asia, however we have limit knowledge about the infection mechanism of the dominant C.heterostrophus race O. We take full advantage of the ideal C.heterostrophus genome database, laser capture microdissection and single-cell level RNA sequencing. Hence, we could avert the artificial influence such as medium, and profile the real gene mobilization strategy in the infection. The results of coding RNA section were accessible, virulence related genes (such as the signal transduction, PKS, SSP) were detected in RNA-seq,which accord with previous reports. However, the results of noncoding RNA was astonished, 2279 circular RNAs (circRNA) and 169 long noncoding RNAs (lncRNA) were revealed in our results. Generally, the function of noncoding RNA was hypothesized in single species, but we boldly guess that the function of circRNA is rather complicated in the pathogen-host interaction. Finally, the circRNA in-situ hybridization (ISH) demonstrate the secretion of pathogen circRNA into the host tissue. By bioinformatic prediction, we found a sole microRNA target, and proved the interaction between circRNA and microRNA. These findings are likely to reveal a novel pathogen effector type: secreted circRNA.

microbiology

Sequence analysis and confirmation of type IV pili-associated proteins PilY1, PilW and PilV in Acidithiobacillus thiooxidans

Acidithiobacillus thiooxidans is an acidophilic chemolithoautotrophic bacterium widely used in the mining industry due to its metabolic sulfur-oxidizing capability. The biooxidation of sulfide minerals is enhanced through the attachment of A. thiooxidans cells to the mineral surface. The Type IV pili (TfP) of At. thiooxidans may play an important role in the bacteria attachment, since among other functions, TfP play a key adhesive role in the attachment to and colonization of different surfaces. In this work, we reported for the first time the confirmed mRNA sequences of three TfP proteins from At. thiooxidans, the protein PilY1 and the TfP pilins PilW and PilV. The nucleotide sequences of these TfP proteins show changes of some nucleotide positions with respect to the corresponding annotated sequences. The bioinformatic analyses and 3D-modeling of protein structures sustain their classification as TfP proteins, as structural homologs of the corresponding proteins of P. aeruginosa, results that sustain the role of PilY1, PilW and PilV in pili assembly. Also, that PilY1 comprises the conserved Neisseria-PilC (superfamily) domain of the tip-associated adhesin, while PilW of the superfamily of putative TfP assembly proteins and PilV belongs to the superfamily of TfP assembly protein. Also, the analyses suggested the presence of specific functional domains involved in adhesion, energy transduction and signaling functions. The phylogenetic analysis indicated that the PilY1 of Acidithiobacillus genus forms a cohesive group linked with iron- and/or sulfur-oxidizing microorganisms from acid mine drainage or mine tailings. This work enriches knowledge regarding colonization, adhesion and biooxidation of inorganic sulfurs by A. thiooxidans.

microbiology

An Equivariant Bayesian Convolutional Network predicts recombination hotspots and accurately resolves binding motifs

MotivationConvolutional neural networks (CNNs) have been trememdously successful in many contexts, particularly where training data is abundant and signal-to-noise ratios are large. However, when predicting noisily observed biological phenotypes from DNA sequence, each training instance is only weakly informative, and the amount of training data is often fundamentally limited, emphasizing the need for methods that make optimal use of training data and any structure inherent in the model.\n\nResultsHere we show how to combine equivariant networks, a general mathematical framework for handling exact symmetries in CNNs, with Bayesian dropout, a version of MC dropout suggested by a reinterpretation of dropout as a variational Bayesian approximation, to develop a model that exhibits exact reverse-complement symmetry and is more resistant to overtraining. We find that this model has increased power and generalizability, resulting in significantly better predictive accuracy compared to standard CNN implementations and state-of-art deep-learning-based motif finders. We use our network to predict recombination hotspots from sequence, and identify high-resolution binding motifs for the recombination-initiation protein PRDM9, which were recently validated by high-resolution assays. The network achieves a predictive accuracy comparable to that attainable by a direct assay of the H3K4me3 histone mark, a proxy for PRDM9 binding.\n\nAvailabilityhttps://github.com/luntergroup/EquivariantNetworks\n\nContactrichard.brown@well.ox.ac.uk, gerton.lunter@well.ox.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

genomics

Nitrogen regulator GlnR directly controls transcription of prpDBC operon involved in methylcitrate cycle in Mycobacterium smegmatis

Mycobacterium tuberculosis utilizes the fatty acids of the host as the carbon source. While the metabolism of odd chain fatty acids produces propionyl-CoA. Methylcitrate cycle is essential for Mycobacteria to utilize the propionyl-CoA to persist and grow on these fatty acids. In M. smegmatis, methylcitrate synthase, methylcitrate dehydratase, and methylisocitrate lyase involved in methylcitrate cycle were respectively encoded by prpC, prpD, and prpB in operon prpDBC. In this study, we found that the nitrogen regulator GlnR directly binds to the promoter region of prpDBC operon and inhibits its transcription. The typical binding sequence of GlnR was identified by bioinformatics analysis and electrophoretic mobility shift assay. The GlnR-binding motif was seperated by 164 bp with the binding site of PrpR which was a pathway-specific transcriptional activator of methylcitrate cycle. Moreover, the affinity constant of GlnR was much stronger than that of PrpR to prpDBC. The deletion of glnR resulted in poor growth in propionate or cholesterol medium comparing with wild-type strain. The {Delta}glnR mutant strain also showed a higher survival in macrophages. These results illustrated that the nitrogen regulator GlnR regulated methylcitrate cycle through directly repressing the transcription of prpDBC operon. The finding reveals an unprecedented link between nitrogen metabolism and methylcitrate pathway, and provides a potential application for controlling populations of pathogenic mycobacteria.\n\nAuthor SummaryNutrients are crucial for the survival and pathogenicity of Mycobacterium tuberculosis. The success of this pathogen survival in macrophage due to its ability to assimilate fatty acids and cholesterol from host. The cholesterol and fatty acids are catabolized via {beta}-oxidation to generate propionyl-CoA, which is then mainly metabolized via the methylcitrate cycle. The assimilation of propionyl-CoA needs to be tightly regulated to prevent its accumulation and alleviate toxicity in cell. Here, we identified a new regulator GlnR (the nitrogen transcriptional regulator) that repressed the transcription of prp operon involved in methylcitrate cycle in M. smegmatis. In this study, we found a typical GlnR binding box in prp operon, and the affinity is much stronger than that of PrpR which is known as a pathway-specific transcriptional activator of methylcitrate cycle. In addition, deletion of glnR obviously affect the growth of mutant in propionate or cholesterol medium, and show a better viability in macrophage. The findings not only provide the insights into the regulatory mechanism underlying crosstalk of nitrogen metabolism and carbon metabolism, but also reveal a potential application for controlling populations of pathogenic mycobacteria.

microbiology

Over 2.5 million COI sequences in GenBank and growing

The increasing popularity of cytochrome c oxidase subunit 1 (COI) DNA metabarcoding warrants a careful look at the underlying reference databases used to make high-throughput taxonomic assignments. The objectives of this study are to document trends and assess the future usability of COI records for metabarcode identification. Over 2.5 million COI sequences were found in GenBank, half of which were fully identified to the species rank. From 2003 to 2017, the number of COI Eukaryote records deposited has grown by two orders of magnitude representing a nearly 42-fold increase in unique species. For fully identified records, 92% are at least 500 bp in length, 74% have a country annotation, and 51% have latitude-longitude annotations. To ensure the future usability of COI records in GenBank we suggest: 1) Improving the geographic representation of COI records 2) Improving the cross-referencing of COI records in the Barcode of Life Data System and GenBank to facilitate consolidation and incorporation into existing bioinformatic pipelines, 3) Adherence to the minimum information about a marker gene sequence guidelines, and 4) Integrating metabarcodes from eDNA and mixed community studies with existing sequences. COI metabarcoders are normally considered consumers of taxonomic data. Here we discuss the potential for taxonomists to reverse this pattern and instead mine metabarcode data to guide species discovery. The growth of COI reference records over the past 15 years has been substantial and is likely to be a resource across many fields for years to come.

ecology

TP53 mutations promote immunogenic activity in breast cancer

BackgroundAlthough immunotherapy has recently achieved clinical successes in a variety of cancers, thus far there is no any immunotherapeutic strategy for breast cancer (BC). Thus, it is important to discover biomarkers for identifying the BC patients responsive to immunotherapy. TP53 mutations were often associated with worse clinical outcome in BC, of which the triple-negative BC (TNBC) has a high TP53 mutation rate (approximately 80%). TNBC is high-risk due to its high invasiveness, and lack of targeted therapy. To explore a potentially promising therapeutic option for the TP53-mutated BC subtype, we studied the associations between TP53 mutations and immunogenic activity in BC.\n\nMethodsWe compared enrichment levels of 26 immune gene-sets that indicated activities of diverse immune cells, functions, and pathways between TP53-mutated and TP53-wildtype BCs based on two large-scale BC multi-omics data. Moreover, we explored the molecular cues that were associated with the differences in immunogenic activity between TP53-mutated and TP53-wildtype BCs. Furthermore, we performed experimental validation of the findings from bioinformatics analysis.\n\nResultsWe found that almost all analyzed immune gene-sets had significantly higher enrichment levels in TP53-mutated BCs compared to TP53-wildtype BCs. Moreover, our experiments confirmed that mutant p53 could increase BC immunogenicity. Furthermore, our computational and experimental results showed that TP53 mutations could promote BC immunogenicity via regulation of the p53-mediated pathways including cell cycle, apoptosis, Wnt, Jak-STAT, NOD-like receptor, and glycolysis. Interestingly, we found that elevated immune activities were likely to be associated with better survival prognosis in TP53-mutated BCs, but not necessarily in TP53-wildtype BCs.\n\nConclusionsTP53 mutations promote immunogenic activity in breast cancer. This finding demonstrates a different effect of p53 dysfunction on tumor immunogenicity from that of previous studies, suggesting that the TP53 mutation status could be a useful biomarker for stratifying BC patients responsive to immunotherapy.

cancer biology

Genome-wide association analysis with a 50K transcribed gene SNP-chip identifies QTL affecting muscle yield in rainbow trout

Detection of coding/functional SNPs that change the biological function of a gene may lead to identification of putative causative alleles within QTL regions and discovery of genetic markers with large effects on phenotypes. Two bioinformatics pipelines, GATK and SAMtools, were used to identify ~21K transcribed SNPs with allelic imbalances associated with important aquaculture production traits including body weight, muscle yield, muscle fat content, shear force, and whiteness in addition to resistance/susceptibility to bacterial cold-water disease (BCWD). SNPs were identified from pooled RNA-Seq data collected from ~620 fish, representing 98 families from growth- and 54 families from BCWD-selected lines with divergent phenotypes. In addition, ~29K transcribed SNPs without allelic-imbalances were strategically added to build a 50K Affymetrix SNP-chip. SNPs selected included two SNPs per gene from 14K genes and ~5K non-synonymous SNPs. The SNP-chip was used to genotype 1728 fish. The average SNP calling-rate for samples passing quality control (QC; 1,641 fish) was [≥] 98.5%. Genome-wide association (GWA) study on 878 fish (representing 197 families from 2 consecutive generations) with muscle yield phenotypes and genotyped for 35K polymorphic markers (passing QC) identified several QTL regions explaining together up to 28.40% of the additive genetic variance for muscle yield in this rainbow trout population. The most significant QTLs were on chromosomes 14 and 16 with 12.71% and 10.49% of the genetic variance, respectively. Many of the annotated genes in the QTL regions were previously reported as important regulators of muscle development and cell signaling. No major QTLs were identified in a previous GWA study using a 57K genomic SNP chip on the same fish population. These results indicate improved detection power of the transcribed gene SNP-chip in the target trait and population, allowing identification of large-effect QTLs for important traits in rainbow trout.

genomics

G-quadruplex stabilization in the ions and maltose transporters inhibit Salmonella enterica growth and virulence.

The G-quadruplex structure forming motifs have recently emerged as a novel therapeutic drug target in various human pathogens. Herein, we report three highly conserved G-quadruplex motifs (SE-PGQ-1, 2, and3) in genome of all the 412 strains of Salmonella enterica. Bioinformatics analysis inferred the presence of SE-PGQ-1 in the regulatory region of mgtA, presence of SE-PGQ-2 in the open reading frame of entA and presence of SE-PGQ-3 in the promoter region of malE and malK genes. The products of mgtA and entA are involved in transport and homeostasis of Mg2+ and Fe3+ ion and thereby required for bacterial survival in the presence of reactive nitrogen/oxygen species produced by the host macrophages, whereas, malK and malE genes are involved in transport of maltose sugar, that is one of the major carbon source in the gastrointestinal tract of human. The formation of stable intramolecular G-quadruplex structures by SE-PGQs was confirmed by employing CD, EMSA and NMR spectroscopy. Cellular studies revealed the inhibitory effect of 9-amino acridine on Salmonella enterica growth. Next, CD melting analysis demonstrated the stabilizing effect of 9-amino acridine on SE-PGQs. Further, polymerase inhibition and RT-qPCR assays emphasize the biological relevance of predicted G-quadruplex in the expression of PGQ possessing genes and demonstrate the G-quadruplexes as a potential drug target for the devolping novel therapeutics for combating Salmonella enterica infection.\n\nAuthor SummarySince last several decades scientific community has witnessed a rapid increase in number of such human pathogenic bacterial species that acquired resistant to multiple antibacterial agents. Currently, emergence of multidrug-resistant strains remain a major public health concern for clinical investigators that rings a global alarm to search for novel and highly conserved drug targets. Recently, G-quadruplex structure forming nucleic acid sequences were endorsed as highly conserved Drug target for preventing infection of several human pathogens including viral and protozoan species. Therefore, here we explored the presence G-quadruplex forming motif in genome of Salmonella enterica bacteria that causes food poisoning, and enteric fever in human. The formation of intra molecular G-quadruplex structure in four genes (mgtA, entA, malE and malK) was confirmed by NMR, CD and EMSA. The 9-amino acridine, a known G-quadruplex binder has been shown to stabilize the predicted G-quadruplex motif and decreases the expressioin of G-quadruplex hourbouring genes using RT-PCR and cellular toxicity assay. This study concludes the presence of G-quadruplex motifs in essential genes of Salmonella enterica genome as a novel and conserved drug target and 9-amino acridine as candidate small molecule for preventing the infection of Salmonella enterica using a G4 mediated inhibition mechanism.

genomics

Identification of circulating protein biomarkers for pancreatic cancer cachexia

BackgroundOver 80% of patients with pancreatic ductal adenocarcinoma (PDAC) suffer from cachexia, characterized by severe muscle and fat loss. Although various model systems have improved our understanding of cachexia, translating the findings to human cachexia has remained a challenge. In this study, our objectives were to i) identify circulating protein biomarkers using serum for human PDAC cachexia, (ii) identify the ontological functions of the identified biomarkers and (iii) identify new pathways associated with human PDAC cachexia by performing protein co-expression analysis.\n\nMethodsSerum from 30 patients with PDAC was collected. Body composition measurements of skeletal muscle index (SMI), skeletal muscle density (SMD), total adipose index (TAI) were obtained from computed tomography scans (CT). Cancer associated weight loss (CAWL), an ordinal classification of history of weight loss and body mass index (BMI) was obtained from medical record. Serum protein profiles and concentrations were generated using SOMAscan, a quantitative aptamer-based assay. Ontological analysis of the proteins correlated with clinical variables (r[&ge;] 0.5 and p<0.05) was performed using DAVID Bioinformatics. Protein co-expression analysis was determined using pairwise Spearmans correlation.\n\nResultsOverall, 111 proteins of 1298 correlated with these clinical measures, 48 proteins for CAWL, 19 for SMI, 14 for SMD, and 30 for TAI. LYVE1, a homolog of CD44 implicated in tumor metastasis, was the top CAWL-associated protein (r= 0.67, p=0.0001). Other proteins such as INHBA, MSTN/GDF11, and PIK3R1 strongly correlated with CAWL. Proteins correlated with cachexia included those associated with proteolysis, acute inflammatory response, as well as B cell and T cell activation. Protein co-expression analysis identified networks such as activation of immune related pathways such as B-cell signaling, Th1 and Th2 pathways, natural killer cell signaling, IL6 signaling, and mitochondrial dysfunction.\n\nConclusionTaken together, these data both identify immune system molecules and additional secreted factors and pathways not previously associated with PDAC and confirm the activation of previously identified pathways. Identifying altered secreted factors in serum of PDAC patients may assist in developing minimally invasive laboratory tests for clinical cachexia as well as identifying new mediators.

cancer biology

Compound heterozygous ZP1 mutations cause empty follicle syndrome in infertile sisters

PurposeEmpty follicle syndrome (EFS) is a condition in which no oocyte is retrieved from mature follicles after proper ovarian stimulation in an in vitro fertilization (IVF) procedure. Genetic evidence accumulates for the etiology of recurrent EFS even with improved medical treatment which had avoided the pharmacological or iatrogenic problems. Here, this study investigated the genetic cause of recurrent EFS in a family with two infertile sisters.\n\nMethodsIn this work, we present two infertile sisters in a family with recurrent EFS after three cycles of standard ovarian stimulation with hCG and/or GnRHa therapy. We performed whole-exome sequencing and targeted sequencing in the core members of this family, and further bioinformatics analysis to identify pathogenesis of gene.\n\nResultsWe identified compound heterozygous variants, c.161_165del (p.54fs) and c.1166_1173del (p.389fs), on zona pellucida glycoprotein 1 (ZP1) gene, which were shared with two infertile sisters. Cosegregation tests on the affected and unaffected members of this family confirmed that the allelic mutants were transmitted from either parent.\n\nConclusionsThis EFS phenotype was distinct from the previously reported disruption of zona pellucida due to homozygous ZP1 defects. We thus propose that the specific mutations in ZP1 gene may render a causality for the recurrent EFS.

genetics

Alzheimer’s disease risk SNPs show no strong effect on miRNA expression in human lymphoblastoid cell lines

The role of microRNAs (miRNAs) in the pathogenesis of Alzheimers disease (AD) is currently extensively investigated. In this study, we assessed the potential impact of AD genetic risk variants on miRNA expression by performing large-scale bioinformatic data integration. Our analysis was based on genetic variants from three AD genome-wide association studies (GWAS). Association with miRNA expression was tested by expression quantitative trait loci (eQTL) analysis using next-generation miRNA sequencing data generated in lymphoblastoid cell lines (LCL). While, overall, we did not identify a strong effect of AD GWAS variants on miRNA expression in this cell type we highlight two notable outliers, i.e. miR-29c-5p and miR-6840-5p. MiR-29c-5p was recently reported to be involved in the regulation of BACE1 and SORL1 expression. In conclusion, despite two exceptions our large-scale assessment provides only limited support for the hypothesis that AD GWAS variants act as miRNA eQTLs.

genetics

Lmx1a drives Cux2 expression in the cortical hem through activation of a conserved intronic enhancer.

During neocortical development, neurons are produced by a diverse pool of neural progenitors. A subset of progenitors express the Cux2 gene and are fate-restricted to produce certain neuronal subtypes, but the upstream pathways that specify these progenitor fates remain unknown. To uncover the transcriptional networks that regulate Cux2 expression in the forebrain, we characterized a conserved Cux2 enhancer that we find recapitulates Cux2 expression specifically in the cortical hem. Using a bioinformatic approach, we found several potential transcription factor (TF) binding sites for cortical hem-patterning TFs. We found that the homeobox transcription factor, Lmx1a, can activate the Cux2 enhancer in vitro. Furthermore, we show that multiple Lmx1a binding sites required for enhancer activity in the cortical hem in vivo. Mis-expression of Lmx1a in neocortical progenitors caused an increase in Cux2+-lineage cells. Finally, we compared several conserved human enhancers with cortical hem-restricted activity and found that recurrent Lmx1a binding sites are a top shared feature. Uncovering the network of TFs involved in regulating Cux2 expression will increase our understanding of the mechanisms pivotal in establishing Cux2-lineage fates in the developing forebrain.\n\nSummary StatementAnalysis of a cortical hem-specific Cux2 enhancer reveals role for Lmx1a as a critical upstream regulator of Cux2 expression patterns in neural progenitors during early forebrain development.

neuroscience

Identification of novel RNA viruses associated to bird’s-foot trefoil (Lotus corniculatus)

Birds-foot trefoil (Lotus corniculatus) is a nutritious forage crop, employed for livestock foraging around the world. Here, we report the identification and characterization of two novel viruses associated with birds-foot trefoil. Virus sequences with affinity to enamoviruses (ssRNA (+); Luteoviridae; Enamovirus) and nucleorhabdoviruses (ssRNA (-); Rhabdoviridae; Nucleorhabdovirus) were detected in L. corniculatus transcriptome data. The tentatively named birds-foot trefoil associated virus 1 (BFTV-1) genome organization is characterized by 13,626 nt long negative-sense ssRNA. BFTV-1 presents in its antigenome orientation six predicted gene products in the canonical order 3'-N-P-P3-M-G-L-5'. The proposed birds-foot trefoil associated virus 2 (BFTV-2) 5,736 nt virus sequence presents a typical 5'-PO-P1-2-IGS-P3-P5-3' enamovirus genome structure. Phylogenetic analysis suggests that BFTV-1 is closely related to Datura yellow vein nucleorhabdovirus, and that BFTV-2 clusters into a monophyletic cluster of legumes-associated enamoviruses. This sub-clade of highly related and co-divergent legume associated viruses provides insights on the evolutionary history of the enamoviruses. The bioinformatic reanalysis of SRA libraries deposited in the NCBI database constitutes an emerging approach to the discovery of novel plant viruses which should be important for both quarantine purposes and disease management.

plant biology

The structural complexity of the Gammaproteobacteria flagellar motor is related to the type of its torque-generating stators

The bacterial flagellar motor is a cell-envelope-embedded macromolecular machine that functions as a propeller to move the cell. Rather than being an invariant machine, the flagellar motor exhibits significant variability between species, allowing bacteria to adapt to, and thrive in, a wide range of environments. For instance, different torque-generating stator modules allow motors to operate in conditions with different pH and sodium concentrations and some motors are adapted to drive motility in high-viscosity environments. How such diversity evolved is unknown. Here we use electron cryo-tomography to determine the in situ macromolecular structures of the flagellar motors of three Gammaproteobacteria species: Legionella pneumophila, Pseudomonas aeruginosa, and Shewanella oneidensis MR-1, providing the first views of intact motors with dual stator systems. Complementing our imaging with bioinformatics analysis, we find a correlation between the stator system of the motor and its structural complexity. Motors with a single H+-driven stator system have only the core P- and L-rings in their periplasm; those with dual H+-driven stator systems have an extra component elaborating their P-ring; and motors with Na+- (or dual Na+-H+)- driven stator systems have additional rings surrounding both their P- and L-rings. Our results suggest an evolution of structural complexity that may have enabled pathogenic bacteria like L. pneumophila and P. aeruginosa to colonize higher-viscosity environments in animal hosts.

microbiology

The landscape of S100B+ and HLA-DR+ dendritic cell subsets in tonsils at the single cell level via high-parameter mapping

Dendritic cells (DC) (classic, plasmacytoid, inflammatory) are an intense focus of interest because of their role in inflammation, autoimmunity, vaccination and cancer. We present a tissue-based classification of human DC subsets in tonsils with a high-parameter (>40 markers) immunofluorescent approach, cell type-specific image segmentation and the use of bioinformatics platforms. Through this deep phenotypic and spatial examination, classic cDC1, cDC2, pDC subsets have been further refined and a novel subset of DC co-expressing IRF4 and IRF8 identified. Based on unique tissue locations within the tonsil, and close interactions with T cells (cDC1) or B cells (cDC2), DC subsets can be further subdivided by correlative phenotypic changes associated with these interactions. In addition, monocytes and macrophages expressing HLA-DR or S100AB are identified and localized in the tissue. This study thus provides a whole tissue in situ catalog of human DC subsets and their cellular interactions within spatially defined niches.

immunology

sigfit: flexible Bayesian inference of mutational signatures

Mutational signature analysis aims to infer the mutational spectra and relative exposures of processes that contribute mutations to genomes. Different models for signature analysis have been developed, mostly based on non-negative matrix factorisation or non-linear optimisation. Here we present sigfit, an R package for mutational signature analysis that applies Bayesian inference to perform fitting and extraction of signatures from mutation data. We compare the performance of sigfit to prominent existing software, and find that it compares favourably. Moreover, sigfit introduces novel probabilistic models that enable more robust, powerful and versatile fitting and extraction of mutational signatures and broader biological patterns. The package also provides user-friendly visualisation routines and is easily integrable with other bioinformatic packages.

cancer biology