bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Systematic identification of circular RNAs and corresponding regulatory networks unveil their potential roles in the midgut of Apis cerana cerana workers

BackgroundCircular RNAs (circRNAs) are newly discovered noncoding RNAs (ncRNAs) that play key roles in various biological functions, such as the regulation of gene expression and alternative splicing. CircRNAs have been identified in some species, including western honeybees. However, the understanding of honeybee circRNA is still very limited, and to date, no study on eastern honeybee circRNA has been conducted. Here, the circRNAs in the midguts of Apis cerana cerana workers were identified and validated, and the regulatory networks were constructed. Differentially expressed circRNAs (DEcircRNAs) and the corresponding competitively endogenous RNA (ceRNA) networks in the development of the workers midgut were further investigated. ResultsHere, 7- and 10-day-old A. c. cerana workers midguts (Ac1 and Ac2) were sequenced using RNA-seq, and a total of 9589 circRNAs were predicted using bioinformatics. These circRNAs were approximately 201-800 nt in length and could be classified into six types; the annotated exonic circRNAs were the most abundant. Additionally, five novel A. c. cerana circRNAs were confirmed by PCR amplification and Sanger sequencing, indicating the authenticity of A. c. cerana circRNAs. Interestingly, novel_circ_003723, novel_circ_002714, novel_circ_002451 and novel_circ_001980 were the most highly expressed circRNAs in both Ac1 and Ac2, which is indicative of their key roles in the development of the midgut. Moreover, 55 DEcircRNAs were identified in the Ac1 vs Ac2 comparison group, including 34 upregulated and 21 downregulated circRNAs. Further investigation showed that the source genes of circRNAs were classified into 34 GO terms and were involved in 141 KEGG pathways. In addition, the source genes of DEcircRNAs were categorized into 10 GO terms and 15 KEGG pathways, which demonstrated that the corresponding DEcircRNAs may affect the growth, development, and material and energy metabolisms of the workers midgut by regulating the expression of the related source genes. Additionally, the circRNA-miRNA regulatory networks were constructed and analyzed, and the results demonstrated that 1060 circRNAs can bind to 74 miRNAs and that 71.51% of circRNAs can be linked to only one miRNA. Furthermore, the DEcircRNA-miRNA-mRNA networks were constructed and explored, and the results indicate that the 13 downregulated circRNAs can bind to eight miRNAs and to 29 target genes. In addition, the results indicate that the 16 upregulated circRNAs can bind to 9 miRNAs and to 29 target genes, demonstrating that DEcircRNAs are likely involved in the regulation of midgut development via ceRNA mechanisms. Moreover, the regulatory networks of miR-6001-y-targeted DEcircRNAs were analyzed, and the results showed that eight DEcircRNAs may affect the development of A. c. cerana workers midguts by targeting miR-6001-y. Finally, four randomly selected DEcircRNAs were verified via RT-qPCR, confirming the reliability of our sequencing data. ConclusionThis is the first systematic investigation of circRNAs and their corresponding regulatory networks in eastern honeybees. The identified circRNAs from the A. c. cerana workers midgut will enrich the known reservoir of honeybee ncRNAs. DEcircRNAs may play a comprehensive role during the development of the workers midgut via the regulation of source genes and the interaction with miRNAs by acting as ceRNAs. The eight DEcircRNAs that targeted miR-6001-y were likely to be vital for the development of the workers midgut. Our results provide a valuable resource for the future studies of A. c. cerana circRNA and lay a foundation to reveal the molecular mechanisms underlying the regulatory networks of circRNAs responsible for the workers midgut development; in addition, these findings facilitate a functional study on the key circRNAs involved in the developmental process. Graphical Abstract O_FIG_DISPLAY_L [Figure 1] M_FIG_DISPLAY C_FIG_DISPLAY

molecular biology

Rapid, multiplexed, whole genome and plasmid sequencing of foodborne pathogens using long-read nanopore technology

United States public health agencies are focusing on next-generation sequencing (NGS) to quickly identify and characterize foodborne pathogens. Here, the MinION nanopore, long-read sequencer was used to simultaneously sequence the entire chromosome and plasmids of Salmonella enterica subsp. enterica serovar Bareilly and Escherichia coli O157:H7. A rapid, random sequencing approach, coupled with de novo genome assembly within a customized data analysis workflow, that can resolve highly-repetitive genomic regions, was developed. In sequencing runs, as short as four hours, using nanopore data alone, full-length genomes were obtained with an average identity of 99.87% for Salmonella Bareilly and 99.89% for E. coli in comparison to the respective MiSeq references. These long-read assemblies provided information on serotype, virulence factors, and antimicrobial resistance genes. Using a custom-developed, SNP-selection workflow, the potential of the nanopore-only assemblies (after only 30 minutes of sequencing) for rapid phylogenetic inference, with identical topology compared to the published dataset, was demonstrated. To achieve maximum quality assemblies, the developed bioinformatics workflow employed additional polishing steps to correct the systematic errors produced by the nanopore-only assemblies. Nanopore sequencing provided a shorter (10 hours library preparation and sequencing) turnaround time compared to other NGS technologies.

microbiology

Structure-guided function discovery of an NRPS-like glycine betaine reductase for choline biosynthesis in fungi

Nonribosomal peptide synthetases (NRPS) and NRPS-like enzymes have diverse functions in primary and secondary metabolism. By using a structure-guided approach, we uncovered the function of an NRPS-like enzyme with unusual domain architecture, catalyzing two sequential two-electron reductions of glycine betaine to choline. Structural analysis based on homology model suggests cation-{pi} interactions as the major substrate specificity determinant, which was verified using substrate analogs and inhibitors. Bioinformatic analysis indicates this NRPS-like glycine betaine reductase is highly conserved and widespread in fungi kingdom. Genetic knockout experiments confirmed its role in choline biosynthesis and maintaining glycine betaine homeostasis in fungi. Our findings demonstrate that the oxidative choline-glycine betaine degradation pathway can operate in a fully reversible fashion and provide new insights in understanding fungal choline metabolism. The use of an NRPS-like enzyme for reductive choline formation is energetically efficient compared to known pathways. Our discovery also underscores the capabilities of structure-guided approach in assigning function of uncharacterized multidomain proteins, which can potentially aid functional discovery of new enzymes by genome mining.

biochemistry

OxyR senses reactive sulfane sulfur and activates genes for its removal in Escherichia coli

Reactive sulfane sulfur species such as hydrogen polysulfide and organic persulfide are newly recognized as normal cellular components, involved in signaling and protecting cells from oxidative stress. Their production is extensively studied, but their removal is less characterized. Herein, we showed that reactive sulfane sulfur is toxic at high levels, and it is mainly removed via reduction by thioredoxin and glutaredoxin with the release of H2S in Escherichia coli. OxyR is best known to respond to H2O2, and it also played an important role in responding to reactive sulfane sulfur under both aerobic and anaerobic conditions. It was modified by hydrogen polysulfide to OxyR C199-SSH, which activated the expression of thioredoxin 2 and glutaredoxin 1. This is a new type of OxyR modification. Bioinformatics analysis showed that OxyRs are widely present in bacteria, including strict anaerobic bacteria. Thus, the OxyR sensing of reactive sulfane sulfur may represent a conserved mechanism for bacteria to deal with sulfane sulfur stress.

microbiology

Quantitative proteomic profiling of tumor-associated vascular endothelial cells in colorectal cancer

SummeryTo investigate the global proteomic profiles of vascular endothelial cells (VECs) in the tumor microenvironment and antiangiogenic therapy for colorectal cancer (CRC), matched pairs of normal (NVECs) and tumor-associated VECs (TVECs) were purified from CRC tissues by laser capture microdissection and subjected to iTRAQ based quantitative proteomics analysis. Here, 216 differentially expressed proteins (DEPs) were identified and performed bioinformatics analysis. Interestingly, these proteins were implicated in epithelial mesenchymal transition (EMT), ECM-receptor interaction, focal adhesion, PI3K-Akt signaling pathway, angiogenesis and HIF-1 signaling pathway, which may play important roles in CRC angiogenesis. Among these DEPs, Tenascin-C (TNC) was found to upregulated in the TVECs of CRC and be correlate with CRC multistage carcinogenesis and metastasis. Furthermore, the reduction of tumor-derived TNC could attenuate human umbilical vein endothelial cell (HUVEC) proliferation, migration and tube formation through ITGB3/FAK/Akt signaling pathway. Based on the present work, we provided a large-scale proteomic profiling of VECs in CRC with quantitative information, a certain number of potential antiangiogenic targets and a novel vision in the angiogenesis bio-mechanism of CRC. Summery statementWe provided large-scale proteomic profiling of vascular endothelial cells in colorectal cancer with quantitative information, a number of potential antiangiogenic targets and a novel vision in the angiogenesis bio-mechanism of CRC.

cancer biology

High-coverage, long-read sequencing of Han Chinese trio reference samples.

Single-molecule long-read sequencing datasets were generated for a son-father-mother trio of Han Chinese descent that is part of the Genome In a Bottle (GIAB) consortium portfolio. The dataset was generated using the Pacific Biosciences Sequel System. The son and each parent were sequenced to an average coverage of 60 and 30, respectively, with N50 subread lengths between 16 and 18 kb. Raw reads and reads aligned to both the GRCh37 and GRCh38 are available at the NCBI GIAB ftp site (ftp://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/ChineseTrio/) and the raw read data is archived in NCBI SRA (SRX4739017, SRX4739121, and SRX4739122). This dataset is available for anyone to develop and evaluate long-read bioinformatics methods.

genomics

MicroRNA-138 negatively regulates the hypoxia-inducible factor 1α to suppress melanoma growth and metastasis

Melanoma with rapid progression towards metastasis becomes the deadliest form of skin cancer. However, the mechanism of melanoma growth and metastasis is still unclear. Here, we found that miRNA-138 was low expression and hypoxia-inducible factor 1 (HIF1) was high expression in the patients melanoma tissue, and they had a significant negative correlation (r=-0.937, P < 0.001). Patients with miRNA-138low/HIF1high signature were predominant in late stage. Further, bioinformatic analysis demonstrated that miRNA-138 directly targeted HIF1. We found that the introduction of miRNA-138 mimics to A375 cells could reduced HIF1 mRNA expression, and suppressed the cell proliferation, migration and invasion. Overexpression of miRNA-138 or inhibition of HIF1 significantly suppressed the growth and metastasis of melanoma in vivo. Our study demonstrates the role and clinical relevance of miRNA-138 and HIF1 in melanoma cell growth and metastasis, providing a novel therapeutic target for suppression of melanoma growth and metastasis.

cancer biology

Single-Cell RNA Sequencing Reveals Regulatory Mechanism for Trophoblast Cell-Fate Divergence in Human Peri-Implantation Embryo

Multipotent trophoblasts undergo dynamic morphological movement and cellular differentiation after embryonic implantation to generate placenta. However, the mechanism controlling trophoblast development and differentiation during peri-implantation development remains elusive. In this study, we modeled human embryo peri-implantation development from blastocyst to early post-implantation stages by using an in vitro coculture system, and profiled the transcriptome of individual trophoblast cells from these embryos. We revealed the genetic networks regulating peri-implantation trophoblast development. While determining when trophoblast differentiation happens, our bioinformatic analysis identified T-box transcription factor 3 (TBX3) as a key regulator for the differentiation of cytotrophoblast into syncytiotrophoblast. The function of TBX3 in trophoblast differentiation is then validated by a loss-of-function experiment. In conclusion, our results provided a valuable resource to study the regulation of trophoblasts development and differentiation during human peri-implantation development.

cell biology

Pulcherrimin formation controls growth arrest of the Bacillus subtilis biofilm

Biofilm formation by Bacillus subtilis is a communal process that culminates in the formation of architecturally complex multicellular communities. Here we reveal that the transition of the biofilm into a non-expanding phase constitutes a distinct step in the process of biofilm development. Using genetic analysis we show that B. subtilis strains lacking the ability to synthesize pulcherriminic acid form biofilms that sustain the expansion phase, thereby linking pulcherriminic acid to growth arrest. However, production of pulcherriminic acid is not sufficient to block expansion of the biofilm. It needs to be secreted into the extracellular environment where it chelates Fe3+ from the growth medium in a non-enzymatic reaction. Utilizing mathematical modelling and a series of experimental methodologies we show that when the level of freely available iron in the environment drops below a critical threshold, expansion of the biofilm stops. Bioinformatics analysis allows us to identify the genes required for pulcherriminic acid synthesis in other Firmicutes but the patchwork presence both within and across closely related species suggests loss of these genes through multiple independent recombination events. The seemingly counterintuitive self-restriction of growth led us to explore if there were any benefits associated pulcherriminic acid production. We identified that pulcherriminic acid producers can prevent invasion from neighbouring communities through the generation of an \"iron free\" zone thereby addressing the paradox of pulcherriminic acid production by B. subtilis.\n\nSignificanceUnderstanding the processes that underpin the mechanism of biofilm formation, dispersal, and inhibition are critical to allow exploitation and to understand how microbes thrive in the environment. Here, we reveal that the formation of an extracellular iron chelate restricts the expansion of a biofilm. The countering benefit to self-restriction of growth is protection of an environmental niche. These findings highlight the complex options and outcomes that bacteria need to balance in order to modulate their local environment to maximise colonisation, and therefore survival.

microbiology

Identification and characterization of putative Aeromonas spp. T3SS effectors

The genetic determinants of bacterial pathogenicity are highly variable between species and strains. However, a factor that is commonly associated with virulent Gram-negative bacteria, including many Aeromonas spp., is the type 3 secretion system (T3SS), which is used to inject effector proteins into target eukaryotic cells. In this study, we developed a bioinformatics pipeline to identify T3SS effector proteins, applied this approach to the genomes of 105 Aeromonas strains isolated from environmental, mutualistic, or pathogenic contexts and evaluated the cytotoxicity of the identified effectors through their heterologous expression in yeast. The developed pipeline uses a two-step approach, where candidate families are initially selected using HMM profiles with minimal similarity scores against the Virulence Factors DataBase (VFDB), followed by strict comparisons against positive and negative control datasets, greatly reducing the number of false positives. Using our approach, we identified 21 Aeromonas T3SS likely effector groups, of which 8 represented known or characterized effectors, while the remaining 13 had not previously been described in Aeromonas. We experimentally validated our in silico findings by assessing the cytotoxicity of representative effectors in Saccharomyces cerevisiae BY4741, with 15 out of 21 assayed proteins eliciting a cytotoxic effect in yeast. The results of this study demonstrate the utility of our approach, combining a novel in silico search method with in vivo experimental validation, and will be useful in future research aimed at identifying and authenticating bacterial effector proteins from other genera.

microbiology

An ancient lineage of highly divergent parvoviruses infects both vertebrate and invertebrate hosts.

Chapparvoviruses are a highly divergent group of parvoviruses (family Parvoviridae) first identified in 2013. Interest in these poorly characterized viruses has been raised by recent studies indicating that they are the cause of chronic kidney disease that arises spontaneously in laboratory mice. In this study, we investigate the biological and evolutionary characteristics of chapparvoviruses via comparative analysis of genome sequence data. Our analysis, which incorporates sequences derived from endogenous viral elements (EVEs) as well as exogenous viruses, reveals that chapparvoviruses are an ancient lineage within the family Parvoviridae, clustering separately from members of both currently established parvoviral subfamilies. Consistent with this, they exhibit a number of characteristic genomic and structural features, i.e. a large number of putative auxiliary protein-encoding genes, capsid protein genes non-homologous to any hitherto parvoviral cap, as well as a putative capsid structure lacking the canonical fifth strand of the ABIDG sheet comprising the luminal side of the jelly roll. Our findings demonstrate that the chapparvovirus lineage infects an exceptionally broad range of host species, including both vertebrates and invertebrates. Furthermore, we observe that chapparvoviruses found in fish are more closely related to those from invertebrates than they are to those that infect amniote vertebrates. This suggests that transmission between distantly related host species may have occurred in the past. Our study provides the first integrated overview of the chapparvovirus group, and revises current views of parvovirus evolution\n\nAUTHOR SUMMARYChapparvoviruses are a recently identified group of viruses about which relatively little is known. However, recent studies have shown that these viruses cause disease in laboratory mice and are prevalent in the fecal virome of pigs and poultry, raising interest in their potential impact as pathogens, and utility as experimental tools. We examined the genomes of chapparvoviruses and endogenous viral elements ( fossilized virus sequences derived from ancestral viruses) using a variety of bioinformatics-based approaches. We show that the chapparvoviruses have an ancient origin and are evolutionarily distinct from all other related viruses. Accordingly, their genomes and virions exhibit a range of distinct characteristic features. We examine the distribution of these features in the light of chapparvovirus evolutionary history (which we can also infer from genomic data), revealing new insights into chapparvovirus biology.

evolutionary biology

CRISPR disruption and UK Biobank analysis of a highly conserved polymorphic enhancer suggests its role in anxiety and male alcohol intake.

Excessive alcohol intake is associated with 5.9% of global deaths. However, this figure is especially acute in men such that 7.6% of deaths can be attributed to alcohol intake. Previous studies identified a significant interaction between genotypes of the galanin (GAL) gene with anxiety and alcohol abuse in different male populations but were unable to define a mechanism. To address these issues the current study analysed the human UK Biobank cohort and identified a significant interaction (n=115,865; p=0.0007) between allelic variation (GG or CA genotypes) in the highly conserved human GAL5.1 enhancer, alcohol intake (AUDIT questionnaire scores) and anxiety in men that was consistent with these previous studies. Critically, disruption of GAL5.1 in mice using CRISPR genome editing significantly reduced GAL expression in the amygdala and hypothalamus whilst producing a corresponding reduction in ethanol intake in KO mice. Intriguingly, we also found evidence of reduced anxiety-like behaviour in male GAL5.1KO animals mirroring that seen in humans. Using bioinformatic analysis and co-transfection studies we further identified the EGR1 transcription factor, that is co expressed with GAL in amygdala and hypothalamus, as being important in the protein kinase C (PKC) supported activity of the GG genotype of GAL5.1 but less so in the CA genotype. Our unique study uses a novel combination of human association analysis, CRISPR genome editing in mice, animal behavioural analysis and cell culture studies to identify a highly conserved regulatory mechanism linking anxiety and alcohol intake that might contribute to increased susceptibility to anxiety and alcohol abuse in men.

genetics

Genetic analysis of TP53 gene mutations in exon 4 and exon 8 among esophageal cancer patients in Sudan

BackgroundEsophageal carcinoma (EC) represents the 1st rank among all gastrointestinal cancers in Sudan. Despite little publications, there is a deep absence of literature about the molecular pathogenesis of EC considering TP53 gene from Sudanese population.\n\nAimsIn this study, we performed the expression analysis on p53 protein level by immunohistochemical staining and examined its overexpression with p53 mutations in exons 4 and 8 among esophageal cancer patients in Sudan.\n\nMaterial and MethodsFixed tissue with 10% buffered formalin was stained by Hematoxlin and Eosin (H&E), Alcian blue-Periodic Acid Schiff (PAS) and Immunohistochemistry stain. PCR-RFLP was used to study the frequencies of p53 codon 72 R/P polymorphism. Conventional PCR and sanger sequencing were applied for exon 4 and exon 8. Then detection and functional analysis of SNPs and mutations were performed using various in bioinformatics tools.\n\nResultNuclear accumulations for p53 protein was detected in all of the esophageal carcinomas examined while no accumulations were observed in normal control sections. Four patients with immune-positive for p53 showed no mutations in p53 gene (exon4 and exon8). The incidence of the homozygous mutant variant Pro/Pro was higher in esophageal cancerous patients comparing to healthy control subject 20(71. 4%) vs. 1(10%), respectively (p=0.0026). In exon 4, no mutation was detected other than NG_017013.2:g. 16397C>G. While in exon 8, g.18783-18784AG>TT, g.18803A>C, g.18860A>C, g.18845A>T and g.18863_ 18864 InsT were observed.\n\nConclusionwe found a significant association between the overexpression of TP53 protein and mutation in exon 4 and 8. A silent mutation P301P was detected in all of examined cases. Two patients who diagnosed with small cell sarcoma have shared the same mutations in exon8. Further studies with large sample size are required to demonstrate the usefulness of these mutations in the screening of EC especially SCCE.

cancer biology

De-novo variants in XIRP1 associated with polydactyly and polysyndactyly in Holstein cattle

Congenital polydactylous cattle are sporadically observed. Impairment of the limb patterning process due to altered control of the zone of polarizing activity (ZPA) was associated in several species with preaxial polydactyly and syndactyly. In cattle, the role of ZPA and other genes involved in limb patterning for polydactyly was not yet elucidated. Herein, we report on a preaxial type II polydactyly and a praeaxial type II+V polysyndactyly in two Holstein calves and screen whole genome sequencing data for associated variants. Using whole genome sequencing data of both affected calves did not show mutations in the candidate regions of ZRS and pZRS or in candidate genes associated with polydactyly, syndactyly and polysyndactyly in other species. Two indels, which are located in XIRP1 within a common haplotype, were highly associated with the two phenotypes. Bioinformatic analyses retrieved an interaction between XIRP1 and FGFR1, CTNNB1 and CTNND1 supporting a link between the XIRP1 variants and embryonic limb patterning. The heterozygous haplotype was highly associated with the present polydactylous phenotypes due to dominant mode of inheritance with an incomplete penetrance in Holstein cattle.

genetics

A validated workflow for rapid taxonomic assignment and monitoring of a national fauna of bees (Apiformes) using high throughput barcoding

Improved taxonomic methods are needed to quantify declining populations of insect pollinators. This study devises a high-throughput DNA barcoding protocol for a regional fauna (United Kingdom) of bees (Apiformes), consisting of reference library construction, a proof-of-concept monitoring scheme, and the deep barcoding of individuals to assess potential artefacts and organismal associations. A reference database of Cytochrome Oxidase subunit 1 (cox1) sequences including 92.4% of 278 bee species known from the UK showed high congruence with morphological taxon concepts, but molecular species delimitations resulted in numerous split and (fewer) lumped entities within the Linnaean species. Double tagging permitted deep illumina sequencing of 762 separate individuals of bees from a UK-wide survey. Extracting the target barcode from the amplicon mix required a new protocol employing read abundance and phylogenetic position, which revealed 180 molecular entities of Apiformes identifiable to species. An additional 72 entities were ascribed to mitochondrial pseudogenes based on patterns of read abundance and phylogenetic relatedness to the reference set. Clustering of reads revealed a range of secondary Operational Taxonomic Units (OTUs) in almost all samples, resulting from traces of insect species caught in the same traps, organisms associated with the insects including a known mite parasite of bees, and the common detection of human DNA, besides evidence for low-level cross-contamination in pan traps and laboratory steps. Custom scripts were generated to conduct critical steps of the bioinformatics protocol. The resources built here will greatly aid DNA-based monitoring to inform management and conservation policies for the protection of pollinators.

molecular biology

Nanopore sequencing of the pharmacogene CYP2D6 allows simultaneous haplotyping and detection of duplications

BackgroundThe accurate genotyping of CYP2D6 is hindered by the very polymorphic nature of the gene, high homology with its pseudogene CYP2D7, and the occurrence of structural variations. Long read sequencing offers the promise of overcoming some of these challenges, along with the advantage of straightforward variant phasing. We have established methods for sequencing and analysis of DNA amplicons containing the whole CYP2D6 gene, using the GridION nanopore sequencer.\n\nMaterials and methodsSeven reference and 25 clinical samples covering various haplotypes including gene duplication were barcoded and sequenced over two sequencing runs. Sequenced raw reads were analyzed using a pipeline of bioinformatics tools including two mapping tools and two variant calling tools.\n\nResultsUsing minimap2 and nanopolish (mapping and variant calling tools respectively) resulted in the most accurate variant detection. Haplotypes of 52 alleles could be matched accurately to known alleles or subvariants, while the remaining 12 alleles being assigned as novel star (*) allele of novel subvariants of known alleles in the PharmVar CYP2D6 haplotype database. Allele duplication could be detected by analyzing the allelic balance between the sample haplotypes.\n\nConclusionNanopore sequencing of CYP2D6 offers a high throughput method for genotyping, accurate haplotyping, and detection of new variants and duplicated alleles.

genetics

Unbiased metagenomic sequencing for pediatric meningitis in Bangladesh reveals neuroinvasive Chikungunya virus outbreak and other unrealized pathogens

The disease burden due to meningitis in low and middle-income countries remains significant and failure to determine an etiology impedes appropriate treatment for patients and evidence-based policy decisions for populations. Broad-range pathogen surveillance using metagenomic next-generation sequencing (mNGS) of RNA isolated from cerebral spinal fluid (CSF) provides an unbiased assessment for possible infectious etiologies. In this study, our objective was to use mNGS to identify etiologies of pediatric meningitis in Bangladesh.\n\nWe conducted a retrospective case-control mNGS study on CSF from patients with known neurologic infections (n=36), idiopathic meningitis (n=25), without infection (n=30) and six environmental samples collected between 2012-2018. Using an open-access, cloud-based bioinformatics pipeline (IDseq) and machine learning, we identified potential pathogens which were confirmed through qPCR and Sanger sequencing. These cases were followed-up through phone/home-visits. The CSF samples were collected from children with WHO-defined meningeal signs during prospective meningitis surveillance at the largest pediatric referral hospital in Bangladesh.\n\nThe 91 participants (42% female) ranged in age from 0-160 months (median: 9 months). In samples with known infectious causes of meningitis and without infections (n=66), there was 83% concordance between mNGS and conventional testing. In idiopathic cases (n=25), mNGS identified a potential etiology in 40% (n=10), including bacterial and viral pathogens. There were three instances of neuroinvasive Chikungunya virus (CHIKV). The CHIKV genomes were >99% identical to each other and to a Bangladeshi strain only previously recognized to cause systemic illness in 2017. CHIKV qPCR of all remaining stored CSF samples from children who presented with idiopathic meningitis in 2017 at the same hospital (n=472) revealed 17 additional CHIKV meningitis cases. Orthogonal molecular confirmation of each mNGS-identified infection, case-based clinical data, and follow-up of patients substantiated the key findings.\n\nUsing mNGS, we obtained a microbiological diagnosis for 40% of idiopathic meningitis cases and identified a previous unappreciated pediatric CHIKV meningitis outbreak. Case-control CSF mNGS surveys can complement conventional diagnostic methods to identify etiologies of meningitis and facilitate informed policy decisions.

genomics

An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species

AO_SCPLOWBSTRACTC_SCPLOWExon capture coupled to high-throughput sequencing constitutes a cost-effective technical solution for addressing specific questions in evolutionary biology by focusing on expressed regions of the genome preferentially targeted by selection. Transcriptome-based capture, a process that can be used to capture the exons of non-model species, is use in phylogenomics. However, its use in population genomics remains rare due to the high costs of sequencing large numbers of indexed individuals across multiple populations. We evaluated the feasibility of combining transcriptome-based capture and the pooling of tissues from numerous individuals for DNA extraction as a cost-effective, generic and robust approach to estimating the variant allele frequencies of any species at the population level. We designed capture probes for [~]5 Mb of chosen de novo transcripts from the Asian ladybird Harmonia axyridis (5,717 transcripts). We called [~]300,000 bi-allelic SNPs for a pool of 36 non-indexed individuals. Capture efficiency was high, and pool-seq was as effective and accurate as individual-seq for detecting variants and estimating allele frequencies. Finally, we also evaluated an approach for simplifying bioinformatic analyses by mapping genomic reads directly to targeted transcript sequences to obtain coding variants. This approach is effective and does not affect the estimation of SNP allele frequencies, except for a small bias close to some exon ends. We demonstrate that this approach can also be used to predict the intron-exon boundaries of targeted de novo transcripts, making it possible to abolish genotyping biases near exon ends.

genomics