bioRxiv ScienceSearch

Biology subjects

Nakagawa, S.

Publications and source records attributed to Nakagawa, S..

12 recordsLinked to original sources

Direct PCR amplification of 16S rRNA genes offers accelerated bacterial identification using the MinION™ nanopore sequencer

Rapid identification of bacterial pathogens is crucial for appropriate and adequate antibiotic treatment, which significantly improves patient outcomes. 16S ribosomal RNA (rRNA) gene amplicon sequencing has proven to be a powerful strategy for diagnosing bacterial infections. We have recently established a sequencing method and bioinformatics pipeline for 16S rRNA gene analysis utilizing the Oxford Nanopore Technologies MinION sequencer. In combination with our taxonomy annotation analysis pipeline, the system enabled the molecular detection of bacterial DNA in a reasonable timeframe for diagnostic purposes. However, purification of bacterial DNA from specimens remains a rate-limiting step in the workflow. To further accelerate the process of sample preparation, we adopted a direct PCR strategy that amplifies 16S rRNA genes from bacterial cell suspensions without DNA purification. Our results indicate that differences in cell wall morphology significantly affect direct PCR efficiency and sequencing data. Notably, mechanical cell disruption preceding direct PCR was indispensable for obtaining an accurate representation of the specimen bacterial composition. Furthermore, 16S rRNA gene analysis of mock polymicrobial samples indicated that primer sequence optimization is required to avoid preferential detection of particular taxa and to cover a broad range of bacterial species. This study establishes a relatively simple workflow for rapid bacterial identification via MinIONTM sequencing, which reduces the turnaround time from sample to result, and provides a reliable method that may be applicable to clinical settings.

microbiology

UPA-Seq: Prediction of Functional LncRNAs Using Differential Sensitivity to UV Crosslinking

While a large number of long noncoding RNAs (lncRNAs) are transcribed from the genome of higher eukaryotes, systematic prediction of their functionality has been challenging due to the lack of conserved sequence motifs or structures. Assuming that lncRNAs function as large ribonucleoprotein complexes and thus are easily crosslinked to proteins upon UV irradiation, we performed RNA-Seq analyses of RNAs recovered from the aqueous phase after UV irradiation and phenol-chloroform extraction (UPA-Seq). As expected, the numbers of UPA-Seq reads mapped to known functional lncRNAs were remarkably reduced upon UV irradiation. Comparison with ENCODE eCLIP data revealed that lncRNAs that exhibited greater decreases upon UV irradiation preferentially associated with proteins containing prion-like domains (PrLDs). Fluorescent in situ hybridization (FISH) analyses revealed the nuclear localization of novel functional lncRNA candidates, including one that accumulated at the site of transcription. We propose that UPA-Seq provides a useful tool for the selection of lncRNA candidates to be analyzed in depth in subsequent functional studies.

molecular biology

Alopecia areata susceptibility variant identified by MHC risk haplotype sequencing reproduces symptomatic patched hair loss in mice

BackgroundAlopecia areata (AA) is a highly heritable multifactorial and complex disease. However, no convincing susceptibility gene has yet been pinpointed in the major histocompatibility complex (MHC), a region in the human genome known to be associated with AA as compared to other regions.\n\nResultsBy sequencing MHC risk haplotypes, we identified a variant (rs142986308, p.Arg587Trp) in the coiled-coil alpha-helical rod protein 1 (CCHCR1) gene as the only non-synonymous variant in the AA risk haplotype. Using CRISPR/Cas9 for allele-specific genome editing, we then phenocopied AA symptomatic patched hair loss in mice engineered to carry the Cchcr1 risk allele. Skin biopsies of these alopecic mice showed strong up-regulation of hair-related genes, including hair keratin and keratin-associated proteins (KRTAPs). Using transcriptomics findings, we further identified CCHCR1 as a novel component of hair shafts and cuticles in areas where the engineered alopecic mice displayed fragile and impaired hair.\n\nConclusionsThese results suggest an alternative mechanism for the aetiology of AA based on aberrant keratinization, in addition to generally well-known autoimmune events.

genomics

Meta-analysis challenges a textbook example of status signalling: evidence for publication bias

The status signalling hypothesis aims to explain conspecific variation in ornamentation by suggesting that some ornaments signal dominance status. Here, we use multilevel meta-analytic models to challenge the textbook example of this hypothesis, the black bib of house sparrows (Passer domesticus). We conducted a systematic review, and obtained raw data from published and unpublished studies to test whether dominance rank is positively associated with bib size across studies. Contrary to previous studies, our meta-analysis did not support this prediction. Furthermore, we found several biases in the literature that further question the support available for the status signalling hypothesis. First, the overall effect size of unpublished studies was zero, compared to the medium effect size detected in published studies. Second, the effect sizes of published studies decreased over time, and recently published effects were, on average, no longer distinguishable from zero. We discuss several explanations including pleiotropic, population- and context-dependent effects. Our findings call for reconsidering this established textbook example in evolutionary and behavioural ecology, raise important concerns about the validity of the current scientific publishing culture, and should stimulate renewed interest in understanding within-species variation in ornamental traits.

evolutionary biology

Planned missing data design: stronger inferences, increased research efficiency and improved animal welfare in ecology and evolution

O_LIEcological and evolutionary research questions are increasingly requiring the integration of research fields along with larger datasets to address fundamental local and global scale problems. Unfortunately, these agendas are often in conflict with limited funding and a need to balance animal welfare concerns.\nC_LIO_LIPlanned missing data design (PMDD), where data are randomly and deliberately missed during data collection, is a simple and effective strategy to working under greater research constraints while ensuring experiments have sufficient power to address fundamental research questions. Here, we review how PMDD can be incorporated into existing experimental designs by discussing alternative design approaches and evaluating how data imputation procedures work under PMDD situations.\nC_LIO_LIUsing realistic examples and simulations of multilevel data we show how a variety of research questions and data types, common in ecology and evolution, can be aided by using a PMDD with data imputation procedures. More specifically, we show how PMDD can improve statistical power in detecting effects of interest even with high levels (50%) of missing data and moderate sample sizes. We also provide examples of how PMDD can facilitate improved animal welfare and potentially alleviate research costs and constraints that would make endeavours for integrative research challenging.\nC_LIO_LIPlanned missing data designs are still in their infancy and we discuss some of the difficulties in their implementation and provide tentative solutions. Nonetheless, data imputation procedures are becoming more sophisticated and more easily implemented and it is likely that PMDD will be an effective and powerful tool for a wide range of experimental designs, data types and problems in ecology and evolution.\nC_LI

ecology

The application of zeta diversity as a continuous measure of compositional change in ecology

Zeta diversity provides the average number of shared species across n sites (or shared operational taxonomic units (OTUs) across n cases). It quantifies the variation in species composition of multiple assemblages in space and time to capture the contribution of the full suite of narrow, intermediate and wide-ranging species to biotic heterogeneity. Zeta diversity was proposed for measuring compositional turnover in plant and animal assemblages, but is equally relevant for application to any biological system that can be characterised by a row by column incidence matrix. Here we illustrate the application of zeta diversity to explore compositional change in empirical data, and how observed patterns may be interpreted. We use 10 datasets from a broad range of scales and levels of biological organisation - from DNA molecules to microbes, plants and birds - including one of the original data sets used by R.H. Whittaker in the 1960s to express compositional change and distance decay using beta diversity. The applications show (i) how different sampling schemes used during the calculation of zeta diversity may be appropriate for different data types and ecological questions, (ii) how higher orders of zeta may in some cases better detect shifts, transitions or periodicity, and importantly (iii) the relative roles of rare versus common species in driving patterns of compositional change. By exploring the application of zeta diversity across this broad range of contexts, our goal is to demonstrate its value as a tool for understanding continuous biodiversity turnover and as a metric for filling the empirical gap that exists on spatial or temporal change in compositional diversity.

ecology

Cell Type Specific Survey of Epigenetic Modifications by Tandem Chromatin Immunoprecipitation Sequencing

BackgroundThe nervous system of higher eukaryotes is composed of numerous types of neurons and glia that together orchestrate complex neuronal responses. However, this complex pool of cells typically poses analytical challenges in investigating gene expression profiles and their epigenetic basis for specific cell types. Here, we developed a novel method that enables cell type-specific analyses of epigenetic modifications using tandem chromatin immunoprecipitation sequencing (tChIP-Seq).\n\nResultsFLAG-tagged histone H2B, a constitutive chromatin component, was first expressed in Camk2a-positive pyramidal cortical neurons and used to purify chromatin in a cell type-specific manner. Subsequent chromatin immunoprecipitation using antibodies against H3K4me3--an active promoter mark--allowed us to survey neuron-specific coding and non-coding transcripts. Indeed, tChIP-Seq identified hundreds of genes associated with neuronal functions and genes with unknown functions expressed in cortical neurons.\n\nConclusionstChIP-Seq thus provides a versatile approach to investigating the epigenetic modifications of particular cell types in vivo.

molecular biology

Fixed effect variance and the estimation of the heritability: Issues and solutions

Linear mixed effects models are frequently used for estimating quantitative genetic parameters, including the heritability, of traits of interest. Heritability is an important metric, because it acts as a filter that determines how efficiently phenotypic selection translates into evolutionary change. As a quantity of biological interest, it is important that the denominator, the phenotypic variance, actually reflects the amount of phenotypic variance in the relevant ecological stetting. The current practice of quantifying heritability from mixed effects models frequently deprives the heritability of variance explained by fixed effects (often leading to upward-bias) and it has been suggested to omit fixed effects when estimating heritabilities. We advocate an alternative option of fitting complex models incorporating all relevant effects, while including the variance explained by fixed effects into the estimation of heritabilities. The approach is easily implemented (an example is provided) and allows corrections for the estimation of heritability, such as the exclusion of variance arising from experimental design effects while still including all biologically relevant sources of variation. We explore the complications arising depending on the nature of the covariates included as fixed effects (e.g. biological or experimental origin, characteristics of biological covariates). Furthermore, we discuss fixed effects in non-linear and generalized linear models when fixed effects. In these cases, the variance parameters depend on the location of the intercept and hence on the scaling of the fixed effects. Integration over the biologically relevant range of fixed effects offers a preferred solution in those situations.

evolutionary biology

Facultative adjustment of paternal care in the face of female infidelity in dunnocks

A much-debated issue is whether or not males should reduce parental care when they lose paternity (i.e. the certainty of paternity hypothesis). While there is general support for this relationship across species, within-population evidence is still contentious. Among the main reasons behind such problem is the confusion discerning between-from within-individual patterns. Here, we tested this hypothesis empirically by investigating the parental care of male dunnocks (Prunella modularis) in relation to paternity. We used a thorough dataset of observations in a wild population, genetic parentage, and a within-subject centring statistical approach to disentangle paternal care adjustment within-male and between males. We found support for the certainty of paternity hypothesis, as there was evidence for within-male adjustment in paternal care when socially monogamous males lost paternity to extra-pair sires. There was little evidence of a between-male effect overall. Our findings show that monogamous males adjust paternal care when paired to the same female partner. We also show that - in monogamous broods - the proportion of provisioning visits made by males yields fitness benefits in terms of fledging success. Our results suggest that socially monogamous females that engage in extra-pair behaviour may suffer fitness costs, as their partners reduction in paternal care can negatively affect fledging success.

animal behavior and cognition

Nanopore-based single molecule sequencing of the D4Z4 array responsible for facioscapulohumeral muscular dystrophy

Subtelomeric macrosatellite repeats are difficult to sequence using conventional sequencing methods owing to the high similarity among repeat units and high GC content. Sequencing these repetitive regions is challenging, even with recent improvements in sequencing technologies. Among these repeats, a haplotype of the telomeric sequence and shortening of the D4Z4 array on human chromosome 4q35 causes one of the most prevalent forms of muscular dystrophy with autosomal-dominant inheritance, facioscapulohumeral muscular dystrophy (FSHD). Here, we applied a nanopore-based ultra-long read sequencer to sequence a BAC clone containing 13 D4Z4 repeats and flanking regions. We successfully obtained the whole D4Z4 repeat sequence, including the pathogenic gene DUX4 in the last D4Z4 repeat. The estimated sequence accuracy of the total repeat region was 99.7% based on a comparison with the reference sequence. Errors were typically observed between purine or between pyrimidine bases. Further, we analyzed the D4Z4 sequence from publicly available ultra-long whole human genome sequencing data obtained by nanopore sequencing. This technology may become a new standard for the molecular diagnosis of FSHD in the future and has the potential to widen our understanding of subtelomeric regions.

bioinformatics

A portable system for metagenomic analyses using nanopore-based sequencer and laptop computers can realize rapid on-site determination of bacterial compositions

We developed a portable system for metagenomic analyses consisting of nanopore technology-based sequencer, MinION, and laptop computers, and assessed its potential ability to determine bacterial compositions rapidly. We tested our protocols using mock bacterial community that contained equimolar 16S rDNA and a pleural effusion from a patient with empyema for time effectiveness and accuracy. MinION sequencing targeting 16S rDNA detected all of 20 bacteria present in the mock bacterial community. Time course analysis indicated that sequence data obtained during the first 5-minute sequencing were enough to detect all 20 bacteria species in the mock sample and determine their compositions with sufficient accuracy. Additionally, using a clinical sample extracted from the pleural effusion of a patient with empyema, we could identify major bacteria in a pleural effusion by rapid sequencing and analysis. All of these results are comparable to or even better than the conventional 16S rDNA sequencing results using IonPGM sequencer. Our results suggest that rapid sequencing and bacterial composition determination is possible within 2 hours.Our integrative system is applicable to rapid diagnostic tests for infectious diseases in near future.

genomics

Coefficient of determination R2 and intra-class correlation coefficient ICC from generalized linear mixed-effects models revisited and expanded

O_LIThe coefficient of determination R2 quantifies the proportion of variance explained by a statistical model and is an important summary statistic of biological interest. However, estimating R2 for (generalized) linear mixed models (GLMMs) remains challenging. We have previously introduced a version of R2 that we called R2GLMM for Poisson and binomial GLMMs, but not for other distributional families.\nC_LIO_LISimilarly, we earlier discussed how to estimate intra-class correlation coefficients ICC (also known as repeatability in the field of ecology and evolution) using Poisson and binomial GLMMs, but not for other distributional families. ICC is related to R2 because they are both ratios of variance components.\nC_LIO_LIIn this article we expand our method to additional non-Gaussian distributions, namely quasi-Poisson, negative binomial and gamma GLMMs. However, in theory, our extension could be applied to any distribution and we include an explanatory calculation for the Tweedie distribution.\nC_LIO_LIWhile expanding our approach, we highlight two useful concepts, Jensens inequality and the delta method, both of which help in understanding the properties of GLMMs. Jensens inequality has important implications for the interpretation GLMMs while the delta method allows a general derivation of distribution-specific variances. We also discuss some special considerations for binomial GLMMs with binary or proportion data.\nC_LIO_LIWe illustrate the implementation of our extension by worked examples in the R environment. However, our method can be used regardless of statistical packages and environments. We finish by referring to two alternative methods to our approach along with a cautionary note.\nC_LI

ecology