bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

A molecular switch for Cdc48 activity and localization during oxidative stress and aging

Control over a healthy proteome begins with the birth of the polypeptide chain and ends with coordinated protein degradation. One of the major players in eukaryotic protein degradation is the essential and highly conserved ATPase, Cdc48 (p97/VCP in mammals). Cdc48 mediates clearance of misfolded proteins from the nucleus, cytosol, ER, mitochondria, and more. Here we dissect the crosstalk between cellular oxidation and Cdc48 activity by identification of a redox-sensitive site, Cys115. By integrating proteomics, biochemistry, microscopy, and bioinformatics, we show that removal of Cys115s redox-sensitive thiol group leads to accumulation of Cdc48 in the nucleus and consequently, results in severe defects in the oxidative stress response, mitochondrial fragmentation, and a decrease in ERAD and sterol biogenesis. We have thus identified a unique redox switch in Cdc48, which may provide a clearer picture of the importance of Cdc48s localization in maintaining a \"healthy\" proteome during oxidative stress and chronological aging in yeast.

molecular biology

Separating the signal from the noise in metagenomic cell-free DNA sequencing

Cell-free DNA (cfDNA) in blood, urine and other biofluids provides a unique window into human health. A proportion of cfDNA is derived from bacteria and viruses, creating opportunities for the diagnosis of infection via metagenomic sequencing. The total biomass of microbial-derived cfDNA in clinical isolates is low, which makes metagenomic cfDNA sequencing susceptible to contamination and alignment noise. Here, we report Low Biomass Background Correction (LBBC), a bioinformatics noise filtering tool informed by the uniformity of the coverage of microbial genomes and the batch variation in the absolute abundance of microbial cfDNA. We demonstrate that LBBC leads to a dramatic reduction in false positive rate while minimally affecting the true positive rate for a cfDNA test to screen for urinary tract infection. We next performed high throughput sequencing of cfDNA in amniotic fluid collected from term uncomplicated pregnancies or those complicated with clinical chorioamnionitis with and without intra-amniotic infection. The data provide unique insight into the properties of fetal and maternal cfDNA in amniotic fluid, demonstrate the utility of cfDNA to screen for intra-amniotic infection, support the view that the amniotic fluid is sterile during normal pregnancy, and reveal cases of intra-amniotic inflammation without infection at term.

genomics

The Metabolic Effects of Angiopoietin-like protein 8 (ANGPLT8) are Differentially Regulated by Insulin and Glucose in Adipose Tissue and Liver and are Controlled by AMPK Signaling

ObjectiveThe angiopoietin-like protein (ANGPTL) family represents a promising therapeutic target for dyslipidemia, which is a feature of obesity and type 2 diabetes (T2DM). The aim of the present study was to determine the metabolic role of ANGPTL8 and to investigate its nutritional, hormonal and molecular regulation in key metabolic tissues.\n\nMethodsThe metabolism of ANGPTL8 knockout mice (ANGPTL8-/-) was examined in mice following chow and high-fat diets (HFD). The regulation of ANGPTL8 expression by insulin and glucose was quantified using a combination of in vivo insulin clamp experiments in mice and in vitro experiments in hepatocytes and adipocytes. The role of AMPK signaling was examined, and the transcriptional control of ANGPTL8 was determined using bioinformatic and luciferase reporter approaches.\n\nResultsThe ANGPTL8-/-mice had improved glucose tolerance and displayed reduced fed and fasted plasma triglycerides. However, there was no reduction in steatosis in ANGPTL8-/-mice after the HFD. Insulin acutely activated ANGPTL8 expression in liver and adipose tissue, which was mediated by C/EBP{beta}. Using insulin clamp experiments we observed that glucose further enhanced ANGPTL8 expression in the presence of insulin in adipocytes only. The activation of AMPK signaling potently suppressed the effect of insulin on ANGPTL8 expression in hepatocytes.\n\nConclusionThese data show that ANGPTL8 plays an important metabolic role in mice that may extend beyond triglyceride metabolism. The finding that insulin and glucose have distinct roles in regulating ANGPTL8 expression in liver and adipose tissue may provide important clues about the function of ANGPTL8 in these tissues.

physiology

Regulation of amino acid and nucleotide metabolism by crustacean hyperglycemic hormone in the muscle and hepatopancreas of the crayfish Procambarus clarkii

To comprehensively characterize the metabolic roles of crustacean hyperglycemic hormone (CHH), metabolites in two CHH target tissues of the crayfish Procambarus clarkii, whose levels were significantly different between CHH-silenced and saline-treated control animals, were analyzed using bioinformatics tools provided by an on-line analysis suite (MetaboAnalyst). Analysis with Metabolic Pathway Analysis (MetPA) indicated that in the muscle Glyoxylate and dicarboxylate metabolism, Nicotinate and nicotinamide metabolism, Alanine, aspartate and glutamate metabolism, Pyruvate metabolism, and Nitrogen metabolism were significantly affected by silencing of CHH gene expression at 24 hours post injection (hpi), while only Nicotinate and nicotinamide metabolism remained significantly affected at 48 hpi. In the hepatopancreas, silencing of CHH gene expression significantly impacted, at 24 hpi, Pyruvate metabolism and Glycolysis or gluconeogenesis, and at 48 hpi, Glycine, serine and threonine metabolism. Moreover, analysis using Metabolite Set Enrichment Analysis (MSEA) showed that many metabolite sets were significantly affected in the muscle at 24hpi, including Ammonia recycling, Nicotinate and nicotinamide metabolism, Pyruvate metabolism, Purine metabolism, Warburg effect, Citric acid cycle, and metabolism of several amino acids, and at 48 hpi only Nicotinate and nicotinamide metabolism, Glycine and serine metabolism, and Ammonia recycling remained significantly affected. In the hepatopancreas, MSEA analysis showed that Fatty acid biosynthesis was significantly impacted at 24 hpi. Finally, in the muscle, levels of several amino acids decreased significantly, while those of 5 other amino acids or related compounds significantly increased in response to CHH gene silencing. Levels of metabolites related to nucleotide metabolism significantly decreased across the board at both time points. In the hepatopancreas, the effects were comparatively minor with only levels of thymine and urea being significantly decreased at 24 hpi. The combined results showed that the metabolic effects of silencing CHH gene expression were far more diverse than suggested by previous studies that emphasized on carbohydrate and energy metabolism. Based on the results, metabolic roles of CHH on the muscle and hepatopancreas were summarized and discussed.

zoology

Screening and identification of MicroRNAs expressed in perirenal adipose tissue during rabbit growth

MiRNAs regulate adipose tissue development, which are closely related to subcutaneous and intramuscular fat deposition and adipocyte differentiation. As an important economic and agricultural animal, rabbits have low adipose tissue deposition and are an ideal model to study adipose regulation. However, the miRNAs related to fat deposition during the growth and development of rabbits are poorly defined. In this study, miRNA-sequencing and bioinformatics analyses were used to profile the miRNAs in rabbit perirenal adipose tissue at 35, 85 and 120 days post-birth. Differentially expressed (DE) miRNAs between different stages were identified by DEseq in R. Target genes of DE miRNAs were predicted by TargetScan and miRanda. To explore the functions of identified miRNAs, Gene Ontology (GO) enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses were performed. Approximately 1.6 GB of data was obtained by miRNA-seq. A total of 987 miRNAs (780 known and 207 newly predicted) and 174 DE miRNAs were identified. The miRNAs ranged from 18nt to 26nt. GO enrichment and KEGG pathway analyses revealed that the target genes of the DE miRNAs were mainly involved in zinc ion binding, regulation of cell growth, MAPK signaling pathway, and other adipose hypertrophy-related pathways. Six DE miRNAs were randomly selected and their expression profiles were validated by q-PCR. In summary, we provide the first report of the miRNA profiles of rabbit adipose tissue during different growth stages. Our data provide a theoretical reference for subsequent studies on rabbit genetics, breeding and the regulatory mechanisms of adipose development.

developmental biology

Population genomic SNPs from epigenetic RADs: gaining genetic and epigenetic data from a single established next-generation sequencing approach

O_LIEpigenetics is increasingly recognised as an important molecular mechanism underlying phenotypic variation. To study DNA methylation in ecological and evolutionary contexts, epiRADseq is a cost-effective next-generation sequencing technique based on reduced representation sequencing of genomic regions surrounding non-/methylated sites. EpiRADseq for genome-wide methylation abundance and ddRADseq for genome-wide SNP genotyping follow very similar library and sequencing protocols, but to date these two types of dataset have been handled separately. Here we test the performance of using epiRADseq data to generate SNPs for population genomic analyses.\nC_LIO_LIWe tested the robustness of using epiRADseq data for population genomics with two independent datasets: a newly generated single-end dataset for the European whitefish Coregonus lavaretus, and a re-analysis of publicly available, previously published paired-end data on corals. Using standard bioinformatic pipelines with a reference genome and without (i.e. de novo catalogue loci), we compared the number of SNPs retained, population genetic summary statistics, and population genetic structure between data drawn from ddRADseq and epiRADseq library preparations.\nC_LIO_LIWe find that SNPs drawn from epiRADseq are similar in number to those drawn from ddRADseq, with a 55-83% of SNPs being identified by both methods. Genotyping error rate was <5% in both approaches. For summary statistics such as heterozygosity and nucleotide diversity, there is a strong correlation between methods (Spearmans rho > 0.88). Furthermore, identical patterns of population genetic structure were recovered using SNPs from epiRADseq and ddRADseq approaches.\nC_LIO_LIWe show that SNPs obtained from epiRADseq are highly similar to those from ddRADseq and are equivalent for estimating genetic diversity and population structure. This finding is particularly relevant to researchers interested in genetics and epigenetics on the same individuals because using a single epigenomic approach to generate two datasets greatly reduces the time and financial costs compared to using these techniques separately. It also efficiently enables correction of epigenetic estimates with population genetic data. Many studies will benefit from a combinatorial approach with genetic and epigenetic markers and this demonstrates a single, efficient method to do so.\nC_LI

genomics

Unravelling the virome in birch: RNA-Seq reveals a complex of known and novel viruses

High-throughput sequencing (HTS), combined with bioinformatics for de novo discovery and assembly of plant virus or viroid genome reads, has promoted the discovery of abundant novel DNA and RNA viruses and viroids. However, the elucidation of a viral population in a single plant is rarely reported. In five birch trees of German and Finnish origin exhibiting symptoms of birch leaf-roll disease (BRLD), we identified in total five viruses, among which three are novel. The number of identified virus variants in each transcriptome ranged from one to five. The novel species are genetically - fully or partially - characterized, they belong to the genera Carlavirus, Idaeovirus and Capillovirus and they are tentatively named birch carlavirus, birch idaeovirus, and birch capillovirus, respectively. The only virus systematically detected by HTS in symptomatic trees affected by the BRLD was the recently discovered birch leafroll-associated virus. The role of the new carlavirus in BLRD etiology seems at best weak, as it was detected only in one of three symptomatic trees. Continuing studies have to clarify the impact of the carlavirus to the BLRD. The role of the Capillovirus and the Idaeovirus within the BLRD complex and whether they influence plant vitality need to be investigated. Our study reveals the viral population in single birch trees and provides a comprehensive overview for the diversities of the viral communities they harbor.

molecular biology

Massive parallel variant characterization identifies NUDT15 alleles associated with thiopurine toxicity

As a prototype of genomics-guided precision medicine, individualized thiopurine dosing based on pharmacogenetics is a highly effective way to mitigate hematopoietic toxicity of this class of drugs. Recently, NUDT15 deficiency was identified as a novel genetic cause of thiopurine toxicity, and NUDT15-informed preemptive dose reduction is quickly adopted in clinical settings. To exhaustively identify pharmacogenetic variants in this gene, we developed massively parallel NUDT15 function assays to determine variants effect on protein abundance and thiopurine cytotoxicity. Of the 3,097 possible missense variants, we characterized the abundance of 2,922 variants and found 54 hotspot residues at which variants resulted in complete loss of protein stability. Analyzing 2,935 variants in the thiopurine cytotoxicity-based assay, we identified 17 additional residues where variants altered NUDT15 activity without affecting protein stability. We identified structural elements key to NUDT15 stability and/or catalytical activity with single amino-acid resolution. Functional effects for NUDT15 variants accurately predicted toxicity risk alleles in 2,398 patients treated with thiopurines, with 100% sensitivity and specificity, in contrast with poor performance of bioinformatic prediction algorithms. In conclusion, our massively parallel variant function assays identified 1,103 deleterious NUDT15 variants, providing a comprehensive reference of variant function and vastly improving the ability to implement pharmacogenetics-guided thiopurine treatment individualization.

genomics

A systemic approach provides insights into the salt stress adaptation mechanisms of contrasting bread wheat genotypes

Bread wheat is one of the most important crops for human diet but the increasing soil salinization is causing yield reductions worldwide. Physiological, genetic, transcriptomics and bioinformatics analyses were integrated to study the salt stress adaptation response in bread wheat. A comparative analysis to uncover the dynamic transcriptomic response of contrasting genotypes from two wheat populations was performed at both osmotic and ionic phases in time points defined by physiologic measurements. The differential stress effect on the expression of photosynthesis, calcium binding and oxidative stress response genes in the contrasting genotypes supported the greater photosynthesis inhibition observed in the susceptible genotype at the osmotic phase. At the ionic phase genes involved in metal ion binding and transporter activity were up-regulated and down-regulated in the tolerant and susceptible genotypes, respectively. The stress effect on mechanisms related with protein synthesis and breakdown was identified at both stress phases. Based on the linkage disequilibrium blocks it was possible to select salt-responsive genes as potential components operating in the salt stress response pathways leading to salt stress resilience specific traits. Therefore, the implementation of a systemic approach provided insights into the adaptation response mechanisms of contrasting bread wheat genotypes at both salt stress phases.\n\nHighlightThe implementation of a systemic approach provided insights into salt stress adaptation response mechanisms of contrasting bread wheat genotypes from two mapping populations at both osmotic and ionic phases.

systems biology

Diffusion of DNA-binding species in the nucleus: A transient anomalous subdiffusion model

Single-particle tracking experiments have measured the distribution of escape times of DNA-binding species diffusing in living cells: CRISPR-Cas9, TetR, and LacI. The observed distribution is a truncated power law. One important property of this distribution is that it is inconsistent with a Gaussian distribution of binding energies. Another is that it leads to transient anomalous subdiffusion, in which diffusion is anomalous at short times and normal at long times, here only mildly anomalous. Monte Carlo simulations are used to characterize the time-dependent diffusion coefficient D(t) in terms of the anomalous exponent , the crossover time t(cross), and the limits D(0) and D({infty}), and to relate these quantities to the escape time distribution. The simplest interpretations identifSubdiffusion of DNA-binding speciesy the escape time as the actual binding time to DNA, or the period of 1D diffusion on DNA in the standard model combining 1D and 3D search, but a more complicated interpretation may be required. The model has several implications for cell biophysics. (a), The initial anomalous regime represents the search of the DNA-binding species for its target DNA sequence. (b), Non-target DNA sites have a significant effect on search kinetics. False positives in bioinformatic searches of the genome are potentially rate-determining in vivo. For simple binding, the search would be speeded if false-positive sequences were eliminated from the genome. (c), Both binding and obstruction affect diffusion. Obstruction ought to be measured directly, using as the primary probe the DNA-binding species with the binding site inactivated, and eGFP as a calibration standard among laboratories and cell types. (d), Overexpression of the DNA-binding species reduces anomalous subdiffusion because the deepest binding sites are occupied and unavailable. (e), The model provides a coarse-grained phenomenological description of diffusion of a DNA-binding species, useful in larger-scale modeling of kinetics, FCS, and FRAP.\n\nSIGNIFICANCEDNA-binding proteins such as transcription factors diffuse in the nucleus until they find their biological target and bind to it. A protein may bind to many false-positive sites before it reaches its target, and the search process is a research topic of considerable interest. Experimental results from the Dahan lab show a truncated power law distribution of escape times at these sites. We show by Monte Carlo simulations that this escape time distribution implies that the protein shows transient anomalous subdiffusion, defined as anomalous subdiffusion at short times and normal diffusion at long times. Implications of the model for experiments, controls, and interpretation of experiments are discussed.

biophysics

Genomic dissection of 43 serum urate-associated loci provides multiple insights into molecular mechanisms of urate control.

Serum urate is the end-product of purine metabolism. Elevated serum urate is causal of gout and a predictor of renal disease, cardiovascular disease and other metabolic conditions. Genome-wide association studies (GWAS) have reported dozens of loci associated with serum urate control, however there has been little progress in understanding the molecular basis of the associated loci. Here we employed trans-ancestral meta-analysis using data from European and East Asian populations to identify ten new loci for serum urate levels. Genome-wide colocalization with cis-expression quantitative trait loci (eQTL) identified a further five new loci. By cis- and trans-eQTL colocalization analysis we identified 24 and 20 genes respectively where the causal eQTL variant has a high likelihood that it is shared with the serum urate-associated locus. One new locus identified was SLC22A9 that encodes organic anion transporter 7 (OAT7). We demonstrate that OAT7 is a very weak urate-butyrate exchanger. Newly implicated genes identified in the eQTL analysis include those encoding proteins that make up the dystrophin complex, a scaffold for signaling proteins and transporters at the cell membrane; MLXIP that, with the previously identified MLXIPL, is a transcription factor that may regulate serum urate via the pentose-phosphate pathway; and MRPS7 and IDH2 that encode proteins necessary for mitochondrial function. Trans-ancestral functional fine-mapping identified six loci (RREB1, INHBC, HLF, UBE2Q2, SFMBT1, HNF4G) with colocalized eQTL that contained putative causal SNPs (posterior probability of causality > 0.8). This systematic analysis of serum urate GWAS loci has identified candidate causal genes at 19 loci and a network of previously unidentified genes likely involved in control of serum urate levels, further illuminating the molecular mechanisms of urate control.\n\nAuthor SummaryHigh serum urate is a prerequisite for gout and a risk factor for metabolic disease. Previous GWAS have identified numerous loci that are associated with serum urate control, however, only a small handful of these loci have known molecular consequences. The majority of loci are within the non-coding regions of the genome and therefore it is difficult to ascertain how these variants might influence serum urate levels without tangible links to gene expression and / or protein function. We have applied a novel bioinformatic pipeline where we combined population-specific GWAS data with gene expression and genome connectivity information to identify putative causal genes for serum urate associated loci. Overall, we identified 15 novel serum urate loci and show that these loci along with previously identified loci are linked to the expression of 44 genes. We show that some of the variants within these loci have strong predicted regulatory function which can be further tested in functional analyses. This study expands on previous GWAS by identifying further loci implicated in serum urate control and new causal mechanisms supported by gene expression changes.

genetics

Genome-wide analysis of GATA factors in moso bamboo (Phyllostachys edulis) unveils that PeGATAs regulate shoot rapid-growth and rhizome development

BackgroundMoso bamboo is well-known for its rapid-growth shoots and widespread rhizomes. However, the regulatory genes of these two processes are largely unexplored. GATA factors regulate many developmental processes, but its role in plant height control and rhizome development remains unclear.\n\nResultsHere, we found that bamboo GATA factors (PeGATAs) are involved in the growth regulation of bamboo shoots and rhizomes. Bioinformatics and evolutionary analysis showed that there are 31 PeGATA factors in bamboo, which can be divided into three subfamilies. Light, hormone, and stress-related cis-elements were found in the promoter region of the PeGATA genes. Gene expression of 12 PeGATA genes was regulated by phytohormone-GA but there was no correlation between auxin and PeGATA gene expression. More than 27 PeGATA genes were differentially expressed in different tissues of rhizomes, and almost all PeGATAs have dynamic gene expression level during the rapid-growth of bamboo shoots. These results indicate that PeGATAs regulate rhizome development and bamboo shoot growth partially via GA signaling pathway. In addition, PeGATA26, a rapid-growth negative regulatory candidate gene modulated by GA treatment, was overexpressed in Arabidopsis, and over-expression of PeGATA26 significantly repressed Arabidopsis primary root length and plant height. The PeGATA26 overexpressing lines were also resistant to exogenous GA treatment, further emphasizing that PeGATA26 inhibits plant height from Arabidopsis to moso bamboo via GA signaling pathway.\n\nConclusionsOur results provide an insight into the function of GATA transcription factors in regulating shoot rapid-growth and rhizome development, and provide genetic resources for engineering plant height.

plant biology

Novel genetic determinants of telomere length from a multi-ethnic analysis of 75,000 whole genome sequences in TOPMed

Telomeres shorten in replicating somatic cells, and telomere length (TL) is associated with age-related diseases 1,2. To date, 17 genome-wide association studies (GWAS) have identified 25 loci for leukocyte TL 3-19, but were limited to European and Asian ancestry individuals and relied on laboratory assays of TL. In this study from the NHLBI Trans-Omics for Precision Medicine (TOPMed) program, we used whole genome sequencing (WGS) of whole blood for variant genotype calling and the bioinformatic estimation of TL in n=109,122 trans-ethnic (European, African, Asian and Hispanic/Latino) individuals. We identified 59 sentinel variants (p-value <5x10-9) from 36 loci (20 novel, 13 replicated in external datasets). There was little evidence of effect heterogeneity across populations, and 10 loci had >1 independent signal. Fine-mapping at OBFC1 indicated the independent signals colocalized with cell-type specific eQTLs for OBFC1 (STN1). We further identified two novel genes, DCLRE1B (SNM1B) and PARN, using a multi-variant gene-based approach.

genetics

Large-scale Genetic Analysis Identifies 66 Novel Loci for Asthma

We carried out a genome-wide association study (GWAS) for asthma in UK Biobank, followed by a meta-analysis with results from the Trans-National Asthma Genetic Consortium (TAGC). 66 novel genomic regions were identified, bringing the number of known asthma susceptibility loci to 211. Significant gene-sex interactions were also observed where susceptibility alleles, either individually or as a function of polygenic risk scores, increased asthma risk to a greater extent in men than women. Bioinformatics analyses demonstrated that asthma-associated variants were enriched for colocalizing to regions of open chromatic in immune cells and identified candidate causal genes at 52 of the novel loci, including CD52. An anti-CD52 (-CD52) antibody mimicked the immune cell-depleting effects of an FDA-approved human -CD52 antibody and reduced allergen-induced airway hyperreactivity in mice. These results further elucidate the genetic architecture of asthma, provide evidence that the immune system plays a prominent role in its pathogenesis, and suggest that CD52 represents a potentially novel therapeutic target for treating asthma.

genetics

Undulating changes in human plasma proteome across lifespan are linked to disease

Aging is the predominant risk factor for numerous chronic diseases that limit healthspan. Mechanisms of aging are thus increasingly recognized as therapeutic targets. Blood from young mice reverses aspects of aging and disease across multiple tissues, pointing to the intriguing possibility that age-related molecular changes in blood can provide novel insight into disease biology. We measured 2,925 plasma proteins from 4,331 young adults to nonagenarians and developed a novel bioinformatics approach which uncovered profound non-linear alterations in the human plasma proteome with age. Waves of changes in the proteome in the fourth, seventh, and eighth decades of life reflected distinct biological pathways, and revealed differential associations with the genome and proteome of age-related diseases and phenotypic traits. This new approach to the study of aging led to the identification of unexpected signatures and pathways of aging and disease and offers potential pathways for aging interventions.

systems biology

Integrated annotations and analyses of small RNA-producing loci from 47 diverse plants

Plant endogenous small RNAs (sRNAs) are important regulators of gene expression. There are two broad categories of plant sRNAs: microRNAs (miRNAs) and endogenous short interfering RNAs (siRNAs). MicroRNA loci are relatively well-annotated but comprise only a small minority of the total sRNA pool; siRNA locus annotations have lagged far behind. Here, we used a large dataset of published and newly generated sRNA sequencing data (1,333 sRNA-seq libraries containing over 20 billion reads) and a uniform bioinformatic pipeline to produce comprehensive sRNA locus annotations of 47 diverse plants, yielding over 2.7 million sRNA loci. The two most numerous classes of siRNA loci produced mainly 24 nucleotide and 21 nucleotide siRNAs, respectively. 24 nucleotide-dominated siRNA loci usually occurred in intergenic regions, especially at the 5-flanking regions of protein-coding genes. In contrast, 21 nucleotide-dominated siRNA loci were most often derived from double-stranded RNA precursors copied from spliced mRNAs. Genic 21 nucleotide-dominated loci were especially common from disease resistance genes, including from a large number of monocots. Individual siRNA sequences of all types showed very little conservation across species, while mature miRNAs were more likely to be conserved. We developed a web server where our data and several search and analysis tools are freely accessible at http://plantsmallrnagenes.science.psu.edu.

plant biology

Cell-type diversity and regionalized gene expression in the planarian intestine revealed by laser-capture microdissection transcriptome profiling

Organ regeneration requires precise coordination of new cell differentiation and remodeling of uninjured tissue to faithfully re-establish organ morphology and function. An atlas of gene expression and cell types in the uninjured state is therefore an essential pre-requisite for understanding how damage is repaired. Here, we use laser-capture microdissection (LCM) and RNA-Seq to define the transcriptome of the intestine of Schmidtea mediterranea, a planarian flatworm with exceptional regenerative capacity. Bioinformatic analysis of 1,844 intestine-enriched transcripts suggests extensive conservation of digestive physiology with other animals, including humans. Comparison of the intestinal transcriptome to purified absorptive intestinal cell (phagocyte) and published single-cell expression profiles confirms the identities of known intestinal cell types, and also identifies hundreds of additional transcripts with previously undetected intestinal enrichment. Furthermore, by assessing the expression patterns of 143 transcripts in situ, we discover unappreciated mediolateral regionalization of gene expression and cell-type diversity, especially among goblet cells. Demonstrating the utility of the intestinal transcriptome, we identify 22 intestine-enriched transcription factors, and find that several have distinct functional roles in the regeneration and maintenance of goblet cells. Furthermore, depletion of goblet cells inhibits planarian feeding and reduces viability. Altogether, our results show that LCM is a viable approach for assessing tissue-specific gene expression in planarians, and provide a new resource for further investigation of digestive tract regeneration, the physiological roles of intestinal cell types, and axial polarity.

developmental biology

Identification and quantification of meat product ingredients by whole-genome metagenomics (All-Food-Seq)

Complex food matrices bear the risk of intentional or accidental admixture of non-declared species. Moreover, declared components can be present in false proportions, since expensive taxa might be exchanged for cheaper ones. We have previously reported that PCR-free metagenomic sequencing of total DNA extracted from sausage samples combined with bioinformatic analysis (termed All-Food-Seq, AFS), can be a valuable screening tool to identify the taxon composition of food ingredients. Here we illustrate this principle by analysing regional Doner kebap samples, which revealed unexpected and unlabelled poultry and plant components in three of five cases. In addition, we systematically apply AFS to a broad set of reference meat material of known composition (i.e. reference sausages) to evaluate quantification accuracy and potential limitations. We include a detailed analysis of the effect of different food matrices and the possibility of false-positive sequence read assignment to closely related species, and we compare AFS quantification results to quantitative real-time PCR (qPCR) and droplet digital PCR (ddPCR). AFS emerges as a potent PCR-free screening tool, which can detect multiple target species of different kingdoms of life within a single assay. Mathematical calibration accounting for pronounced matrix effects can significantly improves AFS quantification accuracy. In comparison, AFS performs better than classical qPCR, and is on par with ddPCR.

genomics