bioRxiv ScienceSearch

Biology subjects

Johnson, R.

Publications and source records attributed to Johnson, R..

16 recordsLinked to original sources

Quantitative multi-locus metabarcoding and waggle dance interpretation reveal honey bee spring foraging patterns in Midwest agroecosystems

We explored the pollen foraging behavior of honey bee colonies situated in the corn and soybean dominated agroecosystems of central Ohio over a month-long period using both pollen metabarcoding and waggle dance inference of spatial foraging patterns. For molecular pollen analysis we developed simple and cost-effective laboratory and bioinformatics methods. Targeting four plant barcode loci (ITS2, rbcL, trnL and trnH), we implemented metabarcoding library preparation and dual-indexing protocols designed to minimize amplification biases and index mis-tagging events. We constructed comprehensive, curated reference databases for hierarchical taxonomic classification of metabarcoding data and used these databases to train the Metaxa2 DNA sequence classifier. Comparisons between morphological and molecular palynology provide strong support for the quantitative potential of multi-locus metabarcoding. Results revealed consistent foraging habits between locations and show clear trends in the phenological progression of honey bee spring foraging in these agricultural areas. Our data suggest that three key taxa, woody Rosaceae such as pome fruits and hawthorns, Salix, and Trifolium provided the majority of pollen nutrition during the study. Spatially, these foraging patterns were associated with a significant preference for forests and tree lines relative to crop fields and herbaceous land cover.

ecology

The wheat Sr22, Sr33, Sr35 and Sr45 genes confer resistance against stem rust in barley

In the last 20 years, stem rust caused by the fungus Puccinia graminis f. sp. tritici (Pgt), has re-emerged as a major threat to wheat and barley cultivation in Africa and Europe. In contrast to wheat with 82 designated stem rust (Sr) resistance genes, barleys genetic variation for stem rust resistance is very narrow with only seven resistance genes genetically identified. Of these, only one locus consisting of two genes is effective against Ug99, a strain of Pgt which emerged in Uganda in 1999 and has since spread to much of East Africa and parts of the Middle East. The objective of this study was to assess the functionality, in barley, of cloned wheat Sr genes effective against Ug99. Sr22, Sr33, Sr35 and Sr45 were transformed into barley cv. Golden Promise using Agrobacterium-mediated transformation. All four genes were found to confer effective stem rust resistance. The barley transgenics remained susceptible to the barley leaf rust pathogen Puccinia hordei, indicating that the resistance conferred by these wheat Sr genes was specific for Pgt. Cloned Sr genes from wheat are therefore a potential source of resistance against wheat stem rust in barley.

plant biology

Nearly all new protein-coding predictions in the CHESS database are not protein-coding

In a 2018 paper posted to bioRxiv, Pertea et al. presented the CHESS database, a new catalog of human gene annotations that includes 1,178 new protein-coding predictions. These are based on evidence of transcription in human tissues and homology to earlier annotations in human and other mammals. Here, we reanalyze the evidence used by CHESS, and find that nearly all protein-coding predictions are false positives. We find that 86% overlap transposons marked by RepeatMasker that are known to frequently result in false positive protein-coding predictions. More than half are homologous to only nine Alu-derived primate sequences corresponding to an erroneous and previously withdrawn Pfam protein domain. The entire set shows poor evolutionary conservation and PhyloCSF protein-coding evolutionary signatures indistinguishable from noncoding RNAs, indicating lack of protein-coding constraint. Only four predictions are supported by mass spectrometry evidence, and even those matches are inconclusive. Overall, the new protein-coding predictions are unsupported by any credible experimental or evolutionary evidence of function, result primarily from homology to genes incorrectly classified as protein-coding, and are unlikely to encode functional proteins.

genomics

Preoperative predictions of in-hospital mortality using electronic medical record data

BackgroundPredicting preoperative in-hospital mortality using readily-available electronic medical record (EMR) data can aid clinicians in accurately and rapidly determining surgical risk. While previous work has shown that the American Society of Anesthesiologists (ASA) Physical Status Classification is a useful, though subjective, feature for predicting surgical outcomes, obtaining this classification requires a clinician to review the patients medical records. Our goal here is to create an improved risk score using electronic medical records and demonstrate its utility in predicting in-hospital mortality without requiring clinician-derived ASA scores.\n\nMethodsData from 49,513 surgical patients were used to train logistic regression, random forest, and gradient boosted tree classifiers for predicting in-hospital mortality. The features used are readily available before surgery from EMR databases. A gradient boosted tree regression model was trained to impute the ASA Physical Status Classification, and this new, imputed score was included as an additional feature to preoperatively predict in-hospital post-surgical mortality. The preoperative risk prediction was then used as an input feature to a deep neural network (DNN), along with intraoperative features, to predict postoperative in-hospital mortality risk. Performance was measured using the area under the receiver operating characteristic (ROC) curve (AUC).\n\nResultsWe found that the random forest classifier (AUC 0.921, 95%CI 0.908-0.934) outperforms logistic regression (AUC 0.871, 95%CI 0.841-0.900) and gradient boosted trees (AUC 0.897, 95%CI 0.881-0.912) in predicting in-hospital post-surgical mortality. Using logistic regression, the ASA Physical Status Classification score alone had an AUC of 0.865 (95%CI 0.848-0.882). Adding preoperative features to the ASA Physical Status Classification improved the random forest AUC to 0.929 (95%CI 0.915-0.943). Using only automatically obtained preoperative features with no clinician intervention, we found that the random forest model achieved an AUC of 0.921 (95%CI 0.908-0.934). Integrating the preoperative risk prediction into the DNN for postoperative risk prediction results in an AUC of 0.924 (95%CI 0.905-0.941), and with both a preoperative and postoperative risk score for each patient, we were able to show that the mortality risk changes over time.\n\nConclusionsFeatures easily extracted from EMR data can be used to preoperatively predict the risk of in-hospital post-surgical mortality in a fully automated fashion, with accuracy comparable to models trained on features that require clinical expertise. This preoperative risk score can then be compared to the postoperative risk score to show that the risk changes, and therefore should be monitored longitudinally over time.\n\nAuthor summaryRapid, preoperative identification of those patients at highest risk for medical complications is necessary to ensure that limited infrastructure and human resources are directed towards those most likely to benefit. Existing risk scores either lack specificity at the patient level, or utilize the American Society of Anesthesiologists (ASA) physical status classification, which requires a clinician to review the chart. In this manuscript we report on using machine-learning algorithms, specifically random forest, to create a fully automated score that predicts preoperative in-hospital mortality based solely on structured data available at the time of surgery. This score has a higher AUC than both the ASA physical status score and the Charlson comorbidity score. Additionally, we integrate this score with a previously published postoperative score to demonstrate the extent to which patient risk changes during the perioperative period.

bioinformatics

Spc110 N-Terminal Domains Act Independently to Mediate Stable γ-Tubulin Small Complex Binding and γ-Tubulin Ring Complex Assembly

Microtubule (MT) nucleation in vivo is regulated by the {gamma}-tubulin ring complex ({gamma}TuRC), an approximately 2-megadalton complex conserved from yeast to humans. In Saccharomyces cerevisiae, {gamma}TuRC assembly is a key point of regulation over the MT cytoskeleton. Budding yeast {gamma}TuRC is composed of seven {gamma}-tubulin small complex ({gamma}TuSC) subassemblies which associate helically to form a template from which microtubules grow. This assembly process requires higher-order oligomers of the coiled-coil protein Spc110 to bind multiple {gamma}TuSCs, thereby stabilizing the otherwise low-affinity interface between {gamma}TuSCs. While Spc110 oligomerization is critical, its N-terminal domain (NTD) also plays a role that is poorly understood both functionally and structurally. In this work, we sought a mechanistic understanding of Spc110 NTD using a combination of structural and biochemical analyses. Through crosslinking-mass spectrometry (XL-MS), we determined that a segment of Spc110 coiled-coil is a major point of contact with {gamma}TuSC. We determined the structure of this coiled-coil segment by X-ray crystallography and used it in combination with our XL-MS dataset to generate an integrative structural model of the {gamma}TuSC-Spc110 complex. This structural model, in combination with biochemical analyses of Spc110 heterodimers lacking one NTD, suggests that the two NTDs within an Spc110 dimer act independently, one stabilizing association between Spc110 and {gamma}TuSC and the other stabilizing the interface between adjacent {gamma}TuSCs.

biochemistry

HIF-2α is essential for carotid body development and function

Mammalian adaptation to oxygen flux occurs at many levels, from shifts in cellular metabolism to physiological adaptations facilitated by the sympathetic nervous system and carotid body (CB). Interactions between differing forms of adaptive response to hypoxia, including transcriptional responses orchestrated by the Hypoxia Inducible transcription Factors (HIFs), are complex and clearly synergistic. We show here that there is an absolute developmental requirement for HIF-2a, one of the HIF isoforms, for growth and survival of oxygen sensitive glomus cells of the carotid body. The loss of these cells renders mice incapable of ventilatory responses to hypoxia, and this has striking effects on processes as diverse as arterial pressure regulation, exercise performance, and glucose homeostasis. We show that the expansion of the glomus cells is correlated with mTORC1 activation, and is functionally inhibited by rapamycin treatment. These findings demonstrate the central role played by HIF-2a in carotid body development, growth and function.

bioinformatics

Resistance gene discovery and cloning by sequence capture and association genetics

Genetic resistance is the most economic and environmentally sustainable approach for crop disease protection. Disease resistance (R) genes from wild relatives are a valuable resource for breeding resistant crops. However, introgression of R genes into crops is a lengthy process often associated with co-integration of deleterious linked genes1, 2 and pathogens can rapidly evolve to overcome R genes when deployed singly3. Introducing multiple cloned R genes into crops as a stack would avoid linkage drag and delay emergence of resistance-breaking pathogen races4. However, current R gene cloning methods require segregating or mutant progenies5-10, which are difficult to generate for many wild relatives due to poor agronomic traits. We exploited natural pan-genome variation in a wild diploid wheat by combining association genetics with R gene enrichment sequencing (AgRenSeq) to clone four stem rust resistance genes in <6 months. RenSeq combined with diversity panels is therefore a major advance in isolating R genes for engineering broad-spectrum resistance in crops.

genomics

Novel phosphorylation states of the yeast spindle pole body

Phosphorylation regulates yeast spindle pole body (SPB) duplication and separation and likely regulates microtubule nucleation. We report a phosphoproteomic analysis using tandem mass spectrometry of purified Saccharomyces cerevisiae SPBs for two cell cycle arrests, G1/S and the mitotic checkpoint, expanding on previously reported phosphoproteomic data sets. We present a novel phosphoproteomic state of SPBs arrested in G1/S by a cdc4-1 temperature sensitive mutation, with particular interest in phosphorylation events on the {gamma}-tubulin small complex ({gamma}-TuSC). The cdc4-1 arrest is the earliest arrest at which microtubule nucleation has occurred at the newly duplicated SPB. Several novel phosphorylation sites were identified in G1/S and during mitosis on the microtubule nucleating {gamma}-TuSC. These sites were analyzed in vivo by fluorescence microscopy and were shown to be required for proper regulation of spindle length. Additionally, in vivo analysis of two mitotic sites in Spc97 found that phosphorylation of at least one of these sites is required for progression through the cell cycle. This phosphoproteomic data set not only broadens the scope of the phosphoproteome of SPBs, it also identifies several {gamma}-TuSC phosphorylation sites influencing microtubule regulation.

biochemistry

Mantis: A Fast, Small, and Exact Large-Scale Sequence Search Index

MotivationSequence-level searches on large collections of RNA-seq experiments, such as the NIH Sequence Read Archive (SRA), would enable one to ask many questions about the expression or variation of a given transcript in a population. Bloom filter-based indexes and variants, such as the Sequence Bloom Tree, have been proposed in the past to solve this problem. However, these approaches suffer from fundamental limitations of the Bloom filter, resulting in slow build and query times, less-than-optimal space usage, and large numbers of false positives.\n\nResultsThis paper introduces Mantis, a space-efficient data structure that can be used to index thousands of rawread experiments and facilitate large-scale sequence searches on those experiments. Mantis uses counting quotient filters instead of Bloom filters, enabling rapid index builds and queries, small indexes, and exact results, i.e., no false positives or negatives. Furthermore, Mantis is also a colored de Bruijn graph representation, so it supports fast graph traversal and other topological analyses in addition to large-scale sequence-level searches.\n\nIn our performance evaluation, index construction with Mantis is 4.4x faster and yields a 20% smaller index than the state-of-the-art split sequence Bloom tree (SSBT). For queries, Mantis is 6x -108x faster than SSBT and has no false positives or false negatives. For example, Mantis was able to search for all 200,400 known human transcripts in an index of 2652 human blood, breast, and brain RNA-seq experiments in one hour and 22 minutes; SBT took close to 4 days and AllSomeSBT took about eight hours.\n\nMantis is written in C++11 and is available at https://github.com/splatlab/mantis.

bioinformatics

Ancient exapted transposable elements drive nuclear localisation of lncRNAs

The sequence domains underlying long noncoding RNA (lncRNA) activities, including their characteristic nuclear enrichment, remain largely unknown. It has been proposed that these domains can originate from neofunctionalised fragments of transposable elements (TEs), otherwise known as RIDLs (Repeat Insertion Domains of Long Noncoding RNA), although just a handful have been identified. It is challenging to distinguish functional RIDL instances against a numerous genomic background of neutrally-evolving TEs. We here show evidence that a subset of TE types experience evolutionary selection in the context of lncRNA exons. Together these comprise an enrichment group of 5374 TE fragments in 3566 loci. Their host lncRNAs tend to be functionally validated and associated with disease. This RIDL group was used to explore the relationship between TEs and lncRNA subcellular localisation. Using global localisation data from ten human cell lines, we uncover a dose-dependent relationship between nuclear/cytoplasmic distribution, and evolutionarily-conserved L2b, MIRb and MIRc elements. This is observed in multiple cell types, and is unaffected by confounders of transcript length or expression. Experimental validation with engineered transgenes shows that these TEs drive nuclear enrichment in a natural sequence context. Together these data reveal a role for TEs in regulating the subcellular localisation of lncRNAs.

genomics

A family-based method for leveraging random genetic variation to identify variance-controlling loci

The propensity of a trait to vary within a population may have evolutionary, ecological, or clinical significance. In the present study we deploy sibling models to offer a novel and unbiased way to ascertain loci associated with the extent to which phenotypes vary (variance-controlling quantitative trait loci, or vQTLs). Previous methods for vQTL-mapping either exclude genetically related individuals or treat genetic relatedness among individuals as a complicating factor addressed by adjusting estimates for non-independence in phenotypes. The present method uses genetic relatedness as a tool to obtain unbiased estimates of variance effects rather than as a nuisance. The family-based approach, which utilizes random variation between siblings in minor allele counts at a locus, also allows controls for parental genotype, mean effects, and non-linear (dominance) effects that may spuriously appear to generate variation.\n\nSimulations show that the approach performs equally well as two existing methods (squared Z-score and DGLM) in controlling type I error rates when there is no unobserved confounding, and performs significantly better than these methods in the presence of confounding. Using height and BMI as empirical applications, we investigate SNPs that alter within-family variation in height and BMI, as well as pathways that appear to be enriched. One significant SNP for BMI variability, in the MAST4 gene, replicated. Pathway analysis revealed one gene set, encoding members of several signaling pathways related to gap junction function, which appears significantly enriched for associations with within-family height variation in both datasets (while not enriched in analysis of mean levels). We recommend approximating laboratory random assignment of genotype using family data and more careful attention to the possible conflation of mean and variance effects.

genetics

Unique genomic features and deeply-conserved functions of long non-coding RNAs in the Cancer LncRNA Census (CLC)

Long non-coding RNAs (lncRNAs) that drive tumorigenesis are a growing focus of cancer genomics studies. To facilitate further discovery, we have created the \"Cancer LncRNA Census\" (CLC), a manually-curated and strictly-defined compilation of lncRNAs with causative roles in cancer. CLC has two principle applications: first, as a resource for training and benchmarking de novo identification methods; and second, as a dataset for studying the fundamental properties of these genes.\n\nCLC Version 1 comprises 122 lncRNAs implicated in 29 distinct cancers. LncRNAs are included based on functional or genetic evidence for causative roles in cancer progression. All belong to the GENCODE reference annotation, to enable integration across projects and datasets. For each entry, the evidence type, biological activity (oncogene or tumour suppressor), source reference and cancer type are recorded. Supporting its usefulness, CLC genes are significantly enriched amongst de novo predicted driver genes from PCAWG. CLC genes are distinguished from other lncRNAs by a series of features consistent with biological function, including gene length, high expression and sequence conservation of both exons and promoters. We identify a trend for CLC genes to be co-localised with known protein-coding cancer genes along the human genome. Finally, by integrating data from transposon-mutagenesis functional screens, we show that mouse orthologues of CLC genes tend also to be cancer genes.\n\nThus CLC represents a valuable resource for research into long non-coding RNAs in cancer. Their evolutionary and genomic properties have implications for understanding disease mechanisms and point to conserved functions across ~80 million years of evolution.

bioinformatics

Squeakr: An Exact and Approximate k-mer Counting System

Motivationk-mer-based algorithms have become increasingly popular in the processing of high-throughput sequencing (HTS) data. These algorithms span the gamut of the analysis pipeline from k-mer counting (e.g., for estimating assembly parameters), to error correction, genome and transcriptome assembly, and even transcript quantification. Yet, these tasks often use very different k-mer representations and data structures. In this paper, we set forth the fundamental operations for maintaining multisets of k-mers and classify existing systems from a data-structural perspective. We then show how to build a k-mer-counting and multiset-representation system using the counting quotient filter (CQF), a feature-rich approximate membership query (AMQ) data structure. We introduce the k-mer-counting/querying system Squeakr (Simple Quotient filter-based Exact and Approximate Kmer Representation), which is based on the CQF. This off-the-shelf data structure turns out to be an efficient (approximate or exact) representation for sets or multisets of k-mers.\n\nResultsSqueakr takes 2x-3;4.3x less time than the state-of-the-art to count and perform a random-point-query workload. Squeakr is memory-efficient, consuming 1.5X-4.3X less memory than the state-of-the-art. It offers competitive counting performance, and answers point queries (i.e. queries for the abundance of a particular k-mer) over an order-of-magnitude faster than other systems. The Squeakr representation of the k-mer multiset turns out to be immediately useful for downstream processing (e.g., de Bruijn graph traversal) because it supports fast queries and dynamic k-mer insertion, deletion, and modification.\n\nAvailabilityhttps://github.com/splatlab/squeakr\n\nContact ppandey@cs.stonybrook.edu

bioinformatics

LncATLAS database for subcellular localisation of long noncoding RNAs

BackgroundThe subcellular localisation of long noncoding RNAs (lncRNAs) holds valuable clues to their molecular function. However, measuring localisation of newly-discovered lncRNAs involves time-consuming and costly experimental methods.\n\nResultsWe have created \"LncATLAS\", a comprehensive resource of lncRNA localisation in human cells based on RNA-sequencing datasets. Altogether, 6768 GENCODE-annotated lncRNAs are represented across various compartments of 15 cell lines. We introduce \"Relative concentration index\" (RCI) as a useful measure of localisation derived from ensemble RNAseq measurements. LncATLAS is accessible through an intuitive and informative webserver, from which lncRNAs of interest are accessed using identifiers or names. Localisation is presented across cell types and organelles, and may be compared to the distribution of all other genes. Publication-quality figures and raw data tables are automatically generated with each query, and the entire dataset is also available to download.\n\nConclusionsLncATLAS makes lncRNA subcellular localisation data available to the widest possible number of researchers. It is available at lncATLAS.crg.eu.

bioinformatics

High-throughput annotation of full-length long noncoding RNAs with Capture Long-Read Sequencing (CLS)

Accurate annotations of genes and their transcripts is a foundation of genomics, but no annotation technique presently combines throughput and accuracy. As a result, reference gene collections remain incomplete: many gene models are fragmentary, while thousands more remain uncatalogued-particularly for long noncoding RNAs (lncRNAs). To accelerate lncRNA annotation, the GENCODE consortium has developed RNA Capture Long Seq (CLS), combining targeted RNA capture with third-generation long-read sequencing. We present an experimental re-annotation of the GENCODE intergenic lncRNA population in matched human and mouse tissues, resulting in novel transcript models for 3574 / 561 gene loci, respectively. CLS approximately doubles the annotated complexity of targeted loci, outperforming existing short-read techniques. Full-length transcript models produced by CLS enable us to definitively characterize the genomic features of lncRNAs, including promoter- and gene-structure, and protein-coding potential. Thus CLS removes a longstanding bottleneck of transcriptome annotation, generating manual-quality full-length transcript models at high-throughput scales.\n\nAbbreviations

genomics

Nus factors have a widespread regulatory function in bacteria

Nus factors are broadly conserved across bacterial species, and are often essential for viability. A complex of five Nus factors (NusB, NusE, NusA, NusG and SuhB) is considered to be a dedicated regulator of ribosomal RNA folding, and has been shown to prevent Rho-dependent transcription termination. We have established the first cellular function for the Nus factor complex beyond regulation of ribosomal assembly: repression of the Nus factor-encoding gene, suhB. This repression occurs by translation inhibition followed by Rho-dependent transcription termination. Thus, the Nus factor complex can prevent or promote Rho activity depending on the gene context. Extensive conservation of NusB/E binding sites upstream of nus factor genes suggests that Nus factor autoregulation occurs in many species. Putative NusB/E binding sites are also found upstream of many other genes in diverse species, and we demonstrate Nus factor regulation of one such gene in Citrobacter koseri. We conclude that Nus factors have an evolutionarily widespread regulatory function beyond ribosomal RNA, and that they are often autoregulatory.

microbiology