bioRxiv ScienceSearch

Biology subjects

Shi, J.

Publications and source records attributed to Shi, J..

At least 19 recordsLinked to original sources

The N-end rule E3 ligase UBR2 activates Nlrp1b inflammasomes

Innate immunity relies on the formation of different inflammasomes to initiate immune responses. The recognition of diverse infection and other danger signals by innate immune receptors trigger caspase-1 activation that induces pyroptosis. Anthrax lethal factor (LF) is a secreted bacterial protease that known to potently activate Nlrp1b inflammasomes in mouse macrophages, but the molecular mechanism underlying LF-induced Nlrp1b activation remains unknown. We here carried out both a mouse genome-wide siRNA screen and a CRISPR/Cas9 knockout screen seeking to identify genes that participate in Nlrp1b activation triggered by LF treatment. We found that the N-end rule pathway E3 ligase UBR2 is required for Nlrp1b activation and a ubiquitin conjugating E2 enzyme E2O is also involved in this process via its physically interaction with UBR2. We show that LF triggers activation of Nlrp1b by initiating the degradation of the N-terminal fragment of Nlrp1b itself that produced via an auto-cleavage process. This study deepens our understanding of innate immunity defense against bacterial infection by elucidating the functional role of UBR2-mediated N-end rule pathway in LF-induced Nlrp1b activation.

immunology

Genome-scale sequence disruption following biolistic transformation in rice and maize

We biolistically transformed linear 48 kb phage lambda and two different circular plasmids into rice and maize and analyzed the results by whole genome sequencing and optical mapping. While some transgenic events showed simple insertions, others showed extreme genome damage in the form of chromosome truncations, large deletions, partial trisomy, and evidence of chromothripsis and breakage-fusion bridge cycling. Several transgenic events contained megabase-scale arrays of introduced DNA mixed with genomic fragments assembled by non-homologous or microhomology-mediated joining. Damaged regions of the genome, assayed by the presence of small fragments displaced elsewhere, were often repaired without a trace, presumably by homology-dependent repair (HDR). The results suggest a model whereby successful biolistic transformation relies on a combination of end joining to insert foreign DNA and HDR to repair collateral damage caused by the microprojectiles. The differing levels of genome damage observed among transgenic events may reflect the stage of the cell cycle and the availability of templates for HDR.

plant biology

BodyMap transcriptomes reveal unique circular RNA features across tissue types and developmental stages

Circular RNAs (circRNAs) are a novel class of regulatory RNAs. Here, we present a comprehensive investigation of circRNA expression profiles across 11 tissues and 4 developmental stages in rats, along with cross-species analyses in humans and mice. Although positively correlated, circRNAs exhibit higher tissue specificity than cognate mRNAs. Also, genes with higher expression levels exhibit a larger fraction of spliced circular transcripts than their linear counterparts. Intriguingly, while we observed a monotonic increase of circRNA abundance with age in the rat brain, we further discovered a dynamic, age-dependent pattern of circRNA expression in the testes that is characterized by a dramatic increase with advancing stages of sexual maturity and a decrease with aging. The age-sensitive testicular circRNAs are highly associated with spermatogenesis, independent of cognate mRNA expression. The tissue/age implications of circRNAs suggest that they present unique physiological functions rather than simply occurring as occasional by-products of gene transcription.

bioinformatics

The Therapeutic Antibody Profiler (TAP): Five Computational Developability Guidelines

Therapeutic monoclonal antibodies (mAbs) must not only bind to their target but must also be free from 'developability issues', such as poor stability or high levels of aggregation. While small molecule drug discovery benefits from Lipinski's rule of five to guide the selection of molecules with appropriate biophysical properties, there is currently no in silico analog for antibody design. Here, we model the variable domain structures of a large set of post-Phase I clinical-stage antibody therapeutics (CSTs), and calculate an array of metrics to estimate their typical properties. In each case, we contextualize the CST distribution against a snapshot of the human antibody gene repertoire. We describe guideline values for five metrics thought to be implicated in poor developability: the total length of the Complementarity-Determining Regions (CDRs), the extent and magnitude of surface hydrophobicity, positive charge and negative charge in the CDRs, and asymmetry in the net heavy and light chain surface charges. The guideline cut-offs for each property were derived from the values seen in CSTs, and a flagging system is proposed to identify nonconforming candidates. On two mAb drug discovery sets, we were able to selectively highlight sequences with developability issues. We make available the Therapeutic Antibody Profiler (TAP), an open-source computational tool that builds downloadable homology models of variable domain sequences, tests them against our five developability guidelines, and reports potential sequence liabilities and canonical forms. TAP is freely available at http://opig.stats.ox.ac.uk/webapps/sabdab-sabpred/TAP.php.

immunology

SummaryAUC: a tool for evaluating the performance of polygenic risk prediction models in validation datasets with only summary level statistics

MotivationPolygenic risk score (PRS) methods based on genome-wide association studies (GWAS) have a potential for predicting the risk of developing complex diseases and are expected to become more accurate with larger training data sets and innovative statistical methods. The area under the ROC curve (AUC) is often used to evaluate the performance of PRSs, which requires individual genotypic and phenotypic data in an independent GWAS validation dataset. We are motivated to develop methods for approximating AUC of PRSs based on the summary level data of the validation dataset, which will greatly facilitate the development of PRS models for complex diseases.\n\nResultsWe develop statistical methods and an R package SummaryAUC for approximating the AUC and its variance of a PRS when only the summary level data of the validation dataset are available. SummaryAUC can be applied to PRSs with SNPs either genotyped or imputed in the validation dataset. We examined the performance of SummaryAUC using a large-scale GWAS of schizophrenia. SummaryAUC provides accurate approximations to AUCs and their variances. The bias of AUC is typically less than 0.5% in most analyses. SummaryAUC cannot be applied to PRSs that use all SNPs in the genome because it is computationally prohibitive.\n\nAvailabilityhttps://github.com/lsncibb/SummaryAUC\n\nContactJianxin.Shi@nih.gov

bioinformatics

Verification of the phenylpropanoid pinoresinol biosynthetic pathway and its glycosides in Phomopsis sp. XP-8 using 13C stable isotope labeling and liquid chromatography coupled with time-of-flight mass spectrometry

Phomopsis sp. XP-8, an endophytic fungus from the bark of Tu-Chung (EucommiaulmoidesOliv), revealed the pinoresinol diglucoside (PDG) biosynthetic pathway after precursor feeding measurements and genomic annotation. To verify the pathway more accurately, [13C6]-labeled glucose and [13C6]-labeled phenylalanine were separately fed to the strain as sole substrates and [13C6]-labeled products were detected by ultra-high performance liquid chromatography-quantitative time of flight mass spectrometry. As results, [13C6]-labeled phenylalanine was found as [13C6]-cinnamylic acid and p-coumaric acid, and [13C12]-labeled pinoresinol revealed that the pinoresinol benzene ring came from phenylalanine via the phenylpropane pathway. [13C6]-Labeled cinnamylic acid and p-coumaric acid, [13C12]-labeled pinoresinol, [13C18]-labeled pinoresinol monoglucoside (PMG), and [13C18]-labeled PDG products were found when [13C6]-labeled glucose was used, demonstrating that the benzene ring and glucoside of PDG originated from glucose. It was also determined that PMG was not the direct precursor of PDG in the biosynthetic pathway. The study verified the occurrence of the plant-like phenylalanine and lignan biosynthetic pathway in fungi.\n\nImportanceVerify the phenylpropanoid-pinoresinol biosynthetic pathway and its glycosides in an endophytic fungi.

microbiology

CFTR misfolds during native-centric simulations due to entropic penalties of native state formation

Cystic fibrosis (CF) is a common genetic disorder that affects approximately 70,000 people worldwide. It is caused by mutation-induced defects in synthesis, folding, processing, or function of the Cystic Fibrosis Transmembrane conductance Regulator protein (CFTR), a chloride-selective ion channel required for the proper functioning of secretory epithelia in tissues such as the lung, pancreas, and skin. The most common cause of CF is the single-residue deletion of F508 (F508del), a mutation present in one or both alleles in 90% of patients that induces severe folding defects and results in greatly reduced expression of the protein. Despite its medical importance, high-resolution mechanistic information about CFTR folding is lacking. In this study, we used molecular dynamics simulation with a native-centric force field to examine the folding and assembly of both full-length CFTR and the isolated first nucleotide-binding domain (NBD1). We observed that the protein was capable of substantial misfolding on both the intradomain and interdomain scale due to entropically favorable kinetic traps that exist on CFTRs folding free energy surface. These results suggest that even wild type CFTR, in the absence of any disease-related mutations, has suboptimal folding efficiency. We speculate that such entropically-driven misfolding also occurs in disease-prone mutants such as F508del and contributes to the proteins poor in vivo activity.

molecular biology

Seasonal changes of metabolites in phloem sap from Broussonetia papyrifera

Gas chromatography-Mass spectrometry (GC-MS) were employed to analyze the whole metabolites in phloem sap of Broussonetia papyrifera and the seasonal changes of content of these metabolites were also investigated. Thirty-eight metabolites were detected in BP phloem exudates. The highest content (44.59mg g-1) of total metabolites was presented in March. High contents of organic acids and sugars were detected in BP phloem exudates from all growing months. Smaller amounts of fatty acids and alcohols were also detected in BP phloem exudates. Interestingly, some metabolites, such as PI3 kinase inhibitor, Chlorogenic acid, Chelerythrine and palmitic acid, which have properties of bioactivity to anticancer and anti-inflammation, were also detected. Quininic acid was the most abundant organic acid, representing up to 86.3% (average value) of all organic acids. D-fructose, D-glucose, and sucrose were the major soluble sugars in phloem saps and the maximum of sugars content was 19.76mg g-1 (average value) in November. Seasonal changes of contents of metabolites were different among individuals. The metabolites analysis double confirmed that the BP phloem sap can be serviced as an important resource for synthesis of pharmaceutical and human health products.

plant biology

An optimized toolkit for precision base editing

CRISPR base editing is a potentially powerful technology that enables the creation of genetic mutations with single base pair resolution. By re-engineering both DNA and protein sequences, we developed a collection of constitutive and inducible base editing vector systems that dramatically improve the ease and efficiency by which single nucleotide variants can be created. This new toolkit is effective in a wide range of model systems, and provides a means for efficient in vivo somatic base editing.

bioengineering

A multi-scale model of the yeast chromosome-segregation system

In dividing cells, depolymerizing spindle microtubules move chromosomes by pulling at their kinetochores. While kinetochore subcomplexes have been studied extensively in vitro, little is known about their in vivo structure and interactions with microtubules or their response to spindle damage. Here we combine electron cryotomography of serial cryosections with genetic and pharmacological perturbation to study the yeast chromosome-segregation machinery at molecular resolution in vivo. Each kinetochore microtubule has one (rarely, two) Dam1C/DASH outer-kinetochore assemblies.\n\nDam1C/DASH only contacts the flat surface of the microtubule and does so with its flexible \"bridges\". In metaphase, 40% of the Dam1C/DASH assemblies are complete rings; the rest are partial rings. Ring completeness and binding position along the microtubule are sensitive to kinetochore attachment and tension, respectively. Our study supports a model in which each kinetochore must undergo cycles of conformational change to couple microtubule depolymerization to chromosome movement.

cell biology

SPORTS1.0: a tool for annotating and profiling non-coding RNAs optimized for rRNA- and tRNA- derived small RNAs

High-throughput RNA-seq has revolutionized the process of small RNA (sRNA) discovery, leading to a rapid expansion of sRNA categories. In addition to the previously well-characterized sRNAs such as microRNAs (miRNAs), Piwi-interacting RNA (piRNAs), and small nucleolar RNA (snoRNAs), recent emerging studies have spotlighted on tRNA-derived sRNAs (tsRNAs) and rRNA-derived sRNAs (rsRNAs) as new categories of sRNAs that bear versatile functions. Since existing software and pipelines for sRNA annotation are mostly focused on analyzing miRNAs or piRNAs, here we developed the sRNA annotation pipeline optimized for rRNA- and tRNA- derived sRNAs (SPORTS1.0). SPORTS1.0 is optimized for analyzing tsRNAs and rsRNAs from sRNA-seq data, in addition to its capacity to annotate canonical sRNAs such as miRNAs and piRNAs. Moreover, SPORTS1.0 can predict potential RNA modification sites based on nucleotide mismatches within sRNAs. SPORTS1.0 is precompiled to annotate sRNAs for a wide range of 68 species across bacteria, yeast, plant, and animal kingdoms, while additional species for analyses could be readily expanded upon end users input. For demonstration, by analyzing sRNA datasets using SPORTS1.0, we reveal that distinct signatures are present in tsRNAs and rsRNAs from different mouse cell types. We also find that compared to other sRNA species, tsRNAs bear the highest mismatch rate which is consistent with their highly modified nature. SPORTS1.0 is an open-source software and can be publically accessed at https://github.com/junchaoshi/sports1.0.

bioinformatics

A Novel QconCAT-Based Proteomics Method for Determining Allele-Specific Protein Expression (ASPE): a New Approach to Identify Cis-acting Genetic Variants

Measuring allele-specific expression (ASE) is a powerful approach for identifying cis-regulatory genetic variants. Here we developed a novel targeted proteomics method for quantification of allele-specific protein expression (ASPE) based on scheduled high resolution multiple reaction monitoring (sMRM-HR) with a heavy stable isotope-labeled quantitative concatamer (QconCAT) internal protein standard. This strategy was applied to the determination of the ASPE of UGT2B15 in human livers using the common UGT2B15 nonsynonymous variant rs1902023 (i.e. Y85D) as the marker to differentiate expressions from the two alleles. The QconCAT standard contains both the wild type tryptic peptide and the Y85D mutant peptide at a ratio of 1:1 to ensure accurate measurement of the ASPE of UGT2B15. The results from 18 UGT2B15 Y85D heterozygotes revealed that the ratios between wild type Y allele and mutant D allele varied from 0.60 to 1.46, indicating the presence of cis-regulatory variants. In addition, we observed no significant correlations between the ASPE and mRNA ASE of UGT2B15, suggesting the involvement of different cis-acting variants in regulating the transcription and translation processes of the gene. This novel ASPE approach provides a powerful tool for capturing cis-genetic variants involved in post-transcription processes, an important yet understudied area of research.

molecular biology

Quantifying Waddington’s epigenetic landscape: a comparison of single-cell potency measures

Over 60 years ago Waddington proposed an epigenetic landscape model of cellular differentiation, whereby cell-fate transitions are modelled as canalization events, with stable cell states occupying the basins or attractor states1, 2. A key ingredient of this landscape is the energy potential, or height3, which correlates with cell-potency. To date, very few explicit biophysical models for estimating single-cell potency have been proposed. Using 9 independent experiments, encompassing over 6,600 high-quality single-cell RNA-Seq profiles, we here demonstrate that single-cell potency can be approximated as the graph entropy of a Markov Chain process on a model signaling network. Our analysis highlights that other proposed single-cell potency measures are not robust, whilst also revealing that integration with orthogonal systems-level information improves potency estimates. Thus, this study provides a foundation for an improved systems-level understanding of single-cell potency, which may have profound implications for the discovery of novel stem-and progenitor cell phenotypes.

bioinformatics

Generation of a novel growth-enhanced and reduced environmental impact transgenic pig strain

In pig production, insufficient feed digestion causes excessive nutrients such as phosphorus and nitrogen, which are then released to the environment. To address the issue of environmental emissions, we have established transgenic pigs harboring a single-copy quad-cistronic transgene and simultaneously expressing three microbial enzymes, {beta}-glucanase, xylanase, and phytase in the salivary glands. All the transgenic enzymes were successfully expressed, and the digestion of non-starch polysaccharides (NSPs) and phytate in the feedstuff was enhanced. Fecal nitrogen and phosphate outputs were reduced by 23%-46%, and growth rate improved by 23.4% (gilts) and 24.4% (boars) when the pigs were fed on a corn and soybean-based diet and high-NSP diet. The transgenic pigs showed a 11.5%- 14.5% improvement in feed conversion rate compared to the age-matched wild-type littermates. These findings indicate that transgenic pigs are promising resources for improving feed efficiency and reducing nutrient emissions to the environment.

biochemistry

HNF1A is a Novel Oncogene and Central Regulator of Pancreatic Cancer Stem Cells

The biological properties of pancreatic cancer stem cells (PCSCs) remain incompletely defined and the central regulators are unknown. By bioinformatic analysis of a PCSC-enriched gene signature, we identified the transcription factor HNF1A as a putative central regulator of PCSC function. Levels of HNF1A and its target genes were found to be elevated in PCSCs and tumorspheres, and depletion of HNF1A resulted in growth inhibition, apoptosis, impaired tumorsphere formation, PCSC depletion, and downregulation of OCT4 expression. Conversely, HNF1A overexpression increased PCSC numbers and tumorsphere formation in pancreatic cancer cells and drove PDA cell growth. Importantly, depletion of HNF1A in primary tumor xenografts impaired tumor growth and depleted PCSCs in vivo. Finally, we established an HNF1A-dependent gene signature in PDA cells that significantly correlated with reduced survivability in patients. These findings identify HNF1A as a central transcriptional regulator of the PCSC state and novel oncogene in pancreatic ductal adenocarcinoma.

cancer biology

Cell-type specific eQTL of primary melanocytes facilitates identification of melanoma susceptibility genes

Most expression quantitative trait loci (eQTL) studies to date have been performed in heterogeneous tissues as opposed to specific cell types. To better understand the cell-type specific regulatory landscape of human melanocytes, which give rise to melanoma but account for <5% of typical human skin biopsies, we performed an eQTL analysis in primary melanocyte cultures from 106 newborn males. We identified 597,335 cis-eQTL SNPs prior to LD-pruning and 4,997 eGenes (FDR<0.05), which are higher numbers than in any GTEx tissue type with a similar sample size. Melanocyte eQTLs differed considerably from those identified in the 44 GTEx tissues, including skin. Over a third of melanocyte eGenes, including key genes in melanin synthesis pathways, were not observed to be eGenes in two types of GTEx skin tissues or TCGA melanoma samples. The melanocyte dataset also identified cell-type specific trans-eQTLs with a pigmentation-associated SNP for four genes, likely through its cis-regulation of IRF4, encoding a transcription factor implicated in human pigmentation phenotypes. Melanocyte eQTLs are enriched in cis-regulatory signatures found in melanocytes as well as melanoma-associated variants identified through genome-wide association studies (GWAS). Co-localization of melanoma GWAS variants and eQTLs from melanocyte and skin eQTL datasets identified candidate melanoma susceptibility genes for six known GWAS loci including unique genes identified by the melanocyte dataset. Further, a transcriptome-wide association study using published melanoma GWAS data uncovered four new loci, where imputed expression levels of five genes (ZFP90, HEBP1, MSC, CBWD1, and RP11-383H13.1) were associated with melanoma at genome-wide significant P-values. Our data highlight the utility of lineage-specific eQTL resources for annotating GWAS findings and present a robust database for genomic research of melanoma risk and melanocyte biology.

genetics

An Essential Role for Argonaute 2 in EGFR-KRAS Signaling in Pancreatic Cancer Development

KRAS and EGFR are known essential mediators of pancreatic cancer development. In addition, KRAS and EGFR have both been shown to interact with and perturb the function of Argonaute 2 (AGO2), a key regulator of RNA-mediated gene silencing. Here, we employed a genetically engineered mouse model of pancreatic cancer to define the effects of conditional loss of AGO2 in KRASG12D-driven pancreatic cancer. Genetic ablation of AGO2 does not interfere with development of the normal pancreas or KRASG12D-driven early precursor pancreatic intraepithelial neoplasia (PanIN) lesions. Remarkably, however, AGO2 is required for progression from early to late PanIN lesions, development of pancreatic ductal adenocarcinoma (PDAC), and metastasis. AGO2 ablation permits PanIN initiation driven by the EGFR-RAS axis, but rather than progressing to PDAC, these lesions undergo profound oncogene-induced senescence (OIS). Loss of Trp53 (p53) in this model obviates the requirement of AGO2 for PDAC development. In mouse and human pancreatic tissues, increased expression of AGO2 and elevated co-localization with RAS at the plasma membrane is associated with PDAC progression. Furthermore, phosphorylation of AGO2Y393 by EGFR disrupts the interaction of wild-type RAS with AGO2 at the membrane, but does not affect the interaction of mutant KRAS with AGO2. ARS-1620, a G12C-specific inhibitor, disrupts the KRASG12C-AGO2 interaction specifically in pancreatic cancer cells harboring this mutant, demonstrating that the oncogenic KRAS-AGO2 interaction can be pharmacologically targeted. Taken together, our study supports a biphasic model of pancreatic cancer development: an AGO2-independent early phase of PanIN formation reliant on EGFR-RAS signaling, and an AGO2-dependent phase wherein the mutant KRAS-AGO2 interaction is critical to prevent OIS in PanINs and allow progression to PDAC.

cancer biology

Variational Autoencoder: An Unsupervised Model for Modeling and Decoding fMRI Activity in Visual Cortex

Goal-driven convolutional neural networks (CNN) have been shown to be able to predict and decode cortical responses to natural images or videos. Here, we explored an alternative deep neural network, variational auto-encoder (VAE), as a computational model of the visual cortex. We trained a VAE with a five-layer encoder and a five-layer decoder to learn visual representations from a diverse set of unlabeled images. Inspired by the \"free-energy principle\" in neuroscience, we modeled the brains bottom-up and top-down pathways using the VAEs encoder and decoder, respectively. Following such conceptual relationships, we found that the VAE was able to predict cortical activities observed with functional magnetic resonance imaging (fMRI) from three human subjects watching natural videos. Compared to CNN, VAE resulted in relatively lower prediction accuracies, especially for higher-order ventral visual areas. On the other hand, fMRI responses could be decoded to estimate the VAEs latent variables, which in turn could reconstruct the visual input through the VAEs decoder. This decoding strategy was more advantageous than alternative decoding methods based on partial least square regression. This study supports the notion that the brain, at least in part, bears a generative model of the visual world.

neuroscience