bioRxiv ScienceSearch

Biology subjects

Zou, J.

Publications and source records attributed to Zou, J..

15 recordsLinked to original sources

A genetically encoded fluorescent sensor for rapid and specific in vivo detection ofnorepinephrine

Norepinephrine (NE) and epinephrine (Epi), two key biogenic monoamine neurotransmitters, are involved in a wide range of physiological processes. However, their precise dynamics and regulation remain poorly characterized, in part due to limitations of available techniques for measuring these molecules in vivo. Here, we developed a family of GPCR Activation-Based NE/Epi (GRABNE) sensors with a 230% peak {Delta}F/F0 response to NE, good photostability, nanomolar-to-micromolar sensitivities, sub-second rapid kinetics, high specificity to NE vs. dopamine. Viral- or transgenic- mediated expression of GRABNE sensors were able to detect electrical-stimulation evoked NE release in the locus coeruleus (LC) of mouse brain slices, looming-evoked NE release in the midbrain of live zebrafish, as well as optogenetically and behaviorally triggered NE release in the LC and hypothalamus of freely moving mice. Thus, GRABNE sensors are a robust tool for rapid and specific monitoring of in vivo NE/Epi transmission in both physiological and pathological processes.

neuroscience

Genome-wide selection footprints and deleterious variations in young Asian allotetraploid rapeseed

Brassica napus (AACC, 2n=38), is an important oilseed crop grown worldwide. However, little is known about the population evolution of this species, the genomic difference between its major genetic clusters, such as European and Asian rapeseed, and impacts of historical large-sale introgression events in this young tetraploid. In this study, we reported the de novo assembly of the genome sequences of an Asian rapeseed (B. napus), Ningyou 7 and its four progenitors and carried out de novo assembly-based comparison, pedigree and population analysis with other available genomic data from diverse European and Asian cultivars. Our results showed that Asian rapeseed originally derived from European rapeseed, but it had subsequently significantly diverged, with rapid genome differentiation after intensive local breeding selection. The first historical introgression of B. rapa dramatically broadened the allelic pool of Asian B. napus, but decreased their deleterious variations. The secondary historical introgression of European rapeseed (canola-quality) has reshaped Asian rapeseed into two groups, accompanied by an increase in genetic load. This study demonstrates distinctive genomic footprints by recent intra- and inter-species introgression events for local adaptation, and provide novel insights for understanding the rapid genome evolution of a young allopolyploid crop.

genomics

Systematic characterization of genome editing in primary T cells reveals proximal genomic insertions and enables machine learning prediction of CRISPR-Cas9 DNA repair outcomes

The Streptococcus pyogenes Cas9 (SpCas9) nuclease has become a ubiquitous genome editing tool due to its ability to target almost any location in DNA and create a double-stranded break1,2. After DNA cleavage, the break is fixed with endogenous DNA repair machinery, either by non-templated mechanisms (e.g. non-homologous end joining (NHEJ) or microhomology-mediated end joining (MMEJ)), or homology directed repair (HDR) using a complementary template sequence3,4. Previous work has shown that the distribution of repair outcomes within a cell population is non-random and dependent on the targeted sequence, and only recent efforts have begun to investigate this further5-11. However, no systematic work to date has been validated in primary human cells5,7. Here, we report DNA repair outcomes from 1,521 unique genomic locations edited with SpCas9 ribonucleoprotein complexes (RNPs) in primary human CD4+ T cells isolated from multiple healthy blood donors. We used targeted deep sequencing to measure the frequency distribution of repair outcomes for each guide RNA and discovered distinct features that drive individual repair outcomes after SpCas9 cleavage. Predictive features were combined into a new machine learning model, CRISPR Repair OUTcome (SPROUT), that predicts the length and probability of nucleotide insertions and deletions with R2 greater than 0.5. Surprisingly, we also observed large insertions at more than 90% of targeted loci, albeit at a low frequency. The inserted sequences aligned to diverse regions in the genome, and are enriched for sequences that are physically proximal to the break site due to chromatin interactions. This suggests a new mechanism where sequences from three-dimensionally neighboring regions of the genome can be inserted during DNA repair after Cas9-induced DNA breaks. Together, these findings provide powerful new predictive tools for Cas9-dependent genome editing and reveal new outcomes that can result from genome editing in primary T cells.

cell biology

eTumorRisk, an algorithm predicts cancer risk based on comutated gene networks in an individual’s germline genome

Early cancer detection has potentials to reduce cancer burden. A prior identification of the high-risk population of cancer will facilitate cancer early detection. Traditionally, cancer predisposition genes such as BRCA1/2 have been used for identifying high-risk population of developing breast and ovarian cancers. However, such high-risk genes have only a few. Moreover, the complexity of cancer hints multiple genes involved but also prevents from identifying such predictors for predicting high-risk subpopulation. Therefore, we asked if the germline genomes could be used to identify high-risk cancer population. So far, none of such predictive models has been developed. Here, by analyzing of the germline genomes of 3,090 cancer patients representing 12 common cancer types and 25,701 non-cancer individuals, we discovered significantly differential co-mutated gene pairs between cancer and non-cancer groups, and even between cancer types. Based on these findings, we developed a network-based algorithm, eTumorRisk, which enables to predict individuals cancer risk of six genetic-dominant cancers including breast, colon, brain, leukemia, ovarian and endometrial cancers with the prediction accuracies of 74.1-91.7% and have 1-3 false-negatives out of the validating samples (n=14,701). The eTumorRisk which has a very low false-negative rate might be useful in screening of general population for identifying high-risk cancer population.

bioinformatics

Modeling Spatial Correlation of Transcripts With Application to Developing Pancreas

Recently high-throughput image-based transcriptomic methods were developed and enabled researchers to spatially resolve gene expression variation at the molecular level for the first time. In this work, we develop a general analysis tool to quantitatively study the spatial correlations of gene expression in fixed tissue sections. As an illustration, we analyze the spatial distribution of single mRNA molecules measured by in situ sequencing on human fetal pancreas at three developmental time points 80, 87 and 117 days post-fertilization. We develop a density profile-based method to capture the spatial relationship between gene expression and other morphological features of the tissue sample such as position of nuclei and endocrine cells of the pancreas. In addition, we build a statistical model to characterize correlations in the spatial distribution of the expression level among different genes. This model enables us to infer the inhibitory and clustering effects throughout different time points. Our analysis framework is applicable to a wide variety of spatially-resolved transcriptomic data to derive biological insights.

bioinformatics

Auditory and Language Contributions to Neural Encoding of Speech Features in Noisy Environments

Recognizing speech in noisy environments is a challenging task that involves both auditory and language mechanisms. Previous studies have demonstrated noise-robust neural tracking of the speech envelope, i.e., fluctuations in sound intensity, in human auditory cortex, which provides a plausible neural basis for noise-robust speech recognition. The current study aims at teasing apart auditory and language contributions to noise-robust envelope tracking by comparing 2 groups of listeners, i.e., native listeners of the testing language and foreign listeners who do not understand the testing language. In the experiment, speech is mixed with spectrally matched stationary noise at 4 intensity levels and the neural responses are recorded using electroencephalography (EEG). When the noise intensity increases, an increase in neural response gain is observed for both groups of listeners, demonstrating auditory gain control mechanisms. Language comprehension creates no overall boost in the response gain or the envelope-tracking precision but instead modulates the spatial and temporal profiles of envelope-tracking activity. Based on the spatio-temporal dynamics of envelope-tracking activity, the 2 groups of listeners and the 4 levels of noise intensity can be jointly decoded by a linear classifier. All together, the results show that without feedback from language processing, auditory mechanisms such as gain control can lead to a noise-robust speech representation. High-level language processing, however, further modulates the spatial-temporal profiles of the neural representation of the speech envelope.

neuroscience

Reliable Multiplex Sequencing with Rare Index Mis-Assignment on DNB-Based NGS Platform

BackgroundMassively-parallel-sequencing, coupled with sample multiplexing, has made genetic tests broadly affordable. However, intractable index mis-assignments (commonly exceeds 1%) were repeatedly reported on some widely used sequencing platforms.\n\nResultsHere, we investigated this quality issue on BGI sequencers using three library preparation methods: whole genome sequencing (WGS) with PCR, PCR-free WGS, and two-step targeted PCR. BGIs sequencers utilize a unique DNB technology which uses rolling circle replication for DNA-nanoball preparation; this linear amplification is PCR free and can avoid error accumulation. We demonstrated that single index mis-assignment from free indexed oligos occurs at a rate of one in 36 million reads, suggesting virtually no index hopping during DNB creation and arraying. Furthermore, the DNB-based NGS libraries have achieved an unprecedentedly low sample-to-sample mis-assignment rate of 0.0001% to 0.0004% under recommended procedures.\n\nConclusionsSingle indexing with DNB technology provides a simple but effective method for sensitive genetic assays with large sample numbers.

genomics

Hepatitis C virus NS5A inhibitor daclatasvir allosterically impairs NS4B-involved protein-protein interactions within the viral replicase and disrupts the replicase quaternary structure in a replicase assembly surrogate system

Daclatasvir (DCV) is a highly potent direct-acting antiviral that targets the non-structural protein 5A (NS5A) of hepatitis C virus (HCV) and has achieved great clinical successes. Previous studies demonstrate its impact on the viral replication complex assembly. However the precise mechanism by which DCV impairs the replication complex assembly remains elusive. In this study, by using HCV subgenomic replicons and a viral replicase assembly surrogate system that expresses the HCV NS3-5B polyprotein to mimic the viral replicase assembly, we dissected the impacts of DCV on aggregation and tertiary structure of NS5A, the protein-protein interactions within the viral replicase and the quaternary structure of the viral replicase. We found that DCV didnt affect aggregation and tertiary structure of NS5A. DCV induced a quaternary structural change of the viral replicase, evidenced by selectively increasing of the NS4Bs sensitivity to proteinase K digestion. Mechanically, DCV impaired the NS4B-involved protein-protein interactions within the viral replicase. The DCV-resistant mutant Y93H was refractory to the DCV-induced reduction of the NS4B-invoved protein interactions and the quaternary structural change of the viral replicase. In addition, Y93H reduced NS4B-involed protein-protein interactions within the viral replicase and attenuated viral replication. We propose that DCV may induce a position change of NS5A, which allosterically affects the protein interactions within the replicase components and disrupts the replicase assembly.\n\nImportanceThe development of the direct-acting antivirals (DAA) has resulted in great clinical achievements for Hepatitis C Virus (HCV) treatment. Daclatasvir (DCV) is an inhibitor targeting the non-enzymatic NS5A, with the 50% effective concentration values in the picomolar range. Accumulated data suggest that DCV blocks the biogenesis of the HCV replication complex. However the mechanistic actions of DCV are still largely unknown. Insights into the action mechanism of DCV on the viral replication complex assembly of HCV may enlighten the development of next generation of DAAs and new anti-viral strategies for other positive-strand RNA viruses for which there are a scarcity of DAAs. Herein, using HCV subgenomic replicons and a viral replicase assembly surrogate system, we dissected the mechanistic actions of DCV on the viral replicase assembly. We found that DCV allosterically impairs NS4B-involved protein-protein interactions within the viral replicase and disrupts the quaternary structure of the viral replicase.

microbiology

The Clinical Imperative for Inclusivity: Race, Ethnicity, and Ancestry (REA) in Genomics

The Clinical Genome Resource (ClinGen) Ancestry and Diversity Working Group highlights the need to develop guidance on race, ethnicity, and ancestry (REA) data collection and use in clinical genomics. We present quantitative and qualitative evidence to characterize: 1) acquisition of REA data via clinical laboratory requisition forms, and 2) information disparity across populations in the Genome Aggregation Database (gnomAD) at clinically relevant sites as determined by variants in ClinVar. Our requisition form analysis showed substantial heterogeneity in clinical laboratory ascertainment of REA, as well as marked incongruity among terms used to define REA categories. There was also striking disparity across REA populations in the amount of information available about variants at clinically relevant sites in gnomAD. European ancestral populations constituted the majority of observations (55.8%), allele counts (59.7%), and private alleles (56.1%) in gnomAD at 550 loci with \"pathogenic\" and \"likely pathogenic\" expert-reviewed variants in ClinVar. Our findings highlight the importance of implementing and supporting programs to increase diversity in genome sequencing and clinical genomics, as well as measuring uncertainty around population-level datasets that are used in variant interpretation. Finally, we suggest the need for a standardized REA data collection framework to be developed and adopted across clinical genomics.

genomics

Germline genomic landscapes of breast cancer patients significantly predict clinical outcomes

Germline genetic variants such as BRCA1/2 play an important role in tumorigenesis and clinical outcomes of cancer patients. However, only a small fraction (i.e., 5-10%) of inherited variants has been associated with clinical outcomes (e.g., BRCA1/2, APC, TP53, PTEN and so on). The challenge remains in using these inherited germline variants to predict clinical outcomes of cancer patient population. In an attempt to solve this issue, we applied our recently developed algorithm, eTumorMetastasis, which constructs predictive models, on exome sequencing data to ER+ breast (n=755) cancer patients. Gene signatures derived from the genes containing functionally germline genetic variants significantly distinguished recurred and non-recurred patients in two ER+ breast cancer independent cohorts (n=200 and 295, P=1.4x10-3). Furthermore, we found that recurred patients possessed a higher rate of germline genetic variants. In addition, the inherited germline variants from these gene signatures were predominately enriched in T cell function, antigen presentation and cytokine interactions, likely impairing the adaptive and innate immune response thus favoring a pro-tumorigenic environment. Hence, germline genomic information could be used for developing non-invasive genomic tests for predicting patients outcomes (or drug response) in breast cancer, other cancer types and even other complex diseases.

cancer biology

eTumorMetastasis, a network-based algorithm predicts clinical outcomes using whole-exome sequencing data of cancer patients

Continual reduction in sequencing cost is expanding the accessibility of genome sequencing data for routine clinical applications. However, the lack of methods to construct machine learning-based predictive models using these datasets has become a crucial bottleneck for the application of sequencing technology in clinics. Here we developed a new algorithm, eTumorMetastasis, which transforms tumor functional mutations into network-based profiles, and identify network operational gene signatures (NOG signatures) which model the tipping point at which a tumor cell shifts from a state that doesnt favor recurrences to one that does. We showed that NOG signatures derived from genomic mutations of tumor founding clones (i.e., the most recent common ancestor of the cells within a tumor) significantly distinguished recurred and non-recurred breast tumors. These results imply that somatic mutations of tumor founders are association with tumor recurrence and can be used to predict clinical outcomes. Finally, the concepts underlying the eTumorMetastasis pave the way for the application of genome sequencing in predictions for other complex genetic diseases.

bioinformatics

The proteome of the malaria plastid organelle, a key anti-parasitic target

Malaria parasites (Plasmodium spp.) and related apicomplexan pathogens contain a non-photosynthetic plastid called the apicoplast. Derived from an unusual secondary eukaryote-eukaryote endosymbiosis, the apicoplast is a fascinating organelle whose function and biogenesis rely on a complex amalgamation of bacterial and algal pathways. Because these pathways are distinct from the human host, the apicoplast is an excellent source of novel antimalarial targets. Despite its biomedical importance and evolutionary significance, the absence of a reliable apicoplast proteome has limited most studies to the handful of pathways identified by homology to bacteria or primary chloroplasts, precluding our ability to study the most novel apicoplast pathways. Here we combine proximity biotinylation-based proteomics (BioID) and a new machine learning algorithm to generate a high-confidence apicoplast proteome consisting of 346 proteins. Critically, the high accuracy of this proteome significantly outperforms previous prediction-based methods and extends beyond other BioID studies of unique parasite compartments. Half of identified proteins have unknown function, and 77% are predicted to be important for normal blood-stage growth. We validate the apicoplast localization of a subset of novel proteins and show that an ATP-binding cassette protein ABCF1 is essential for blood-stage survival and plays a previously unknown role in apicoplast biogenesis. These findings indicate critical organellar functions for newly discovered apicoplast proteins. The apicoplast proteome will be an important resource for elucidating unique pathways derived from secondary endosymbiosis and prioritizing antimalarial drug targets.

microbiology

Leveraging allele-specific expression to refine fine-mapping for eQTL studies

Many disease risk loci identified in genome-wide association studies are present in non-coding regions of the genome. It is hypothesized that these variants affect complex traits by acting as expression quantitative trait loci (eQTLs) that influence expression of nearby genes. This indicates that many causal variants for complex traits are likely to be causal variants for gene expression. Hence, identifying causal variants for gene expression is important for elucidating the genetic basis of not only gene expression but also complex traits. However, detecting causal variants is challenging due to complex genetic correlation among variants known as linkage disequilibrium (LD) and the presence of multiple causal variants within a locus. Although several fine-mapping approaches have been developed to overcome these challenges, they may produce large sets of putative causal variants when true causal variants are in high LD with many non-causal variants. In eQTL studies, there is an additional source of information that can be used to improve fine-mapping called allele-specific expression (ASE) that measures imbalance in gene expression due to different alleles. In this work, we develop a novel statistical method that leverages both ASE and eQTL information to detect causal variants that regulate gene expression. We illustrate through simulations and application to the Genotype-Tissue Expression (GTEx) dataset that our method identifies the true causal variants with higher specificity than an approach that uses only eQTL information. In the GTEx dataset, our method achieves the median reduction rate of 11% in the number of putative causal variants.\n\nContactJaeHoonSul@mednet.ucla.edu, eeskin@cs.ucla.edu

genetics

Interactome Analysis Reveals Regulator of G Protein Signaling 14 (RGS14) is a Novel Calmodulin (CaM) Effector in Mouse Brain

Regulator of G Protein Signaling 14 (RGS14) is a complex scaffolding protein with an unusual domain structure that allows it to integrate G protein and mitogen-activated protein kinase (MAPK) signaling pathways. RGS14 mRNA and protein are enriched in brain tissue of rodents and primates. In the adult mouse brain, RGS14 is predominantly expressed in postsynaptic dendrites and spines of hippocampal CA2 pyramidal neurons where it naturally inhibits synaptic plasticity and hippocampus-dependent learning and memory. However, the signaling proteins that RGS14 natively interacts with in neurons to regulate plasticity are unknown. Here, we show that RGS14 exists as a component of a high molecular weight protein complex in brain. To identify RGS14 neuronal interacting partners, endogenous RGS14 immunoprecipitated from mouse brain was subjected to mass spectrometry and proteomic analysis. We find that RGS14 interacts with key postsynaptic proteins that regulate neuronal plasticity. Gene ontology analysis reveals that the most enriched RGS14 interacting proteins have functional roles in actin-binding, calmodulin(CaM)-binding, and CaM-dependent protein kinase (CaMK) activity. We validate these proteomics findings using biochemical assays that identify interactions between RGS14 and two previously unknown binding partners: CaM and CaMKII. We report that RGS14 directly interacts with CaM in a calcium-dependent manner and is phosphorylated by CaMKII in vitro. Lastly, we detect that RGS14 associates with CaMKII and with CaM in hippocampal CA2 neurons by proximity ligation assays in mouse brain sections. Taken together, these findings demonstrate that RGS14 is a novel CaM effector and CaMKII phosphorylation substrate thereby providing new insight into cellular mechanisms by which RGS14 controls plasticity in CA2 neurons.

biochemistry

Architecture of TAF11/TAF13/TBP complex suggests novel regulatory state in General Transcription Factor TFIID function

General transcription factor TFIID is a key component of RNA polymerase II transcription initiation. Human TFIID is a megadalton-sized complex comprising TATA-binding protein (TBP) and 13 TBP-associated factors (TAFs). TBP binds to core promoter DNA, recognizing the TATA-box. We identified a ternary complex formed by TBP and the histone fold (HF) domain-containing TFIID subunits TAF11 and TAF13. We demonstrate that TAF11/TAF13 competes for TBP binding with TATA-box DNA, and also with the N-terminal domain of TAF1 previously implicated in TATA-box mimicry. In an integrative approach combining crystal coordinates, biochemical analyses and data from cross-linking mass-spectrometry (CLMS), we determine the architecture of the TAF11/TAF13/TBP complex, revealing TAF11/TAF13 interaction with the DNA binding surface of TBP. We identify a highly conserved C-terminal TBP-binding domain (CTID) in TAF13 which is essential for supporting cell growth. Our results thus have implications for cellular TFIID assembly and suggest a novel regulatory state for TFIID function.

biochemistry