bioRxiv ScienceSearch

Biology subjects

Zhu, X.

Publications and source records attributed to Zhu, X..

At least 19 recordsLinked to original sources

Genome-wide association analysis of excessive daytime sleepiness identifies 42 loci that suggest phenotypic subgroups

Excessive daytime sleepiness (EDS) affects 10-20% of the population and is associated with substantial functional deficits. We identified 42 loci for self-reported EDS in GWAS of 452,071 individuals from the UK Biobank, with enrichment for genes expressed in brain tissues and in neuronal transmission pathways. We confirmed the aggregate effect of a genetic risk score of 42 SNPs on EDS in independent Scandinavian cohorts and on other sleep disorders (restless leg syndrome, insomnia) and sleep traits (duration, chronotype, accelerometer-derived sleep efficiency and daytime naps or inactivity). Strong genetic correlations were also seen with obesity, coronary heart disease, psychiatric diseases, cognitive traits and reproductive ageing. EDS variants clustered into two predominant composite phenotypes - sleep propensity and sleep fragmentation - with the former showing stronger evidence for enriched expression in central nervous system tissues, suggesting two unique mechanistic pathways. Mendelian randomization analysis indicated that higher BMI is causally associated with EDS risk, but EDS does not appear to causally influence BMI.

genomics

Adding function to the genome of African Salmonella ST313

Salmonella Typhimurium ST313 causes invasive nontyphoidal Salmonella (iNTS) disease in sub-Saharan Africa, targeting susceptible HIV+, malarial or malnourished individuals. An in-depth genomic comparison between the ST313 isolate D23580, and the well-characterized ST19 isolate 4/74 that causes gastroenteritis across the globe, revealed extensive synteny. To understand how the 856 nucleotide variations generated phenotypic differences, we devised a large-scale experimental approach that involved the global gene expression analysis of strains D23580 and 4/74 grown in sixteen infection-relevant growth conditions. Comparison of transcriptional patterns identified virulence and metabolic genes that were differentially expressed between D23580 versus 4/74, many of which were validated by proteomics. We also uncovered the S. Typhimurium D23580 and 4/74 genes that showed expression differences during infection of murine macrophages. Our comparative transcriptomic data are presented in a new enhanced version of the Salmonella expression compendium SalComD23580: bioinf.gen.tcd.ie/cgi-bin/salcom_v2.pl. We discovered that the ablation of melibiose utilization was caused by 3 independent SNP mutations in D23580 that are shared across ST313 lineage 2, suggesting that the ability to catabolise this carbon source has been negatively selected during ST313 evolution. The data revealed a novel plasmid maintenance system involving a plasmid-encoded CysS cysteinyl-tRNA synthetase, highlighting the power of large-scale comparative multi-condition analyses to pinpoint key phenotypic differences between bacterial pathovariants.

microbiology

Epigenome-wide association analysis of daytime sleepiness in the Multi-Ethnic Study of Atherosclerosis reveals African-American specific associations

Study ObjectivesExcessive daytime sleepiness (EDS) is a consequence of inadequate sleep, or of a primary disorder of sleep-wake control. Population variability in prevalence of EDS and susceptibility to EDS are likely due to genetic and biological factors as well as social and environmental influences. Epigenetic modifications (such as DNA methylation-DNAm) are potential influences on a range of health outcomes. Here, we explored the association between DNAm and daytime sleepiness quantified by the Epworth Sleepiness Scale (ESS).\n\nMethodsWe performed multi-ethnic and ethnic-specific epigenome-wide association studies for DNAm and ESS in 619 individuals from the Multi-Ethnic Study of Atherosclerosis. Replication was assessed in the Cardiovascular Health Study (CHS). Genetic variants in genes proximal to ESS-associated DNAm were analyzed to identify methylation quantitative trait loci and followed with replication of genotype-sleepiness associations in the UK Biobank.\n\nResults61 methylation sites were associated with ESS (FDR [≤] 0.1) in African Americans only, including an association in KCTD5, a gene strongly implicated in sleep. One association (cg26130090) replicated in CHS African Americans (p-value 0.0004). We identified a sleepiness-associated methylation site in the gene RAI1, a gene associated with sleep and circadian phenotypes. In a follow-up analysis, a genetic variant within RAI1 associated with both DNAm and sleepiness score. The variants association with sleepiness was replicated in the UK Biobank.\n\nConclusionsOur analysis identified methylation sites in multiple genes that may be implicated in EDS. These sleepiness-methylation associations were specific to African Americans. Future work is needed to identify mechanisms driving ancestry-specific methylation effects.\n\nStatement of SignificanceExcessive daytime sleepiness is associated with negative health outcomes such as reduction in quality of life, increased workplace accidents, and cardiovascular mortality. There are race/ethnic disparities in excessive daytime sleepiness, however, the environmental and biological mechanisms for these differences are not yet understood. We performed an association analysis of DNA methylation, measured in monocytes, and daytime sleepiness within a racially diverse study population. We detected numerous DNA methylation markers associated with daytime sleepiness in African Americans, but not in European and Hispanic Americans. Future work is required to elucidate the pathways between DNA methylation, sleepiness, and related behavioral/environmental exposures.

genomics

Microbial contamination screening and interpretation for biological laboratory environments

Advances in microbiome researches have led us to the realization that the composition of microbial communities of indoor environment is profoundly affected by the function of buildings, and in turn may bring detrimental effects to the indoor environment and the occupants. Thus investigation is warranted for a deeper understanding of the potential impact of the indoor microbial communities. Among these environments, the biological laboratories stand out because they are relatively clean and yet are highly susceptible to microbial contaminants. In this study, we assessed the microbial compositions of samples from the surfaces of various sites across different types of biological laboratories. We have qualitatively and quantitatively assessed these possible microbial contaminants, and found distinct differences in their microbial community composition. We also found that the type of laboratories has a larger influence than the sampling site in shaping the microbial community, in terms of both structure and richness. On the other hand, the public areas of the different types of laboratories share very similar sets of microbes. Tracing the main sources of these microbes, we identified both environmental and human factors that are important factors in shaping the diversity and dynamics of these possible microbial contaminations in biological laboratories. These possible microbial contaminants that we have identified will be helpful for people who aim to eliminate them from samples.\n\nImportanceMicrobial communities from biological laboratories might hamper the conduction of molecular biology experiments, yet these possible contaminations are not yet carefully investigated. In this work, a metagenomic approach has been applied to identify the possible microbial contaminants and their sources, from the surfaces of various sites across different types of biological laboratories. We have found distinct differences in their microbial community compositions. We have also identified the main sources of these microbes, as well as important factors in shaping the diversity and dynamics of these possible microbial contaminations. The identification and interpretation of these possible microbial contaminants in biological laboratories would be helpful for alleviate their potential detrimental effects.

microbiology

Microbial Cells Harboring a Mitochondrial Gene Are Capable of CO2 Capture

Global warming is escalating with increased temperatures reported worldwide. Given the enormous land mass on the planet, biological capture of CO2 remains a viable approach to mitigate the crisis as it is economical and easy to implement. In this study, a gene capable of CO2 capture was identified via selection in minimal media. This mitochondrial gene named as OG1 encodes the OK/SW-CL.16 protein and shares homology with cytochrome oxidase subunit III of various species and PII uridylyl-transferase from Loktanella vestfoldensis SKA53. CO2 capture experiments indicate that {delta}13C was substantially higher in the cells harboring the gene OG1 than the control in the nutrition-poor media. This study suggests that CO2 capture using engineered microorganisms in barren land can be exploited to address the soaring CO2 level in the atmosphere, opening up vast land resources to cope with global warming.\n\nIMPORTANCEGlobal warming crisis is deteriorating with increased CO2 levels in the atmosphere each year. Action must be taken before catastrophic consequences occur in the not-so-distant future. Biological capture of CO2 is a feasible approach to alleviate the current crisis. We have identified a mitochondrial gene which demonstrated CO2 utilization capability. Data presented in this study suggest that CO2 capture using engineered microorganisms can be harnessed to address the ever-rising CO2 level in the atmosphere.

bioengineering

GranatumX: A community engaging and flexible software environment for single-cell analysis

We present GranatumX, a next-generation software environment for single-cell data analysis. GranatumX is inspired by the interactive web tool Granatum. It enables biologists to access the latest single-cell bioinformatics methods in a web-based graphical environment. It also offers software developers the opportunity to rapidly promote their own tools with others in customizable pipelines. The architecture of GranatumX allows for easy inclusion of plugin modules, named Gboxes, that wrap around bioinformatics tools written in various programming languages and on various platforms. GranatumX can be run on the cloud or private servers and generate reproducible results. It is a community-engaging, flexible, and evolving software ecosystem for scRNA-Seq analysis, connecting developers with bench scientists. GranatumX is freely accessible at http://garmiregroup.org/granatumx/app.

bioinformatics

Haplotype-resolved and integrated genome analysis of ENCODE cell line HepG2

The HepG2 cancer cell line is one of the most widely-used biomedical research and one of the main cell lines of ENCODE. Vast numbers of functional genomics and epigenomics datasets have been produced to characterize its biology. However, the correct interpretation such data requires an understanding of the cell lines genome sequence and genome structure. Using a variety of sequencing and analysis methods, we identified a wide spectrum of HepG2 genome characteristics: copy numbers of chromosomal segments, SNVs and Indels (corrected for aneuploidy), phased haplotypes extending to entire chromosome arms, loss of heterozygosity, retrotransposon insertions, structural variants (SVs) including complex and somatic genomic rearrangements. We also identified allele-specific expression and DNA methylation genome-wide and assembled an allele-specific CRISPR/Cas9 targeting map.\n\nSIGNIFICANCEHaplotype-resolved and comprehensive whole-genome analysis of a widely-used cell line for cancer research and ENCODE, HepG2, serves as an essential resource for unlocking complex cancer gene regulation using a genome-integrated framework and also provides genomic context for the analysis of ~1,000 functional datasets to date on ENCODE for biological discovery. We also demonstrate how deeper insights into genomic regulatory complexity are gained by adopting a genome-integrated framework.

genomics

Mother centrioles are dispensable for deuterosome formation and function during basal body amplification

Mammalian epithelial cells use a pair of mother centrioles (MCs) and numerous deuterosomes as platforms for efficient basal body assembly during multiciliogenesis. How deuterosomes form and function, however, remain controversial. They are proposed to either arise spontaneously followed by maturation into larger ones with increased procentriole-producing capacity or be assembled solely on the young MC, nucleate procentrioles under the MCs guidance, and released as procentriole-occupied \"halos\". Here we show that both MCs are dispensable for deuterosome formation in multiciliate cells. In both mouse tracheal epithelial and ependymal cells (mTECs and mEPCs), discrete deuterosomes in the cytoplasm were initially procentriole-free and then grew into halos. More importantly, eliminating the young MC or both MCs in proliferating precursor cells through shRNA-mediated depletion of Plk4, a kinase essential to procentriole assembly, did not abolish deuterosome formation when these cells were induced to differentiate into mEPCs. The average deuterosome numbers per cell only reduced by 21% as compared to control mEPCs. Therefore, MC is not essential to the assembly of both deuterosomes and deuterosome-mediated procentrioles.

cell biology

DeepImpute: an accurate, fast and scalable deep neural network method to impute single-cell RNA-Seq data

BackgroundSingle-cell RNA sequencing (scRNA-seq) offers new opportunities to study gene expression of tens of thousands of single cells simultaneously. However, a significant problem of current scRNA-seq data is the large fractions of missing values or \"dropouts\" in gene counts. Incorrect handling of dropouts may affect downstream bioinformatics analysis. As the number of scRNA-seq datasets grows drastically, it is crucial to have accurate and efficient imputation methods to handle these dropouts.\n\nMethodsWe present DeepImpute, a deep neural network based imputation algorithm. The architecture of DeepImpute efficiently uses dropout layers and loss functions to learn patterns in the data, allowing for accurate imputation.\n\nResultsOverall DeepImpute yields better accuracy than other publicly available scRNA-Seq imputation methods on experimental data, as measured by mean squared error or Pearsons correlation coefficient. Moreover, its efficient implementation provides significantly higher performance over the other methods as dataset size increases. Additionally, as a machine learning method, DeepImpute allows to use a subset of data to train the model and save even more computing time, without much sacrifice on the prediction accuracy.\n\nConclusionsDeepImpute is an accurate, fast and scalable imputation tool that is suited to handle the ever increasing volume of scRNA-seq data. The package is freely available at https://github.com/lanagarmire/DeepImpute

bioinformatics

An Atomistic view of Short-chain Antimicrobial Biomimetic peptides in Action

Amphiphilic {beta}-peptides, which are rationally designed synthetic oligomers, are established biomimetic alternatives of natural antimicrobial peptides. The ability of these biomimetic peptides to form helical amphiphilic conformation using small number of residues provides a greater synthetic advantage over the naturally occurring antimicrobial peptides, which is reflected in more potent antimicrobial activity of {beta}-peptides than its naturally occurring counterparts. Here we address whether the distinct molecular architecture of short-chain and rigid synthetic peptides compared to relatively long and flexible natural antimicrobial peptides translates to a distinct mechanistic action with membrane. By simulating the interaction of membrane with antimicrobial 10-residue {beta}-peptides at diverse range of concentrations we reveal spontaneous insertion of {beta}-peptides in the membrane interface at a low concentration and occurrence of partial water leakage in the membrane at a high concentration. Intriguingly, unlike prototypical natural antimicrobial peptides, the water molecules leaked inside the membrane by these biomimetic peptides do not span entire membrane, as supported by free energy analysis. As a major advancement, this work brings into lights the key distinction in the membrane-activity of short synthetic biomimetic oligomers relative to the natural long-chain antimicrobial peptides.

biochemistry

Probing the Acyl Carrier Protein-Enzyme Interactions within Terminal Alkyne Biosynthetic Machinery

The alkyne functionality has attracted much interest due to its diverse chemical and biological applications. We recently elucidated an acyl carrier protein (ACP)-dependent alkyne biosynthetic pathway, however, little is known about ACP interactions with the alkyne biosynthetic enzymes, an acyl-ACP ligase (JamA) and a membrane-bound bi-functional desaturase/acetylenase (JamB). Here, we showed that JamB has a more stringent interaction with ACP than JamA. In addition, site directed mutagenesis of a non-cognate ACP significantly improved its compatibility with JamB, suggesting a possible electrostatic interaction at the ACP-JamB interface. Finally, error-prone PCR and screening of a second non-cognate ACP identified hot spots on the ACP that are important for interacting with JamB and yielded mutants which were better recognized by JamB. Our data thus not only provide insights into the ACP interactions in alkyne biosynthesis, but it also potentially aids in future combinatorial biosynthesis of alkyne-tagged metabolites for chemical and biological applications.\n\nTopical HeadingBiomolecular Engineering, Bioengineering, Biochemicals, Biofuels, and Food

bioengineering

Transcriptome Landscape of Human Oocytes and Granulosa Cells Throughout Folliculogenesis

Folliculogenesis is a highly regulated process that involves bidirectional interactions of the oocytes and surrounding granulosa cells (GCs). Little is unknown, however, about the transcriptomic profiles of human oocytes and GCs throughout folliculogenesis. Here we performed a high resolution RNA-Seq of human oocytes and GCs at each follicular stage, which revealed unique transcriptional profiles, stage-specific signature genes, oocyte- and GC-derived genes that reflect ovarian reserve. We identified reciprocal cell-to-cell interactions between oocytes and GCs, including NOTCH, TGF-{beta} signaling and gap junctions and determined the expression patterns of maternal-effect genes involved in folliculogenesis and early embryogenesis. Finally, we demonstrated robust differences between human and mice oocyte transcriptomes. This is the first comprehensive overview of the transcriptomic signatures governing the stepwise human folliculogenesis in-vivo that provides a valuable resource for basic and translational research in human reproductive biology.

cell biology

Common antigenic motif recognized by human VH5-51/VL4-1 tau antibodies with distinct functionalities

Misfolding and aggregation of tau protein are closely associated with the onset and progression of Alzheimers Disease (AD). By interrogating IgG+ memory B cells from asymptomatic donors with tau peptides, we have identified two somatically mutated VH5-51/VL4-1 antibodies. One of these, CBTAU-27.1, binds to the aggregation motif in the R3 repeat domain and blocks the aggregation of tau into paired helical filaments (PHFs) by sequestering monomeric tau. The other, CBTAU-28.1, binds to the N-terminal insert region and inhibits the spreading of tau seeds and mediates the uptake of tau aggregates into microglia by binding PHFs. Crystal structures revealed that the combination of VH5-51 and VL4-1 recognizes a common Pro-Xn-Lys motif driven by germline-encoded hotspot interactions while the specificity and thereby functionality of the antibodies are defined by the CDR3 regions. Affinity improvement led to improvement in functionality, identifying their epitopes as new targets for therapy and prevention of AD.

neuroscience

Systematic Optimization of Whole Plant Carbon Nitrogen Interaction (WACNI) to Support Crop Design for Greater Yield

On the face of the rapid advances in genome editing technology and greatly expanded knowledge on plant genome and genes, there is a strong demand to develop an effective tool to guide designing crops for higher yields. Here we developed a highly mechanistic model of Whole plAnt Carbon Nitrogen Interaction (WACNI), which predicts crop yield based on major metabolic and biophysical processes in source, sink and transport tissues. WACNI accurately predicted the yield responses of so far reported source, sink and transport related genetic manipulations on rice grain yields. Systematic sensitivity analysis with WACNI was used to classify the source, sink and transport related molecular processes into four categories, i.e. universal yield enhancers, universal yield inhibitors, conditional yield enhancers and weak yield regulators. Simulations using WACNI further show that even without a major change in leaf photosynthetic properties, 54.6% to 73% grain yield increase can be potentially achieved by optimizing these molecular processes during the rice grain filling period while simply combining all the superior molecular modules together cannot achieve the optimal yield level. A common macroscopic feature in all these designed high-yield lines is that they all show a sustained and steady growth of grain sink, which might be used as a generic selection criteria in high-yield rice breeding. Overall, WACNI can serve as a tool to facilitate plant source sink interaction research and guide future crops breeding by design.\n\nOne sentence summaryA mechanistic model of source, sink flow model is developed and used to demonstrate that optimization of the whole plant carbon nitrogen metabolism can dramatically increase crop yield potential.

systems biology

Activity of Antimicrobial Peptides Decreases with Increased Cell Membrane Crossing Free Energy Cost

Antimicrobial peptides (AMPs) are a promising alternative to mitigating bacterial infections in light of increasing bacterial resistance to antibiotics. However, predicting, understanding, and controlling the antibacterial activity of AMPs remains a significant challenge. While peptide intramolecular interactions are known to modulate AMP antimi-crobial activity, peptide intermolecular interactions remain elusive in their impact on peptide bioactivity. Herein, we test the relationship between AMP intermolecular interactions and antibacterial efficacy by controlling AMP intermolecular hydrophobic and hydrogen bonding interactions. Molecular dynamics simulations and Gibbs free energy calculations in concert with experimental assays show that increasing intermolecular interactions via inter-peptide aggregation increases the energy cost for the peptide to cross the bacterial cell membrane, which in turn decreases the AMP antibacterial activity. Our findings provide a route for predicting and controlling the antibacterial activity of AMPs against Gramnegative bacteria via reductions of intermolecular AMP interactions.

biophysics

ClusterMine: a Knowledge-integrated Clustering Approach based on Expression Profiles of Gene Sets

MotivationClustering analysis is essential for understanding complex biological data. In widely used methods such as hierarchical clustering (HC) and consensus clustering (CC), expression profiles of all genes are often used to assess similarity between samples for clustering. These methods output sample clusters, but are not able to provide information about which gene sets (functions) contribute most to the clustering. So interpretability of their results is limited. We hypothesized that integrating prior knowledge of annotated biological processes would not only achieve satisfying clustering performance but also, more importantly, enable potential biological interpretation of clusters.\n\nResultsHere we report ClusterMine, a novel approach that identifies clusters by assessing functional similarity between samples through integrating known annotated gene sets, e.g., in Gene Ontology. In addition to outputting cluster membership of each sample as conventional approaches do, it outputs gene sets that are most likely to contribute to the clustering, a feature facilitating biological interpretation. Using three cancer datasets, two single cell RNA-sequencing based cell differentiation datasets, one cell cycle dataset and two datasets of cells of different tissue origins, we found that ClusterMine achieved similar or better clustering performance and that top-scored gene sets prioritized by ClusterMine are biologically relevant.\n\nImplementation and availabilityClusterMine is implemented as an R package and is freely available at: www.genemine.org/clustermine.php\n\nContactjxwang@csu.edu.cn\n\nSupplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics

Cryo-EM structure of an early precursor of large ribosomal subunit reveals a half assembled intermediate

Assembly of eukaryotic ribosome is a complicated and dynamic process that involves a series of intermediates. How the highly intertwined structure of 60S large ribosomal subunits is established is unknown. Here, we report the structure of an early nucleolar pre-60S ribosome determined by cryo-electron microscopy at 3.7 [A] resolution, revealing a half assembled subunit. Domains I, II and VI of 25S/5.8S rRNA tightly pack into a native-like substructure, but domains III, IV and V are not assembled. The structure contains 12 assembly factors and 19 ribosomal proteins, many of which are required for early processing of large subunit rRNA. The Brx1-Ebp2 complex would interfere with the assembly of domains IV and V. Rpf1, Mak16, Nsa1 and Rrp1 form a cluster that consolidates the joining of domains I and II. Our structure reveals a key intermediate on the path to the establishment of the global architecture of 60S subunits.

biophysics

TRIF is a key inflammatory mediator of acute sickness behavior and cancer cachexia

Hypothalamic inflammation is a key component of acute sickness behavior and cachexia, yet mechanisms of inflammatory signaling in the central nervous system remain unclear. We assessed the role of TRIF signaling in acute inflammation (lipopolysaccharide (LPS) challenge) and in a chronic inflammatory state (cancer cachexia). TRIFKO mice resisted anorexia and weight loss after peripheral (intraperitoneal, IP) or central (intracerebroventricular, ICV) LPS challenge and in a model of pancreatic cancer cachexia. Compared to WT mice, TRIFKO mice showed attenuated upregulation of Il6, Ccl2, Ccl5, Cxcl1, Cxcl2, and Cxcl10 in the hypothalamus after IP LPS treatment, as well as attenuated microglial activation and neutrophil infiltration into the brain after ICV LPS treatment. Our results show that TRIF is an important inflammatory signaling mediator of sickness behavior and cachexia and presents a novel therapeutic target for these conditions.

physiology