bioRxiv ScienceSearch

Biology subjects

Yu, W.

Publications and source records attributed to Yu, W..

15 recordsLinked to original sources

Roles of Polycomb gene EED in pathogenesis and prognosis of acute myeloid leukemia and diffuse large B cell lymphoma

In this study, we performed correlation analysis of polycomb gene EED and hematologic malignancies using the omics and clinical data of acute myeloid leukemia (LAML) and diffuse large B-Cell lymphoma (DLBC) from TCGA database. We found that: (1) High EED mRNA level was associated with poor prognosis and high CALGB cytogenetics risk of LAML patients. (2) EED mRNA level in DLBC cancer cells was higher than control cells. (3) EED gene expression could be regulated by both copy number alterations and DNA methylation. (4) Additionally, there were different EED co-expression genes nets in the two kinds of hematologic malignancies. In all, we confirmed that there are potential clinical significance of EED gene in pathogenesis and prognosis of hematologic malignancies.

systems biology

The Dynamic Conformational Landscapes of the Protein Methyltransferase SETD8

Elucidating conformational heterogeneity of proteins is essential for understanding protein functions and developing exogenous ligands for chemical perturbation. While structural biology methods can provide atomic details of static protein structures, these approaches cannot in general resolve less populated, functionally relevant conformations and uncover conformational kinetics. Here we demonstrate a new paradigm for illuminating dynamic conformational landscapes of target proteins. SETD8 (Pr-SET7/SET8/KMT5A) is a biologically relevant protein lysine methyltransferase for in vivo monomethylation of histone H4 lysine 20 and nonhistone targets. Utilizing covalent chemical inhibitors and depleting native ligands to trap hidden high-energy conformational states, we obtained diverse novel X-ray structures of SETD8. These structures were used to seed massively distributed molecular simulations that generated six milliseconds of trajectory data of SETD8 in the presence or absence of its cofactor. We used an automated machine learning approach to reveal slow conformational motions and thus distinct conformational states of SETD8, and validated the resulting dynamic conformational landscapes with multiple biophysical methods. The resulting models provide unprecedented mechanistic insight into how protein dynamics plays a role in SAM binding and thus catalysis, and how this function can be modulated by diverse cancer-associated mutants. These findings set up the foundation for revealing enzymatic mechanisms and developing inhibitors in the context of conformational landscapes of target proteins.

biophysics

Genomic analysis for heavy metal resistance in S. maltophilia

Stenotrophomonas maltophilia is highly resistant to heavy metals, but the genetic knowledge of metal resistance in S. maltophilia is poorly understood. In this study, the genome of S.maltophilia Pho isolated from the contaminated soil near a metalwork factory was sequenced using PacBio RS II. Its genome is composed of a single chromosome with a GC content of 66.4% and 4434 protein-encoding genes. Comparative analysis revealed high syntney between S.maltophilia Pho and the model strain, S.maltophilia K279a. Then, the type and number of mechanisms for heavy metal uptake were analyzed firstly. Results revealed 7 unspecific ion transporter genes and 13 specific ion transporter genes, most of which were involved in iron transport. But the sulfate permeases belonging to the family of SulT/CysP that can uptake chromate and the high affinity ZnuABC/SitABCD were absent. Secondly, the putative genes controlling metal efflux were identified. Results showed that this bacterium encoded 5 CDFs, 1 copper exporting ATPase and 4 RND systems, including 2 CzcABC efflux pumps. Moreover, the putative metal transformation genes including arsenate and mercury detoxification genes were also identified. This study may provide useful information on the metal resistance mechanisms of S.maltophilia.

microbiology

Sibe: a computation tool to apply protein sequence statistics to folding and design

Statistical analysis plays a significant role in both protein sequences and structures, expanding in recent years from the studies of co-evolution guided single-site mutations to protein folding in silico. Here we describe a computational tool, termed Sibe, with a particular focus on protein sequence analysis, folding and design. Since Sibe has various easy-interface modules, expressive architecture and extensible codes, it is powerful in statistically analyzing sequence data and building energetic potentials in boosting both protein folding and design. In this study, Sibe is used to capture positionally conserved couplings between pairwise amino acids and help rational protein design, in which the pairwise couplings are filtered according to the relative entropy computed from the positional conservations and grouped into several blocks. A human {beta}2-adrenergic receptor ({beta}2AR) was used to demonstrated that those blocks could contribute rational design at functional residues. In addition, Sibe provides protein folding modules based on both the positionally conserved couplings and well-established statistical potentials. Sibe provides various easy to use command-line interfaces in C++ and/or Python. Sibe was developed for compatibility with the big data era, and it primarily focuses on protein sequence analysis, in silico folding and design, but it is also applicable to extend for other modeling and predictions of experimental measurements.

bioinformatics

Chronic inflammatory pain drives alcohol drinking in a sex-dependent manner

Sex differences in chronic pain and alcohol abuse are not well understood. The development of rodent models is imperative for investigating the underlying changes behind these pathological states. However, past attempts have failed to produce drinking outcomes similar to those reported in humans. In the present study, we investigated whether hind paw treatment with the inflammatory agent Complete Freunds Adjuvant (CFA) could generate hyperalgesia and alter alcohol consumption in male and female C57BL/6J mice. CFA treatment led to greater nociceptive sensitivity for both sexes in the Hargreaves test, and increased alcohol drinking for males in a continuous access two-bottle choice (CA2BC) paradigm. Regardless of treatment, female mice exhibited greater alcohol drinking than males. Following a 2-hour terminal drinking session, CFA treatment failed to produce changes in alcohol drinking, blood ethanol concentration (BEC), and plasma corticosterone (CORT) for both sexes. 2-hr alcohol consumption and CORT was higher in females than males, irrespective of CFA treatment. Taken together, these findings have established that male mice are more susceptible to escalations in alcohol drinking when undergoing pain, despite higher levels of total alcohol drinking and CORT in females. Furthermore, the exposure of CFA-treated C57BL/6J mice to the CA2BC drinking paradigm has proven to be a useful model for studying the relationship between chronic pain and alcohol abuse. Future applications of the CFA/CA2BC model should incorporate manipulations of stress signaling and other related biological systems to improve our mechanistic understanding of pain and alcohol interactions.

animal behavior and cognition

De novo protein structure prediction using ultra-fast molecular dynamics simulation

Modern genomics sequencing techniques have provided a massive amount of protein sequences, but experimental endeavor in determining protein structures is largely lagging far behind the vast and unexplored sequences. Apparently, computational biology is playing a more important role in protein structure prediction than ever. Here, we present a system of de novo predictor, termed NiDelta, building on a deep convolutional neural network and statistical potential enabling molecular dynamics simulation for modeling protein tertiary structure. Combining with evolutionary-based residue-contacts, the presented predictor can predict the tertiary structures of a number of target proteins with remarkable accuracy. The proposed approach is demonstrated by calculations on a set of eighteen large proteins from different fold classes. The results show that the ultra-fast molecular dynamics simulation could dramatically reduce the gap between the sequence and its structure at atom level, and it could also present high efficiency in protein structure determination if sparse experimental data is available.

bioinformatics

Understanding the limit of open search in the identification of peptides with post-translational modifications -- A simulation-based study

MotivationAnalyzing tandem mass spectrometry data to recognize peptides in a sample is the fundamental task in computational proteomics. Traditional peptide identification algorithms perform well when identifying unmodified peptides. However, when peptides have post-translational modifications (PTMs), these methods cannot provide satisfactory results. Recently, Chick et al., 2015 and Yu et al., 2016 proposed the spectrum-based and tag-based open search methods, respectively, to identify peptides with PTMs. While the performance of these two methods is promising, the identification results vary greatly with respect to the quality of tandem mass spectra and the number of PTMs in peptides. This motivates us to systematically study the relationship between the performance of open search methods and quality parameters of tandem mass spectrum data, as well as the number of PTMs in peptides.\n\nResultsThrough large-scale simulations, we obtain the performance trend when simulated tandem mass spectra are of different quality. We propose an analytical model to describe the relationship between the probability of obtaining correct identifications and the spectrum quality as well as the number of PTMs. Based on the analytical model, we can quantitatively describe the necessary condition to effectively apply open search methods.\n\nAvailabilitySource codes of the simulation are available at http://bioinformatics.ust.hk/PST.html.\n\nContactboningli@ust.hk or eeyu@ust.hk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Common antigenic motif recognized by human VH5-51/VL4-1 tau antibodies with distinct functionalities

Misfolding and aggregation of tau protein are closely associated with the onset and progression of Alzheimers Disease (AD). By interrogating IgG+ memory B cells from asymptomatic donors with tau peptides, we have identified two somatically mutated VH5-51/VL4-1 antibodies. One of these, CBTAU-27.1, binds to the aggregation motif in the R3 repeat domain and blocks the aggregation of tau into paired helical filaments (PHFs) by sequestering monomeric tau. The other, CBTAU-28.1, binds to the N-terminal insert region and inhibits the spreading of tau seeds and mediates the uptake of tau aggregates into microglia by binding PHFs. Crystal structures revealed that the combination of VH5-51 and VL4-1 recognizes a common Pro-Xn-Lys motif driven by germline-encoded hotspot interactions while the specificity and thereby functionality of the antibodies are defined by the CDR3 regions. Affinity improvement led to improvement in functionality, identifying their epitopes as new targets for therapy and prevention of AD.

neuroscience

Septal Secretion of Protein A in Staphylococcus aureus Requires SecA and Lipoteichoic Acid Synthesis

Surface proteins of Staphylococcus aureus are secreted across septal membranes for assembly into the bacterial cross-wall. This localized secretion requires the YSIRK/GXXS motif signal peptide, however the mechanisms supporting precursor trafficking are not known. We show here that the signal peptide of staphylococcal protein A (SpA) is cleaved at the YSIRK/GXXS motif. A signal peptide mutant defective for cleavage can be crosslinked to SecA, SecDF and LtaS. SecA depletion blocks precursor targeting to septal membranes, whereas deletion of secDF diminishes SpA secretion into the cross-wall. Depletion of LtaS blocks lipoteichoic acid synthesis and promotes precursor trafficking to peripheral membranes. We propose a model whereby SecA directs SpA precursors to lipoteichoic acid-rich septal membranes for YSIRK/GXXS motif cleavage and secretion into the cross-wall.

microbiology

Resting-state connectivity predicts patient-specific effects of deep brain stimulation for Parkinson’s disease

Neural circuit-based guidance for optimizing patient screening, target selection and parameter tuning for deep brain stimulation (DBS) remains limited. To this end, we propose a functional brain connectome-based modeling approach that simulates network-spreading effects of stimulating different brain regions and quantifies rectification of abnormal network topology in silico. We validate these analyses by predicting nuclei in basal-ganglia circuits as top-ranked targets for 43 local patients with Parkinsons disease and 90 patients from public database. However, individual connectome-based predictions demonstrate that globus pallidus and subthalamic nucleus (STN) constituted as the best choice for 21.1% and 19.5% of patients, respectively. Notably, the priority rank of STN significantly correlated with motor symptom severity in the local cohort. By introducing whole-brain network diffusion dynamics, these findings unfold a new dimension of brain connectomics and underscore the importance of neural network modeling for personalized DBS therapy, which warrants experimental investigation to validate its clinical utility.

neuroscience

Interplay between antibiotic efficacy and drug-induced lysis underlie enhanced biofilm formation at subinhibitory drug concentrations

Subinhibitory concentrations of antibiotics have been shown to enhance biofilm formation in multiple bacterial species. While antibiotic exposure has been associated with modulated expression in many biofilm-related genes, the mechanisms of drug-induced biofilm formation remain a focus of ongoing research efforts and may vary significantly across species. In this work, we investigate antibiotic-induced biofilm formation in E. faecalis, a leading cause of nosocomial infections. We show that biofilm formation is enhanced by subinhibitory concentrations of cell wall synthesis inhibitors, but not by inhibitors of protein, DNA, folic acid, or RNA synthesis. Furthermore, enhanced biofilm is associated with increased cell lysis, an increase in extracellular DNA (eDNA), and an increase in the density of living cells in the biofilm. In addition, we observe similar enhancement of biofilm formation when cells are treated with non-antibiotic surfactants that induce cell lysis. These findings suggest that antibiotic-induced biofilm formation is governed by a trade-off between drug toxicity and the beneficial effects of cell lysis. To understand this trade-off, we developed a simple mathematical model that predicts changes to antibiotic-induced biofilm formation due to external perturbations, and we verify these predictions experimentally. Specifically, we demonstrate that perturbations that reduce eDNA (DNase treatment) or decrease the number of living cells in the planktonic phase (a second antibiotic) decrease biofilm induction, while chemical inhibitors of cell lysis increase relative biofilm induction and shift the peak to higher antibiotic concentrations. Overall, our results offer experimental evidence linking cell wall synthesis inhibitors, cell lysis, increased eDNA, and biofilm formation in E. faecalis while also providing a predictive, quantitative model that sheds light on the interplay between cell lysis and antibiotic efficacy in developing biofilms.

microbiology

Mice use robust and common strategies to discriminate natural scenes

Mice use vision to navigate and avoid predators in natural environments. However, the spatial resolution of mouse vision is poor compared to primates, and mice lack a fovea. Thus, it is unclear how well mice can discriminate ethologically relevant scenes. Here, we examined natural scene discrimination in mice using an automated touch-screen system. We estimated the discrimination difficulty using the computational metric structural similarity (SSIM), and constructed psychometric curves. However, the performance of each mouse was better predicted by the population mean than SSIM. This high inter-mouse agreement indicates that mice use common and robust strategies to discriminate natural scenes. We tested several other image metrics to find an alternative to SSIM for predicting discrimination performance. We found that a simple, primary visual cortex (V1)-inspired model predicted mouse performance with fidelity approaching the inter-mouse agreement. The model involved convolving the images with Gabor filters, and its performance varied with the orientation of the Gabor filter. This orientation dependence was driven by the stimuli, rather than an innate biological feature. Together, these results indicate that mice are adept at discriminating natural scenes, and their performance is well predicted by simple models of V1 processing.

neuroscience

Xolik: finding cross-linked peptides with maximum paired scores in linear time

MotivationCross-linking technique coupled with mass spectrometry (MS) is widely used in the analysis of protein structures and protein-protein interactions. In order to identify cross-linked peptides from MS data, we need to consider all pairwise combinations of peptides, which is computationally prohibitive when the sequence database is large. To alleviate this problem, some heuristic screening strategies are used to reduce the number of peptide pairs during the identification. However, heuristic screening criteria may ignore true findings.\n\nResultsWe directly tackle the combination challenge without using any screening strategies. With the additive scoring function and the data structure of double-ended queue, the proposed algorithm reduces the quadratic time complexity of exhaustive searching down to the linear time complexity. We implement the algorithm in a tool named Xolik, and the running time of Xolik is validated using databases with different number of proteins. Experiments using synthetic and empirical datasets show that Xolik outperforms existing tools in terms of running time and statistical power.\n\nAvailabilitySource code and binaries of Xolik are freely available at http://bioinformatics.ust.hk/Xolik.html.\n\nContacteeyu@ust.hk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Genome-Scale Mutational Signatures Of Aflatoxin In Cells, Mice And Human Tumors

Aflatoxin B1 (AFB1) is a mutagen and IARC Group 1 carcinogen that causes hepatocellular carcinoma (HCC). Here we present the first whole genome data on the mutational signatures of AFB1 exposure from a total of > 40,000 mutations in four experimental systems: two different human cell lines, and in liver tumors in wild-type mice and in mice that carried a hepatitis B surface antigen transgene - this to model the multiplicative effects of aflatoxin exposure and hepatitis B in causing HCC. AFB1 mutational signatures from all four experimental systems were remarkably similar. We integrated the experimental mutational signatures with data from newly-sequenced HCCs from Qidong County, China, a region of well-studied aflatoxin exposure. This indicated that COSMIC mutational signature 24, previously hypothesized to stem from aflatoxin exposure, indeed likely represents AFB1 exposure, possibly combined with other exposures. Among published somatic mutation data, we found evidence of AFB1 exposure in 0.7% of HCCs treated in North America, 1% of HCCs from Japan, but 16% of HCCs from Hong Kong. Thus, aflatoxin exposure apparently remains a substantial public health issue in some areas. This aspect of our study exemplifies the promise of future widespread resequencing of tumor genomes in providing new insights into the contribution of mutagenic exposures to cancer incidence.

cancer biology

ECL 2.0: Exhaustively Identifying Cross-Linked Peptides with a Linear Computational Complexity

Chemical cross-linking coupled with mass spectrometry is a powerful tool to study protein-protein interactions and protein conformations. Two linked peptides are ionized and fragmented to produce a tandem mass spectrum. In such an experiment, a tandem mass spectrum contains ions from two peptides. The peptide identification problem becomes a peptide-peptide pair identification problem. Currently, most existing tools dont search all possible pairs due to the quadratic time complexity. Consequently, a significant percentage of linked peptides are missed. In our earlier work, we developed a tool named ECL to search all pairs of peptides exhaustively. While ECL does not miss any linked peptides, it is very slow due to the quadratic computational complexity, especially when the database is large. Furthermore, ECL uses a score function without statistical calibration, while researchers1,2 have demonstrated that using a statistical calibrated score function can achieve a higher sensitivity than using an uncalibrated one.\n\nHere, we propose an advanced version of ECL, named ECL 2.0. It achieves a linear time and space complexity by taking advantage of the additive property of a score function. It can analyze a typical data set containing tens of thousands of spectra using a large-scale database containing thousands of proteins in a few hours. Comparison with other five state-of-the-art tools shows that ECL 2.0 is much faster than pLink, StavroX, ProteinProspector, and ECL. Kojak is the only one tool that is faster than ECL 2.0. But Kojak does not exhaustively search all possible peptide pairs. We also adopt an e-value estimation method to calibrate the original score. Comparison shows that ECL 2.0 has the highest sensitivity among the state-of-the-art tools. The experiment using a large-scale in vivo cross-linking data set demonstrates that ECL 2.0 is the only tool that can find PSMs passing the false discovery rate threshold. The result illustrates that exhaustive search and well calibrated score function are useful to find PSMs from a huge search space.

bioinformatics