bioRxiv ScienceSearch

Biology subjects

Yao, W.

Publications and source records attributed to Yao, W..

5 recordsLinked to original sources

Overview of the SAMPL6 host-guest binding affinity prediction challenge

Accurately predicting the binding affinities of small organic molecules to biological macro-molecules can greatly accelerate drug discovery by reducing the number of compounds that must be synthesized to realize desired potency and selectivity goals. Unfortunately, the process of assessing the accuracy of current computational approaches to affinity prediction against binding data to biological macro-molecules is frustrated by several challenges, such as slow conformational dynamics, multiple titratable groups, and the lack of high-quality blinded datasets. Over the last several SAMPL blind challenge exercises, host-guest systems have emerged as a practical and effective way to circumvent these challenges in assessing the predictive performance of current-generation quantitative modeling tools, while still providing systems capable of possessing tight binding affinities. Here, we present an overview of the SAMPL6 host-guest binding affinity prediction challenge, which featured three supramolecular hosts: octa-acid (OA), the closely related tetra-endo-methyl-octa-acid (TEMOA), and cucurbit[8]uril (CB8), along with 21 small organic guest molecules. A total of 119 entries were received from 10 participating groups employing a variety of methods that spanned from electronic structure and movable type calculations in implicit solvent to alchemical and potential of mean force strategies using empirical force fields with explicit solvent models. While empirical models tended to obtain better performance than first-principle methods, it was not possible to identify a single approach that consistently provided superior results across all host-guest systems and statistical metrics. Moreover, the accuracy of the methodologies generally displayed a substantial dependence on the system considered, emphasizing the need for host diversity in blind evaluations. Several entries exploited previous experimental measurements of similar host-guest systems in an effort to improve their physical-based predictions via some manner of rudimentary machine learning; while this strategy succeeded in reducing systematic errors, it did not correspond to an improvement in statistical correlation. Comparison to previous rounds of the host-guest binding free energy challenge highlights an overall improvement in the correlation obtained by the affinity predictions for OA and TEMOA systems, but a surprising lack of improvement regarding root mean square error over the past several challenge rounds. The data suggests that further refinement of force field parameters, as well as improved treatment of chemical effects (e.g., buffer salt conditions, protonation states) may be required to further enhance predictive accuracy.

biophysics

YAP1 Oncogene is a Context-specific Driver for Pancreatic Ductal Adenocarcinoma

AbstractTranscriptomic profiling classifies pancreatic ductal adenocarcinoma (PDAC) into several molecular subtypes with distinctive histological and clinical characteristics. However, little is known about the molecular mechanisms that define each subtype and their correlation with clinical outcome. Mutant KRAS is the most prominent driver in PDAC, present in over 90% of tumors, but the dependence of tumors on oncogenic KRAS signaling varies between subtypes. In particular, squamous subtype are relatively independent of oncogenic KRAS signaling and typically display much more aggressive clinical behavior versus progenitor subtype. Here, we identified that YAP1 activation is enriched in the squamous subtype and associated with poor prognosis. Activation of YAP1 in progenitor subtype cancer cells profoundly enhanced malignant phenotypes and transformed progenitor subtype cells into squamous subtype. Conversely, depletion of YAP1 specifically suppressed tumorigenicity of squamous subtype PDAC cells. Mechanistically, we uncovered a significant positive correlation between WNT5A expression and the YAP1 activity in human PDAC, and demonstrated that WNT5A overexpression led to YAP1 activation and recapitulated YAP1-dependent but Kras-independent phenotype of tumor progression and maintenance. Thus, our study identifies YAP1 oncogene as a major driver of squamous subtype PDAC and uncovers the role of WNT5A in driving PDAC malignancy through activation of the YAP pathway.

cancer biology

The Juicebox Assembly Tools module facilitates de novo assembly of mammalian genomes with chromosome-length scaffolds for under $1000

Hi-C contact maps are valuable for genome assembly (Lieberman-Aiden, van Berkum et al. 2009; Burton et al. 2013; Dudchenko et al. 2017). Recently, we developed Juicebox, a system for the visual exploration of Hi-C data (Durand, Robinson et al. 2016), and 3D-DNA, an automated pipeline for using Hi-C data to assemble genomes (Dudchenko et al. 2017). Here, we introduce \"Assembly Tools,\" a new module for Juicebox, which provides a point-and-click interface for using Hi-C heatmaps to identify and correct errors in a genome assembly. Together, 3D-DNA and the Juicebox Assembly Tools greatly reduce the cost of accurately assembling complex eukaryotic genomes. To illustrate, we generated de novo assemblies with chromosome-length scaffolds for three mammals: the wombat, Vombatus ursinus (3.3Gb), the Virginia opossum, Didelphis virginiana (3.3Gb), and the raccoon, Procyon lotor (2.5Gb). The only inputs for each assembly were Illumina reads from a short insert DNA-Seq library (300 million Illumina reads, maximum length 2x150 bases) and an in situ Hi-C library (100 million Illumina reads, maximum read length 2x150 bases), which cost <$1000.

genomics

The INO80 Chromatin Remodeler Sustains Metabolic Stability by Promoting TOR Signaling and Regulating Histone Acetylation

Chromatin remodeling complexes are essential for gene expression programs that coordinate cell function with metabolic status. However, how these remodelers are integrated in metabolic stability pathways is not well known. Here, we report an expansive genetic screen with chromatin remodelers and metabolic regulators in Saccharomyces cerevisiae. We found that, unlike the SWR1 remodeler, the INO80 chromatin remodeling complex is composed of multiple distinct functional subunit modules. We identified a strikingly divergent genetic signature for the Ies6 subunit module that links the INO80 complex to metabolic homeostasis, including mitochondrial maintenance. INO80 is also needed to communicate TORC1-mediated signaling to chromatin and maintains histone acetylation at TORC1-responsive genes. Furthermore, computational analysis reveals subunits of INO80 and mTORC1 have high co-occurrence of alterations in human cancers. Collectively, these results demonstrate that the INO80 complex is a central component of metabolic homeostasis that influences histone acetylation and may contribute to disease when disrupted.

genetics

YASS: Yet Another Spike Sorter

Spike sorting is a critical first step in extracting neural signals from large-scale electrophysiological data. This manuscript describes an efficient, reliable pipeline for spike sorting on dense multi-electrode arrays (MEAs), where neural signals appear across many electrodes and spike sorting currently represents a major computational bottleneck. We present several new techniques that make dense MEA spike sorting more robust and scalable. Our pipeline is based on an efficient multi-stage \"triage-then-cluster-then-pursuit\" approach that initially extracts only clean, high-quality waveforms from the electrophysiological time series by temporarily skipping noisy or \"collided\" events (representing two neurons firing synchronously). This is accomplished by developing a neural network detection method followed by efficient outlier triaging. The clean waveforms are then used to infer the set of neural spike waveform templates through nonparametric Bayesian clustering. Our clustering approach adapts a \"coreset\" approach for data reduction and uses efficient inference methods in a Dirichlet process mixture model framework to dramatically improve the scalability and reliability of the entire pipeline. The \"triaged\" waveforms are then finally recovered with matching-pursuit deconvolution techniques. The proposed methods improve on the state-of-the-art in terms of accuracy and stability on both real and biophysically-realistic simulated MEA data. Furthermore, the proposed pipeline is efficient, learning templates and clustering much faster than real-time for a [~=] 500-electrode dataset, using primarily a single CPU core.

neuroscience