bioRxiv ScienceSearch

Biology subjects

Cheng, H.

Publications and source records attributed to Cheng, H..

9 recordsLinked to original sources

BitMapperBS: a fast and accurate read aligner for whole-genome bisulfite sequencing

As a gold-standard technique for DNA methylation analysis, whole-genome bisulfite sequencing (WGBS) helps researchers to study the genome-wide DNA methylation at single-base resolution. However, aligning WGBS reads to the large reference genome is a major computational bottleneck in DNA methylation analysis projects. Although several WGBS aligners have been developed in recent years, it is difficult for them to efficiently process the ever-increasing bisulfite sequencing data. Here we propose BitMapperBS, an ultrafast and memory-efficient aligner that is designed for WGBS reads. To improve the performance of BitMapperBS, we propose various strategies specifically for the challenges that are unique to the WGBS aligners, which are ignored in most existing methods. Our experiments on real and simulated datasets show that BitMapperBS is one order of magnitude faster than the state-of-the-art WGBS aligners, while achieves similar or better sensitivity and precision. BitMapperBS is freely available at https://github.com/chhylp123/BitMapperBS.

bioinformatics

Gata4 drives Hh-signaling for second heart field migration and outflow tract development

Dominant mutations of Gata4, an essential cardiogenic transcription factor (TF), cause outflow tract (OFT) defects in both human and mouse. We investigated the molecular mechanism underlying this requirement. Gata4 happloinsufficiency in mice caused OFT defects including double outlet right ventricle (DORV) and conal ventricular septum defects (VSDs). We found that Gata4 is required within Hedgehog (Hh)-receiving second heart field (SHF) progenitors for normal OFT alignment. Increased Pten-mediated cell-cycle transition, rescued atrial septal defects but not OFT defects in Gata4 heterozygotes. SHF Hh-receiving cells failed to migrate properly into the proximal OFT cushion in Gata4 heterozygote embryos. We find that Hh signaling and Gata4 genetically interact for OFT development. Gata4 and Smo double heterozygotes displayed more severe OFT abnormalities including persistent truncus arteriosus (PTA) whereas restoration of Hedgehog signaling rescued OFT defects in Gata4-mutant mice. In addition, enhanced expression of the Gata6 was observed in the SHF of the Gata4 heterozygotes. These results suggested a SHF regulatory network comprising of Gata4, Gata6 and Hh-signaling for OFT development. This study indicates that Gata4 potentiation of Hh signaling is a general feature of Gata4-mediated cardiac morphogenesis and provides a model for the molecular basis of CHD caused by dominant transcription factor mutations.\n\nAuthor SummaryGata4 is an important protein that controls the development of the heart. Human who possess a single copy of Gata4 mutation display congenital heart defects (CHD), including the double outlet right ventricle (DORV). DORV is an alignment problem in which both the Aorta and Pulmonary Artery originate from the right ventricle, instead of originating from the left and the right ventricles, respectively. To study how Gata4 mutation causes DORV, we used a Gata4 mutant mouse model, which displays DORV. We showed that Gata4 is required in the cardiac precursor cells for the normal alignment of the great arteries. Although Gata4 mutation inhibits the rapid increase in number of the cardiac precursor cells, rescuing this defects does not recover the normal alignment of the great arteries. In addition, there is a movement problem of the cardiac precursor cells when migrating toward the great arteries during development. We further showed that a specific molecular signaling, Hh-signaling, is responsible to the Gata4 action in the cardiac precursor cells. Importantly, over-activating the Hh-signaling rescues the DORV in the Gata4 mutant embryos. This study provides an explanation for the ontogeny of CHD.

developmental biology

EnsembleCNV: An ensemble machine learning algorithm to identify and genotype copy number variation using SNP array data

The associations between diseases/traits and copy number variants (CNVs) have not been systematically investigated in genome-wide association studies (GWASs), primarily due to a lack of robust and accurate tools for CNV genotyping. Herein, we propose a novel ensemble learning framework, ensembleCNV, to detect and genotype CNVs using single nucleotide polymorphism (SNP) array data. EnsembleCNV a) identifies and eliminates batch effects at raw data level; b) assembles individual CNV calls into CNV regions (CNVRs) from multiple existing callers with complementary strengths by a heuristic algorithm; c) re-genotypes each CNVR with local likelihood model adjusted by global information across multiple CNVRs; d) refines CNVR boundaries by local correlation structure in copy number intensities; e) provides direct CNV genotyping accompanied with confidence score, directly accessible for downstream quality control and association analysis. Benchmarked on two large datasets, ensembleCNV outperformed competing methods and achieved a high call rate (93.3%) and reproducibility (98.6%), while concurrently achieving high sensitivity by capturing 85% of common CNVs documented in the 1000 Genomes Project. Given this CNV call rate and accuracy, which are comparable to SNP genotyping, we suggest ensembleCNV holds significant promise for performing genome-wide CNV association studies and investigating how CNVs predispose to human diseases.

bioinformatics

Crystal structure of the MBD domain of MBD3 in complex with methylated CG DNA

MBD3 is a core subunit of the Mi-2/NuRD complex, and has been previously reported to lack methyl-CpG binding ability. However, recent reports show that MBD3 recognizes both mCG and hmCG DNA with a preference for hmCG, and is required for the normal expression of hmCG marked genes in ES cells. Nevertheless, it is not clear how MBD3 recognizes the methylated DNA. In this study, we carried out structural analysis coupled with isothermal titration calorimetry (ITC) binding assay and mutagenesis studies to address the structural basis for the mCG DNA binding ability of the MBD3 MBD domain. We found that the MBD3 MBD domain prefers binding mCG over hmCG through the conserved arginine fingers, and this MBD domain as well as other mCG binding MBD domains can recognize the mCG duplex without orientation selectivity. Furthermore, we found that the tyrosine-to-phenylalanine substitution at Phe34 of MBD3 is responsible for its weaker mCG DNA binding ability compared to other mCG binding MBD domains. In summary, our study demonstrates that the MBD3 MBD domain is a mCG binder, and also illustrates its binding mechanism to the methylated CG DNA.

biochemistry

NDUFAB1 Protects Heart by Coordinating Mitochondrial Respiratory Complex and Supercomplex Assembly

The impairment of mitochondrial bioenergetics, often coupled with exaggerated reactive oxygen species (ROS) production, is emerging as a common mechanism in diseases of organs with a high demand for energy, such as the heart. Building a more robust cellular powerhouse holds promise for protecting these organs in stressful conditions. Here, we demonstrate that NDUFAB1 (NADH:ubiquinone oxidoreductase subunit AB1), acts as a powerful cardio-protector by enhancing mitochondrial energy biogenesis. In particular, NDUFAB1 coordinates the assembly of respiratory complexes I, II, and III and supercomplexes, conferring greater capacity and efficiency of mitochondrial energy metabolism. Cardiac-specific deletion of Ndufab1 in mice caused progressive dilated cardiomyopathy associated with defective bioenergetics and elevated ROS levels, leading to heart failure and sudden death. In contrast, transgenic overexpression of Ndufab1 effectively enhanced mitochondrial bioenergetics and protected the heart against ischemia-reperfusion injury. Our findings identify NDUFAB1 as a central endogenous regulator of mitochondrial energy and ROS metabolism and thus provide a potential therapeutic target for the treatment of heart failure and other mitochondrial bioenergetics-centered diseases.

cell biology

Loss of SDHB reprograms energy metabolisms and inhibits high fat diet induced metabolic syndromes

Mitochondrial respiratory complex II utilizes succinate, key substrate of the Krebs cycle, for oxidative phosphorylation, which is essential for glucose metabolism. Mutations of complex II cause cancers and mitochondrial diseases, raising a critical question of the (patho-)physiological functions. To address the fundamental role of complex II in systemic energy metabolism, we specifically knockout SDHB in mice liver, a key complex II subunit that tethers the catalytic SDHA subunit and transfers the electrons to ubiquinone, and found that SHDB deficiency abolishes the assembly of complex II without affecting other respiration complexes while largely retaining SDHA stability. SHDB ablation reprograms energy metabolism and hyperactivates the glycolysis, Krebs cycle and {beta}-oxidation pathways, leading to catastrophic energy deficit and early death. Strikingly, sucrose supplementation or high fat diet resumes both glucose and lipid metabolism and prevent early death. Also, SDHB deficient mice are completely resistant to high fat diet induced obesity. Our findings reveal that the unanticipated role of complex II orchestrating both lipid and glucose metabolisms, and suggest that SDHB is an ideal therapeutic target for combating obesity.

molecular biology

An evidence-based approach to globally assess the covariate-dependent effect of MTHFR SNP rs1801133 on plasma homocysteine: a systematic review and meta-analysis

BackgroundThe single nucleotide polymorphism (SNP) of the gene Methylenetetrahydrofolate Reductase (MTHFR) C677T (or rs1801133) is the most established genetic factor that increases plasma total homocysteine (tHcy) and consequently results in hyperhomocysteinemia. Yet given the limited penetrance of this genetic variant, it is necessary to individually predict the risk of hyperhomocysteinemia for a rs1801133 carrier.\n\nObjectiveWe hypothesized that variability of this genetic risk is largely due to the presence of factors (covariates) that serve as effect modifiers and/or confounders, such as folic acid (FA) intake, and aimed to assess this risk in the complex context of these covariates.\n\nDesignWe systematically extracted from published studies the data of tHcy, rs1801133, and any previously reported rs1801133 covariates. The resulting meta-dataset was first used to analyze the covariates modifying effect by meta regression and other statistical means. Subsequently, we stratified tHcy data by the rs1801133 genotypes and analyzed under each genotype the variability of the risk resulted from the covariates confounding.\n\nResultsThe dataset contains data of 36 rs1801133 covariates that were collected from 114,448 subjects and 249 qualified studies, among which 6 covariates (sex, age, race, FA intake, smoking, and alcohol consumption) are the most frequently informed and therefore included for statistical analysis. The effect of rs1801133 on tHcy exhibits significant variability that can be attributed to effect modification and, to a larger degree, confounding by these covariates. Via statistical modeling, we predicted the covariate-dependent risk of tHcy elevation and hyperhomocysteinemia in a systematic manner.\n\nConclusionswe demonstrated an evidence-based approach that globally assesses the covariate-dependent effect of rs1801133 on tHcy. The results should assist clinicians in interpreting the rs1801133 data from genetic testing for their patients. Such information is also important for the public that increasingly receives genetic data from commercial services without interpretation of its clinical relevance.

genetics

Parallel Computing to Speed up Whole-Genome Bayesian Regression Analyses Using Orthogonal Data Augmentation

1 AbstractBayesian multiple regression methods are widely used in whole-genome analyses to solve the problem that the number p of marker covariates is usually larger than the number n of observations. Inferences from most Bayesian methods are based on Markov chain Monte Carlo methods, where statistics are computed from a Markov chain constructed to have a stationary distribution equal to the posterior distribution of the unknown parameters. In practice, chains of about fifty thousand steps are typically used in whole-genome Bayesian regression analyses, which is computationally intensive. In this paper, we have shown how the sampling of marker effects can be made independent within each step of the chain. This is done by augmenting the marker covariate matrix by adding p new rows to it such that columns of the augmented marker covariate matrix are orthogonal. The phenotypes corresponding to the augmented rows of marker covariate matrix are considered missing. Ideally, the computations at each step of the MCMC chain, can be speeded up by the number k of computer processors up to the number p of markers. Addressing the heavy computational burden associated with Bayesian methods by parallel computing will lead to greater use of these methods.

genetics

Untangling The Gene-Epigenome Networks: Timing Of Epigenetic Regulation Of Gene Expression In Acquired Cetuximab Resistance

BACKGROUNDTargeted therapies specifically act by blocking the activity of proteins that are encoded by genes critical for tumorigenesis. However, most cancers acquire resistance and long-term disease remission is rarely observed. Understanding the time course of molecular changes responsible for the development of acquired resistance could enable optimization of patients treatment options. Clinically, acquired therapeutic resistance can only be studied at a single time point in resistant tumors. To determine the dynamics of these molecular changes, we obtained high throughput omics data weekly during the development of cetuximab resistance in a head and neck cancer in vitro model.\n\nRESULTSAn unsupervised algorithm, CoGAPS, was used to quantify the evolving transcriptional and epigenetic changes. Applying a PatternMarker statistic to the results from CoGAPS enabled novel heatmap-based visualization of the dynamics in these time course omics data. We demonstrate that transcriptional changes result from immediate therapeutic response or resistance, whereas epigenetic alterations only occur with resistance. Integrated analysis demonstrates delayed onset of changes in DNA methylation relative to transcription, suggesting that resistance is stabilized epigenetically.\n\nCONCLUSIONSGenes with epigenetic alterations associated with resistance that have concordant expression changes are hypothesized to stabilize resistance. These genes include FGFR1, which was associated with EGFR inhibitor resistance previously. Thus, integrated omics analysis distinguishes the timing of molecular drivers of resistance. Our findings provide a relevant towards better understanding of the time course progression of changes resulting in acquired resistance to targeted therapies. This is an important contribution to the development of alternative treatment strategies that would introduce new drugs before the resistant phenotype develops.

cancer biology