bioRxiv ScienceSearch

Biology subjects

Zhu, T.

Publications and source records attributed to Zhu, T..

8 recordsLinked to original sources

Cell and tissue type independent age-associated DNA methylation changes are not rare but common

Age-associated DNA methylation changes have been widely reported across many different tissue and cell types. Epigenetic clocks that can predict chronological age with a surprisingly high degree of accuracy appear to do so independently of tissue and cell-type, suggesting that a component of epigenetic drift is cell-type independent. However, the relative amount of age-associated DNAm changes that are specific to a cell or tissue type versus the amount that occurs independently of cell or tissue type is unclear and a matter of debate, with a recent study concluding that most epigenetic drift is tissue-specific. Here, we perform a novel comprehensive statistical analysis, including matched multi cell-type and multi-tissue DNA methylation profiles from the same individuals and adjusting for cell-type heterogeneity, demonstrating that a substantial amount of epigenetic drift, possibly over 70%, is shared between significant numbers of different tissue/cell types. We further show that ELOVL2 is not unique and that many other CpG sites, some mapping to genes in the Wnt and glutamate receptor signaling pathways, are altered with age across at least 10 different cell/tissue types. We propose that while most age-associated DNAm changes are shared between cell-types that the putative functional effect is likely to be tissue-specific.

bioinformatics

5-Hydroxymethylcytosines from Circulating Cell-free DNA as Diagnostic and Prognostic Markers for Hepatocellular Carcinoma

The lack of highly sensitive and specific diagnostic biomarkers is a major contributor to the poor outcomes of patients with hepatocellular carcinoma (HCC), the second-most common cause of cancer deaths worldwide. We sought to develop a clinically convenient and minimally-invasive approach that can be deployed at scale for the sensitive, specific, and highly reliable diagnosis of HCC, and to evaluate the potential prognostic value of this approach. The study cohort comprised of 2,728 subjects, including HCC patients (n = 1,208), controls (n = 965) (572 healthy individuals and 393 patients with benign lesions), as well as patients with chronic hepatitis B infection (CHB) (n =291), liver cirrhosis (LC) (n = 110), and cholangiocarcinoma (CCC) (n = 154), was recruited from three major liver cancer hospitals in Shanghai, China, from July 2016 to November 2017. Circulating cell-free DNA (cfDNA) were collected from plasma samples from these individuals before surgery or any radical treatment. Applying our 5hmC-Seal technique, the summarized 5-hydroxymethylcytosine (5hmC) profiles in cfDNA were obtained. Molecular annotation analysis suggested that the profiled 5hmC loci in cfDNA were enriched with liver tissue-derived regulatory markers (e.g., H3K4me1). We showed that a weighted diagnostic score (wd-score) based on 117 genes detected using the summarized 5hmC profiles in cfDNA accurately distinguished HCC patients from controls (AUC = 95.1%; 95% CI, 93.6-96.5%) in the validation set, markedly outperformed -fetoprotein (AFP) with superior sensitivity. The wd-scores, which not only detected early BCLC stages (e.g., Stage 0: AUC = 96.2%; 95% CI,94.1-98.4%) and small tumors (e.g., < 2 cm: AUC = 95.7%; 95% CI: 93.6-97.7%), also showed high capacity for distinguishing HCC from non-cancer patients with CHB/LC (AUC = 80.2%; 95% CI, 75.8-84.6%). Moreover, the prognostic value of 5hmC markers in cfDNA was evaluated for HCC recurrence, showing that a weighted prognostic score (wp-score) based on 16 marker genes predicted the recurrence risk (HR = 6.67; 95% CI, 2.81-15.82, p < 0.0001) in 555 patients who have been followed up after surgery. In conclusion, we have developed and validated a robust 5hmC-based diagnostic model that can be applied routinely with clinically feasible amount of cfDNA (e.g., from ~2-5 mL of plasma). Applying this new approach in the clinic could significantly improve the clinical outcomes of HCC patients, for example by early detection of those patients with surgically resectable tumors or as a convenient disease surveillance tool for recurrence.

cancer biology

Germline genetics encode the resistance, risk, and lymphatic metastasis of triple-negative breast cancer in the southern Chinese population

Early identification of the risk for triple-negative breast cancer (TNBC) at the asymptomatic phase could lead to better prognosis. Here we developed a machine learning method to quantify systematic impact of all rare germline mutations on each pathway. We collected 106 TNBC patients and 287 elder healthy women controls. The spectra of activity profiles in multiple pathways were mapped and most pathway activities exhibited globally suppressed by the portfolio of individual germline mutations in TNBC patients. Accordingly, all individuals were delineated into two types: A and B. Type A patients could be differentiated from controls (AUC = 0.89) and sensitive to BRCA1/2 damages; Type B patients can be also differentiated from controls (AUC = 0.69) but probably being protected from BRCA1/2 damages. Further we found that Individuals with the lowest activity of selected pathways had extreme high relative risk (up to 21.67 in type A) and increased lymph node metastasis in these patients. Our study showed that genomic DNA contains information of unimaginable pathogenic factors. And this information is in a distributed form that could be applied to risk assessment for more cancer types. SignificanceWe identified individuals who are more susceptible to triple negative breast cancer. Our method performs much better than previous assessments based on BRCA1/2 damages, even polygenic risk scores. We disclosed previously unimaginable pathogens in a distributed form on genome and extended risk prediction to scenarios for other cancers.

cancer biology

The Spectre of Too Many Species

Recent simulation studies examining the performance of Bayesian species delimitation as implemented in the BPP program have suggested that BPP may detect population splits but not species divergences and that it tends to over-split when data of many loci are analyzed. Here we confirm several of these results and provide their mathematical justifications. We point out that the distinction between population and species splits made in the protracted speciation model has no influence on the generation of gene trees and sequence data, which explains why no method can use such data to distinguish between population splits and speciation. We suggest that the the protracted speciation model is unrealistic and its mechanism for assigning species status contradicts prevailing taxonomic practice. We confirm the suggestion, based on simulation, that in the case of speciation with gene flow, Bayesian model selection as implemented in BPP tends to detect population splits when the amount of data (the number of loci) increases so over-splitting is a legitimate concern. We discuss the use of a recently proposed empirical genealogical divergence index (gdi) for species delimitation and illustrate that parameter estimates produced by a full likelihood analysis as implemented in BPP provide much more reliable inference under the gdi than the approximate method PHRAPL. We suggest that the Bayesian model-selection approach is useful for identifying sympatric cryptic species while Bayesian parameter estimation under the multispecies coalescent can be used to implement empirical criteria for determining species status among allopatric populations.

evolutionary biology

Temperature-induced changes in wheat phosphoproteome reveal temperature-regulated interconversion of phosphoforms

Wheat (Triticum ssp.) is one of the most important human food sources. However, this crop is very sensitive to temperature changes. Specifically, processes during wheat leaf, flower and seed development and photosynthesis, which all contribute to the yield of this crop, are affected by high temperature. While this has to some extent been investigated on physiological, developmental and molecular levels, very little is known about early signalling events associated with an increase in temperature. Phosphorylation-mediated signalling mechanisms, which are quick and dynamic, are associated with plant growth and development, also under abiotic stress conditions. Therefore, we probed the impact of a short-term increase in temperature on the wheat leaf and spikelet phosphoproteome. The resulting data set provides the scientific community with a first large-scale plant phosphoproteome under the control of higher ambient temperature, which will be valuable for future studies. Our analyses also revealed a core set of common proteins between leaf and spikelet, suggesting some level of conserved regulatory mechanisms. Furthermore, we observed temperature-regulated interconversion of phosphoforms, which likely impacts protein activity.

plant biology

Genetic load and mutational meltdown in cancer cell populations

ABSRACTLarge and non-recombining genomes are prone to accumulating deleterious mutations faster than natural selection can purge (Mullers ratchet). A possible consequence would then be the extinction of small populations. Relative to most single-cell organisms, cancer cells, with large and non-recombining genomes, could be particularly susceptible to such \"mutational meltdown\". Curiously, deleterious mutations in cancer cells are rarely noticed despite the strong signals in cancer genome sequences. Here, by monitoring single-cell clones from HeLa cell lines, we characterize deleterious mutations that retard cell proliferation. The main mutational events are copy number variations (CNVs), which happen at an extraordinarily high rate of 0.29 events per cell division. The average fitness reduction, estimated to be 18% per mutation, is also very high. HeLa cell populations therefore have very substantial genetic load and, at this level, natural population would likely experience mutational meltdown. We suspect that HeLa cell populations may avoid extinction only after the population size becomes large. Because CNVs are common in most cell lines and cancer tissues, the observations hint at cancer cells vulnerability, which could be exploited by therapeutic strategies.

evolutionary biology

PrimerServer: a high-throughput primer design and specificity-checking platform

SummaryDesigning specific primers for multiple sites across the whole genome is still challenging, especially in species with complex genomes. Here we present PrimerServer, a high-throughput primer design and specificity-checking platform with both web and command-line interfaces. This platform efficiently integrates site selection, primer design, specificity checking and data presentation. In our case study, PrimerServer achieved high accuracy and a fast running speed for a large number of sites, suggesting its potential for molecular biology applications such as molecular breeding or medical testing.\n\nAvailability and ImplementationSource code for PrimerServer is available at https://github.com/billzt/PrimerServer. A demo server is freely accessible at https://primerserver.org, with all major browsers supported.\n\nContactzhangrui@caas.cn or guosandui@caas.cn

bioinformatics

GFF3sort: An efficient tool to sort GFF3 files for tabix indexing

BackgroundThe traditional method of visualizing gene annotation data in JBrowse is converting GFF3 files to JSON format, which is time-consuming. The latest version of JBrowse supports rendering sorted GFF3 files indexed by tabix, a novel strategy that is more convenient than the original conversion process. However, current tools available for GFF3 file sorting have some limitations and their sorting results would lead to erroneous rendering in JBrowse.\n\nResultsWe developed GFF3sort, a script to sort GFF3 files for tabix indexing. Specifically designed for JBrowse rendering, GFF3sort can properly deal with the order of features that have the same chromosome and start position, either by remembering their original orders or by conducting parent-child topology sorting. Based on our test datasets from seven species, GFF3sort produced accurate sorting results with acceptable efficiency compared with currently available tools.\n\nConclusionsGFF3sort is a novel tool to sort GFF3 files for tabix indexing. We anticipate that GFF3sort will be useful to help with genome annotation data processing and visualization.

bioinformatics