bioRxiv ScienceSearch

Biology subjects

Tang, W.

Publications and source records attributed to Tang, W..

15 recordsLinked to original sources

Evaluation of the causal effect of fibrinogen on incident coronary heart disease via Mendelian randomization

BackgroundFibrinogen is an essential hemostatic factor and cardiovascular disease risk factor. Early attempts at evaluating the causal effect of fibrinogen on coronary heart disease (CHD) and myocardial infraction (MI) using Mendelian randomization (MR) used single variant approaches, and did not take advantage of recent genome-wide association studies (GWAS) or multi-variant, pleiotropy robust MR methodologies.\n\nMethods and FindingsWe evaluated evidence for a causal effect of fibrinogen on both CHD and MI using MR. We used both an allele score approach and pleiotropy robust MR models. The allele score was composed of 38 fibrinogen-associated variants from recent GWAS. Initial analyses using the allele score incorporated data from 11 European-ancestry prospective cohorts to examine incidence CHD and MI. We also applied 2 sample MR methods with data from a prevalent CHD and MI GWAS. Results are given in terms of the hazard ratio (HR) or odds ratio (OR), depending on the study design, and associated 95% confidence interval (CI).\n\nIn single variant analyses no causal effect of fibrinogen on CHD or MI was observed. In multi-variant analyses using incidence CHD cases and the allele score approach, the estimated causal effect (HR) of a 1 g/L higher fibrinogen concentration was 1.62 (CI = 1.12, 2.36) when using incident cases and the allele score approach. In 2 sample MR analyses that accounted for pleiotropy, the causal estimate (OR) was reduced to 1.18 (CI = 0.98, 1.42) and 1.09 (CI = 0.89, 1.33) in the 2 most precise (smallest CI) models, out of 4 models evaluated. In the 2 sample MR analyses for MI, there was only very weak evidence of a causal effect in only 1 out of 4 models.\n\nConclusionsA small causal effect of fibrinogen on CHD is observed using multi-variant MR approaches which account for pleiotropy, but not single variant MR approaches. Taken together, results indicate that even with large sample sizes and multi-variant approaches MR analyses still cannot exclude the null when estimating the causal effect of fibrinogen on CHD, but that any potential causal effect is likely to be much smaller than observed in epidemiological studies.\n\nAuthor SummaryInitial Mendelian Randomization (MR) analyses of the causal effect of fibrinogen on coronary heart disease (CHD) utilized single variants and did not take advantage of modern, multivariant approaches. This manuscript provides an important update to these initial analyses by incorporating larger sample sizes and employing multiple, modern multi-variant MR approaches to account for pleiotropy. We used incident cases to perform a MR study of the causal effect of fibrinogen on incident CHD and the nested outcome of myocardial infarction (MI) using an allele score approach. Then using data from a case-control genome-wide association study for CHD and MI we performed two sample MR analyses with multiple, pleiotropy robust approaches. Overall, the results indicated that associations between fibrinogen and CHD in observational studies are likely upwardly biased from any underlying causal effect. Single variant MR approaches show little evidence of a causal effect of fibrinogen on CHD or MI. Multi-variant MR analyses of fibrinogen on CHD indicate there may be a small positive effect, however this result needs to be interpreted carefully as the 95% confidence intervals were still consistent with a null effect. Multi-variant MR approaches did not suggest evidence of even a small causal effect of fibrinogen on MI.

genetics

New human chromosomal safe harbor sites for genome engineering with CRISPR/Cas9, TAL effector and homing endonucleases

Safe Harbor Sites (SHS) are genomic locations where new genes or genetic elements can be introduced without disrupting the expression or regulation of adjacent genes. We have identified 35 potential new human SHS in order to substantially expand SHS options beyond the three widely used canonical human SHS, AAVS1, CCR5 and hROSA26. All 35 potential new human SHS and the three canonical sites were assessed for SHS potential using 9 different criteria weighted to emphasize safety that were broader and more genomics-based than previous efforts to assess SHS potential. We then systematically compared and rank-ordered our 35 new sites and the widely used human AAVS1, hROSA26 and CCR5 sites, then experimentally validated a subset of the highly ranked new SHS together versus the canonical AAVS1 site. These characterizations included in vitro and in vivo cleavage-sensitivity tests; the assessment of population-level sequence variants that might confound SHS targeting or use for genome engineering; homology-dependent and -independent, SHS-targeted transgene integration in different human cell lines; and comparative transgene integration efficiencies at two new SHS versus the canonical AAVS1 site. Stable expression and function of new SHS-integrated transgenes were demonstrated for transgene-encoded fluorescent proteins, selection cassettes and Cas9 variants including a transcription transactivator protein that were shown to drive large deletions in a PAX3/FOXO1 fusion oncogene and induce expression of the MYF5 gene that is normally silent in human rhabdomyosarcoma cells. We also developed a SHS genome engineering toolkit to enable facile use of the most extensively characterized of our new human SHS located on chromosome 4p. We anticipate our newly identified human SHS, located on 16 chromosomes including both arms of the human X chromosome, will be useful in enabling a wide range of basic and more clinically-oriented human gene editing and engineering.

genomics

bayNorm: Bayesian gene expression recovery, imputation and normalisation for single cell RNA-sequencing data

Normalisation of single cell RNA sequencing (scRNA-seq) data is a prerequisite to their interpretation. The marked technical variability and high amounts of missing observations typical of scRNA-seq datasets make this task particularly challenging. Here, we introduce bayNorm, a novel Bayesian approach for scaling and inference of scRNA-seq counts. The methods likelihood function follows a binomial model of mRNA capture, while priors are estimated from expression values across cells using an empirical Bayes approach. We demonstrate using publicly-available scRNA-seq datasets and simulated expression data that bayNorm allows robust imputation of missing values generating realistic transcript distributions that match single molecule FISH measurements. Moreover, by using priors informed by dataset structures, bayNorm improves accuracy and sensitivity of differential expression analysis and reduces batch effect compared to other existing methods. Altogether, bayNorm provides an efficient, integrated solution for global scaling normalisation, imputation and true count recovery of gene expression measurements from scRNA-seq data.

bioinformatics

Genetic Architecture of Collective Behaviors in Zebrafish

Collective behaviors of groups of animals, such as schooling and shoaling of fish, are central to species survival, but genes that regulate these activities are not known. Here we parsed collective behavior of groups of adult zebrafish using computer vision and unsupervised machine learning into a set of highly reproducible, unitary, several hundred millisecond states and transitions, which together can account for the entirety of relative positions and postures of groups of fish. Using CRISPR-Cas9 we then targeted for knockout 35 genes associated with autism and schizophrenia. We found mutations in three genes had distinctive effects on the amount of time spent in the specific states or transitions between states. Mutation in immp2l (inner mitochondrial membrane peptidase 2-like gene) enhances states of cohesion, so increases shoaling; mutation in in the Nav1.1 sodium channel, scn1lab+/- causes the fish to remain scattered without evident social interaction; and mutation in the adrenergic receptor, adra1aa-/-, keeps fish close together and retards transitions between states, leaving fish motionless for long periods. Motor and visual functions seemed relatively well-preserved. This work shows that the behaviors of fish engaged in collective activities are built from a set of stereotypical states. Single gene mutations can alter propensities to collective actions by changing the proportion of time spent in these states or the tendency to transition between states. This provides an approach to begin dissection of the molecular pathways used to generate and guide collective actions of groups of animals.

animal behavior and cognition

Identification of genes affecting saturated fat acid content in Elaeis guineensis by genome-wide association analysis

Oil palm is the highest yielding oil crop per unit area worldwide. Unfortunately, palm oil is often considered unhealthy. In particular, palmic acid (C16:0) is a major component of palm oil. In this study a total of 1 261 501 SNP markers were produced in a diversity panel of 200 oil palm individuals. Oil content in this population varied from 29.8% to 70.3%, palmic acid varied from 31.3% to 48.8%, and oleic acid varied from 31.3% to 50.1%. We identified 274 SNP markers significantly associated with fatty acid compositions; 44 candidate genes in the flanking regions of these SNPs were involved in fatty acid biosynthesis and metabolism. Among them, two acyl-ACP thioesterase B genes had differential expression patterns between the mesocarp and kernel, tissues which show different oil profiles in oil palm (high palmic acid and high lauric acid respectively). Overexpression of both genes caused a significant increase in palmic acid content, while overexpression of the EgFatB2 gene also caused an accumulation of lauric acid and myristic acid. Our research provides genome-wide SNPs, a set of markers significantly associated with fatty acid content, and validated candidate genes for future targeted breeding of lower saturated fat content in palm oil.

plant biology

An RNAi screen in human cell lines reveals conserved DNA damage repair pathways that mitigate formaldehyde sensitivity

Formaldehyde is a ubiquitous DNA damaging agent, with human exposures occuring from both exogenous and endogenous sources. Formaldehyde can also form DNA-protein crosslinks and is representative of other such DNA damaging agents including ionizing radiation, metals, aldehydes, chemotherapeutics, and cigarette smoke. In order to identify genetic determinants of cell proliferation in response to continuous formaldehyde exposure, we quantified cell proliferation after siRNA-depletion of a comprehensive array of over 300 genes representing all of the major DNA damage response pathways. Three unrelated human cell lines (SW480, U-2 OS and GM00639) were used to identify common or cell line-specific mechanisms. Four cellular pathways were determined to mitigate formaldehyde toxicity in all three cell lines: homologous recombination, double-strand break repair, ionizing radiation response, and DNA replication. Differences between cell lines were further investigated by using exome sequencing and Cancer Cell Line Encyclopedia genomic data. Our results reveal major genetic determinants of formaldehyde toxicity in human cells and provide evidence for the conservation of these formaldehyde responses between human and budding yeast.

cell biology

Single-cell phenotyping and RNA sequencing reveal novel patterns of gene expression heterogeneity and regulation during growth and stress adaptation in a unicellular eukaryote

Cell-to-cell variability is central for microbial populations and contributes to cell function, stress adaptation and drug resistance. Gene-expression heterogeneity underpins this variability, but has been challenging to study genome-wide. Here, we report an integrated approach for imaging of individual fission yeast cells followed by single-cell RNA sequencing (scRNA-seq) and novel Bayesian normalisation. We analyse >2000 single cells and >700 matching RNA controls in various environmental conditions and identify sets of highly variable genes. Combining scRNA-seq with cell-size measurements provides unique insights into genes regulated during cell growth and division in single cells, including genes whose expression does not scale with cell size. We further analyse the heterogeneity and dynamics of gene expression during adaptive and acute responses to changing environments. Entry into stationary phase is preceded by a gradual, synchronised adaptation in gene regulation, followed by highly variable gene expression when growth decreases. Conversely, a sudden and acute heat-shock leads to a stronger and coordinated response and adaptation across cells. This analysis reveals that the extent and dynamics of global gene-expression heterogeneity is regulated in response to different physiological conditions within populations of a unicellular eukaryote. In summary, this works illustrates the potential of combined transcriptomics and imaging analysis in single cells to provide comprehensive and unbiased mechanistic understanding of cell-to-cell variability in microbial communities.

genomics

QTL mapping of natural variation reveals that the developmental regulator bruno reduces tolerance to P-element transposition in the Drosophila female germline

Transposable elements (TEs) are obligate genetic parasites that propagate in host genomes by replicating in germline nuclei, thereby ensuring transmission to offspring. This selfish replication not only produces deleterious mutations---in extreme cases, TE mobilization induces genotoxic stress that prohibits the production of viable gametes. Host genomes could reduce these fitness effects in two ways: resistance and tolerance. Resistance to TE propagation is enacted by germline specific small-RNA-mediated silencing pathways, such as the piRNA pathway, and is studied extensively. However, it remains entirely unknown whether host genomes may also evolve tolerance, by desensitizing gametogenesis to the harmful effects of TEs. In part, the absence of research on tolerance reflects a lack of opportunity, as small-RNA-mediated silencing evolves rapidly after a new TE invades, thereby masking existing variation in tolerance.\n\nWe have exploited the recent the historical invasion of the Drosophila melanogaster genome by P-element DNA transposons in order to study tolerance of TE activity. In the absence of piRNA-mediated silencing, the genotoxic stress imposed by P-elements disrupts oogenesis, and in extreme cases leads to atrophied ovaries that completely lack germline cells. By performing QTL-mapping on a panel of recombinant inbred lines (RILs) that lack piRNA-mediated silencing of P- elements, we uncovered multiple QTL that are associated with differences in tolerance of oogenesis to P-element transposition. We localized the most significant QTL to a small 230 Kb euchromatic region, with the LOD peak occurring in the Bruno locus, which codes for a critical and well-studied developmental regulator of oogenesis. We further demonstrate that multiple bruno loss-of-function alleles are strong dominant suppressors of ovarian atrophy, allowing for the development of mature egg-chambers in the face of P-element activity. Genetic and cytological analyses suggest that bruno tolerance is explained by enhanced retention of germline stem cells in dysgenic ovaries, which are typically lost due to DNA damage. Our observations reveal segregating variation in TE tolerance for the first time, and implicate gametogenic regulators as a source of tolerant variants in natural populations.

genetics

FERONIA’s sensing of cell wall pectin activates ROP GTPase signaling in Arabidopsis

Plant cells need to monitor the cell wall dynamic to control the wall homeostasis required for a myriad of processes in plants, but the mechanisms underpinning cell wall sensing and signaling in regulating these processes remain largely elusive. Here, we demonstrate that receptor-like kinase FERONIA senses the cell wall pectin polymer to directly activate the ROP6 GTPase signaling pathway that regulates the formation of the cell shape in the Arabidopsis leaf epidermis. The extracellular malectin domain of FER directly interacts with de-methylesterified pectin in vivo and in vitro. Both loss-of-FER mutations and defects in the pectin biosynthesis and de-methylesterification caused changes in pavement cell shape and ROP6 signaling. FER is required for the activation of ROP6 by de-methylesterified pectin, and physically and genetically interacts with the ROP6 activator, RopGEF14. Thus, our findings elucidate a cell wall sensing and signaling mechanism that connects the cell wall to cellular morphogenesis via the cell surface receptor FER.

plant biology

Identification of piRNA binding sites reveals the Argonaute regulatory landscape of the C. elegans germline

piRNAs (Piwi-interacting small RNAs) engage Piwi Argonautes to silence transposons and promote fertility in animal germlines. Genetic and computational studies have suggested that C. elegans piRNAs tolerate mismatched pairing and in principle could target every transcript. Here we employ in vivo cross-linking to identify transcriptome-wide interactions between piRNAs and target RNAs. We show that piRNAs engage all germline mRNAs and that piRNA binding follows microRNA-like pairing rules. Targeting correlates better with binding energy than with piRNA abundance, suggesting that piRNA concentration does not limit targeting. In mRNAs silenced by piRNAs, secondary small RNAs accumulate at the center and ends of piRNA binding sites. In germline-expressed mRNAs, however, targeting by the CSR-1 Argonaute correlates with reduced piRNA binding density and suppression of piRNA-associated secondary small RNAs. Our findings reveal physiologically important and nuanced regulation of individual piRNA targets and provide evidence for a comprehensive post transcriptional regulatory step in germline gene expression.

genetics

Multiethnic Meta-analysis Identifies New Loci for Pulmonary Function

Nearly 100 loci have been identified for pulmonary function, almost exclusively in studies of European ancestry populations. We extend previous research by meta-analyzing genome-wide association studies of 1000 Genomes imputed variants in relation to pulmonary function in a multiethnic population of 90,715 individuals of European (N=60,552), African (N=8,429), Asian (N=9,959), and Hispanic/Latino (N=11,775) ethnicities. We identified over 50 novel loci at genome-wide significance in ancestry-specific and/or multiethnic meta-analyses. Recent fine mapping methods incorporating functional annotation, gene expression, and/or differences in linkage disequilibrium between ethnicities identified potential causal variants and genes at known and newly identified loci. Sixteen of the novel genes encode proteins with predicted or established drug targets, including KCNK2 and CDK12.

genetics

CTCF, WAPL and PDS5 proteins control the formation of TADs and loops by cohesin

Mammalian genomes are organized into compartments, topologically-associating domains (TADs) and loops to facilitate gene regulation and other chromosomal functions. Compartments are formed by nucleosomal interactions, but how TADs and loops are generated is unknown. It has been proposed that cohesin forms these structures by extruding loops until it encounters CTCF, but direct evidence for this hypothesis is missing. Here we show that cohesin suppresses compartments but is essential for TADs and loops, that CTCF defines their boundaries, and that WAPL and its PDS5 binding partners control the length of chromatin loops. In the absence of WAPL and PDS5 proteins, cohesin passes CTCF sites with increased frequency, forms extended chromatin loops, accumulates in axial chromosomal positions (vermicelli) and condenses chromosomes to an extent normally only seen in mitosis. These results show that cohesin has an essential genome-wide function in mediating long-range chromatin interactions and support the hypothesis that cohesin creates these by loop extrusion, until it is delayed by CTCF in a manner dependent on PDS5 proteins, or until it is released from DNA by WAPL.

genomics

Meta-analysis of exome array data identifies six novel genetic loci for lung function

Over 90 regions of the genome have been associated with lung function to date, many of which have also been implicated in chronic obstructive pulmonary disease (COPD). We carried out meta-analyses of exome array data and three lung function measures: forced expiratory volume in one second (FEV1), forced vital capacity (FVC) and the ratio of FEV1 to FVC (FEV1/FVC). These analyses by the SpiroMeta and CHARGE consortia included 60,749 individuals of European ancestry from 23 studies, and 7,721 individuals of African Ancestry from 5 studies in the discovery stage, with follow-up in up to 111,556 independent individuals. We identified significant (P<2{middle dot}8x10-7) associations with six SNPs: a nonsynonymous variant in RPAP1, which is predicted to be damaging, three intronic SNPs (SEC24C, CASC17 and UQCC1) and two intergenic SNPs near to LY86 and FGF10. eQTL analyses found evidence for regulation of gene expression at three signals and implicated several genes including TYRO3 and PLAU. Further interrogation of these loci could provide greater understanding of the determinants of lung function and pulmonary disease.

genetics

Tumor Origin Detection with Tissue-Specific miRNA and DNA methylation Markers

MotivationCancer of unknown primary origin constitutes 3-5% of all human malignancies. Patients with these carcinomas present with metastases without an established primary site, which may not be found even by thorough histological search methods. Patients with cancer of unknown primary origin always have poor prognosis and hardly have efficient treatment since most cancers respond well to specific chemotherapy or hormone drugs. Many studies have proposed classifiers based on miRNAs or mRNAs to predict the tumor origins, but few study focus on high-dimensional DNA methylation profiles.\n\nResultsWe introduced three classifiers with novel feature selection algorithm combined with random forest to effectively identify highly tissue-specific epigenetics biomarkers such as microRNAs and CpG sites, which can help us predict the origin site of tumors. This algorithm, incorporating differential analysis and descending dimension algorithm, was applied on 14 histological tissues and over 5000 samples based on miRNA expression and DNA methylation profiles to assign given primary tumor to its origin tissue. Our study shows all of these three classifiers have an overall accuracy of 87.78% (72.55%-97.54%) based on miRNA datasets and an accuracy of 96.43% (MRMD: 87.85%-99.76%) or 97.06% (PCA: 92.44%-100%) based on DNA methylation datasets on predicting the origin of tumors and suggests that the biomarkers we selected can efficiently predict the origin of tumors and allow the clinicians to avoid adjuvant systemic therapy or to choose less aggressive therapeutic options. We also developed a user-friendly webserver which enables users to predict the origin site of tumors by uploading the miRNAs expression or DNA methylation profiles of those cancers.\n\nAvailabilityThe webserver, data, and code are accessible free of charge at http://server.malab.cn/MMCOP/\n\nContactzouquan@nclab.net\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Contribution of mobile elements to the uniqueness of human genome with more than 15,000 human-specific insertions

Mobile elements (MEs) collectively constituted to at least 51% of the human genome. Due to their past incremental accumulation and ongoing DNA transposition of members from certain subfamilies, MEs serve as a significant source for both inter- and intra-species genetic diversity during primate and human evolution. Since MEs can exert direct impact on gene function via a plethora of mechanism, it is believed that the ME-derived genetic diversity has contributed to the phenotypic differences between human and non-human primates, as well as among human populations and individuals. To define the specific contribution of MEs in making Human sapiens as a biologically unique species, we aim to compile a complete list of MEs that are only uniquely present in the human genome, i.e., human-specific MEs (HS-MEs).\n\nBy making use of the most recent reference genome sequences for human and many other primates and a unbiased more robust and integrative multi-way comparative genomic approach, we identified a total of 15,463 HS-MEs. This list of HS-MEs represents a 120% increase from prior studies with over 8,000 being newly identified as HS-MEs. Collectively, these ~15,000 HS-MEs have contributed to a total of 15 million base pair (Mbp) sequence increase through insertion, generation of target site duplications, and transductions, as well as a 0.5 Mbp sequence loss via insertion- mediated deletions, leading to a net total of 14.5 Mbp genome size increase. Other new observations made with these HS-MEs include: 1) identification of several additional ME subfamilies with significant transposition activities not visible with prior smaller datasets (e.g. L1HS, L1PA2, and HERV-K); 2) A clear similarity of the retrotransposition mechanism among L1, Alus, and SVAs that is distinct from HERVs based on the pre- integration site sequence motifs; 3) Y-chromosome as a strikingly hot target for HS-MEs, particularly for LTRs, which showed an insertion rate 15 times higher than the genome average; 4) among the ME types, SVAs seem to show a very strong bias in inserting into existing SVAs. Among the HS-MEs, more than 8,000 elements were integrated into the vicinity of ~4900 unique genes, in regions including CDS, untranslated exon regions, promoters, and introns of protein coding genes, as well as promoters and exons of non- coding RNAs. In seven cases, MEs participate in protein coding. Furthermore, 1,213 HS-MEs contributed to a total of 3,124 experimentally identified binding sites for 146 of the 161 transcriptional factors in association with 622 genes. All these data suggest that these HS-MEs, despite being very young, already showed sufficient sign for their participation in gene function via regulation of transcription, splicing, and protein coding, with more potential for future participation.\n\nIn conclusion, our results demonstrate that the amount of MEs uniquely occurred in the human genome is much higher than previously known, and we predict that the same is true regarding their impact on human genome evolution and function. The comprehensive list of HS-MEs provides an important reference resource for studying the impact of DNA transposition in human genome evolution and gene function.

evolutionary biology