bioRxiv ScienceSearch

Biology subjects

Zhu, X.

Publications and source records attributed to Zhu, X..

29 records · Page 2Linked to original sources

Comprehensive, Integrated, and Phased Whole-Genome Analysis of the Primary ENCODE Cell Line K562

K562 is widely used in biomedical research. It is one of three tier-one cell lines of ENCODE and also most commonly used for large-scale CRISPR/Cas9 screens. Although its functional genomic and epigenomic characteristics have been extensively studied, its genome sequence and genomic structural features have never been comprehensively analyzed. Such information is essential for the correct interpretation and understanding of the vast troves of existing functional genomics and epigenomics data for K562. We performed and integrated deep-coverage whole-genome (short-insert), mate-pair, and linked-read sequencing as well as karyotyping and array CGH analysis to identify a wide spectrum of genome characteristics in K562: copy numbers (CN) of aneuploid chromosome segments at high-resolution, SNVs and Indels (both corrected for CN in aneuploid regions), loss of heterozygosity, mega-base-scale phased haplotypes often spanning entire chromosome arms, structural variants (SVs) including small and large-scale complex SVs and non-reference retrotransposon insertions. Many SVs were phased, assembled, and experimentally validated. We identified multiple allele-specific deletions and duplications within the tumor suppressor gene FHIT. Taking aneuploidy into account, we re-analyzed K562 RNA-seq and whole-genome bisulfite sequencing data for allele-specific expression and allele-specific DNA methylation. We also show examples of how deeper insights into regulatory complexity are gained by integrating genomic variant information and structural context with functional genomics and epigenomics data. Furthermore, using K562 haplotype information, we produced an allele-specific CRISPR targeting map. This comprehensive whole-genome analysis serves as a resource for future studies that utilize K562 as well as a framework for the analysis of other cancer genomes.

genomics

VarExp: Estimating variance explained by Genome-Wide GxE summary statistics

Many genomic analyses, such as genome-wide association studies (GWAS) or genome-wide screening for Gene-Environment (GxE) interactions have been performed to elucidate the underlying mechanisms of human traits and diseases. When the analyzed outcome is quantitative, the overall contribution of identified genetic variants to the outcome is often expressed as the percentage of phenotypic variance explained. In practice, this is commonly estimated using individual genotype data. However, using individual-level data faces practical and ethical challenges when the GWAS results are derived in large consortia through meta-analysis of results from multiple cohorts. In this work, we present a R package, \"VarExp\", that allows for the estimation of the percentage of phenotypic variance explained by variants of interest using summary statistics only. Our package allows for a range of models to be evaluated, including marginal genetic effects, GxE interaction effects, and main genetic and interaction effects jointly. Its implementation integrates all recent methodological developments on the topic and does not need external data to be uploaded by users.\n\nThe R source code, tutorial and associated example are available at https://gitlab.pasteur.fr/statistical-genetics/VarExp.git.

bioinformatics

Genome-wide homology analysis reveals new insights into the origin of the wheat B genome

Wheat is a typical allopolyploid with three homoeologous subgenomes (A, B, and D). The ancestors of the subgenomes A and D had been identified, but not for the subgenome B. The goatgrass Aegilops speltoides (genome SS) has been controversially considered a candidate for the ancestor of the wheat B genome. However, the relationship of the Ae. speltoides S genome with the wheat B genome remains largely obscure, which has puzzled the wheat research community for nearly a century. In the present study, the genome-wide homology analysis identified perceptible homology between wheat chromosome 1B and Ae. speltoides chromosome 1S, but not between other chromosomes in the B and S genomes. An Ae. speltoides-originated segment spanning a genomic region of approximately 10.46 Mb was identified on the long arm of wheat chromosome 1B (1BL). The Ae. speltoides-originated segment on 1BL was found to co-evolve with the rest of the B genome in wheat species. Thereby, we conclude that Ae. speltoides had been involved in the origin of the wheat B genome, but should not be considered an exclusive ancestor of this genome. The wheat B genome might have a polyphyletic origin with multiple ancestors involved, including Ae. speltoides. These novel findings provide significant insights into the origin and evolution of the wheat B genome, and will facilitate polyploid genome studies in wheat and other plants as well.

genetics

Local and global chromatin interactions are altered by large genomic deletions associated with human brain development

BackgroundLarge copy number variants (CNVs) in the human genome are strongly associated with common neurodevelopmental, neuropsychiatric disorders such as schizophrenia and autism. Using Hi-C analysis of long-range chromosome interactions and ChIP-Seq analysis of regulatory histone marks we studied the epigenomic effects of the prominent large deletion CNV on chromosome 22q11.2 and also replicated a subset of the findings for the large deletion CNV on chromosome 1q21.1.\n\nResultsWe found that, in addition to local and global gene expression changes, there are pronounced and multilayered effects on chromatin states, chromosome folding and topological domains of the chromatin, that emanate from the large CNV locus. Regulatory histone marks are altered in the deletion proximal regions, and in opposing directions for activating and repressing marks. There are also significant changes of histone marks elsewhere along chromosome 22q and genome wide. Chromosome interaction patterns are weakened within the deletion boundaries and strengthened between the deletion proximal regions. We detected a change in the manner in which chromosome 22q folds onto itself, namely by increasing the long-range contacts between the telomeric end and the deletion proximal region. Further, the large CNV affects the topological domain that is spanning its genomic region. Finally, there is a widespread and complex effect on chromosome interactions genome-wide, i.e. involving all other autosomes, with some of the effect directly tied to the deletion region on 22q11.2.\n\nConclusionsThese findings suggest novel principles of how such large genomic deletions can alter nuclear organization and affect genomic molecular activity.

genetics

Mechanistic view and genetic control of DNA recombination during meiosis

Meiotic recombination is essential for fertility and allelic shuffling. Canonical recombination models fail to capture the observed complexity of meiotic recombinants. Here we revisit these models by analyzing meiotic heteroduplex DNA tracts genome-wide in combination with meiotic DNA double-strand break (DSB) locations. We provide unprecedented support to the synthesis-dependent strand annealing model and establish estimates of its associated template switching frequency and polymerase processivity. We show that resolution of double Holliday junctions (dHJs) is biased toward cleavage of the pair of strands containing newly synthesized DNA near the junctions. The suspected dHJ resolvase Mlh1-3 as well as Mlh1-2, Exo1 and Sgs1 promote asymmetric positioning of crossover intermediates relative to the initiating DSB and bidirectional conversions. Finally, we show that crossover-biased dHJ resolution depends on Mlh1-3, Exo1, Msh5 and to a lesser extent on Sgs1. These properties are likely conserved in eukaryotes containing the ZMM proteins, which includes mammals.

genetics

A large-scale genome-wide enrichment analysis identifies new trait-associated genes, pathways and tissues across 31 human phenotypes

Genome-wide association studies (GWAS) aim to identify genetic factors that are associated with complex traits. Standard analyses test individual genetic variants, one at a time, for association with a trait. However, variant-level associations are hard to identify (because of small effects) and can be difficult to interpret biologically. \"Enrichment analyses\" help address both these problems by focusing on sets of biologically-related variants. Here we introduce a new model-based enrichment analysis method that requires only GWAS summary statistics, and has several advantages over existing methods. Applying this method to interrogate 3,913 biological pathways and 113 tissue-based gene sets in 31 human phenotypes identifies many previously-unreported enrichments. These include enrichments of the endochondral ossification pathway for adult height, the NFAT-dependent transcription pathway for rheumatoid arthritis, brain-related genes for coronary artery disease, and liver-related genes for late-onset Alzheimers disease. A key feature of our method is that inferred enrichments automatically help identify new trait-associated genes. For example, accounting for enrichment in lipid transport genes yields strong evidence for association between MTTP and low-density lipoprotein levels, whereas conventional analyses of the same data found no significant variants near this gene.

genomics

Granatum: a graphical single-cell RNA-seq analysis pipeline for genomics scientists

BackgroundSingle-cell RNA sequencing (scRNA-Seq) is an increasingly popular platform to study heterogeneity at the single-cell level.\n\nComputational methods to process scRNA-Seq have limited accessibility to bench scientists as they require significant amounts of bioinformatics skills.\n\nResultsWe have developed Granatum, a web-based scRNA-Seq analysis pipeline to make analysis more broadly accessible to researchers. Without a single line of programming code, users can click through the pipeline, setting parameters and visualizing results via the interactive graphical interface Granatum conveniently walks users through various steps of scRNA-Seq analysis. It has a comprehensive list of modules, including plate merging and batch-effect removal, outlier-sample removal, gene filtering, geneexpression normalization, cell clustering, differential gene expression analysis, pathway/ontology enrichment analysis, protein-networ interaction visualization, and pseudo-time cell series construction.\n\nConclusionsGranatum enables broad adoption of scRNA-Seq technology by empowering the bench scientists with an easy-to-use graphical interface for scRNA-Seq data analysis. The package is freely available for research use at http://garmiregroup.org/granatum/app

bioinformatics

A probabilistic approach to discovering dynamic full-brain functional connectivity patterns

Recent research shows that the covariance structure of functional magnetic resonance imaging (fMRI) data - commonly described as functional connectivity - can change as a function of the participants cognitive state (for review see [35]). Here we present a Bayesian hierarchical matrix factorization model, termed hierarchical topographic factor analysis (HTFA), for efficiently discovering full-brain networks in large multi-subject neuroimaging datasets. HTFA approximates each subjects network by first re-representing each brain image in terms of the activities of a set of localized nodes, and then computing the covariance of the activity time series of these nodes. The number of nodes, along with their locations, sizes, and activities (over time) are learned from the data. Because the number of nodes is typically substantially smaller than the number of fMRI voxels, HTFA can be orders of magnitude more efficient than traditional voxel-based functional connectivity approaches. In one case study, we show that HTFA recovers the known connectivity patterns underlying a collection of synthetic datasets. In a second case study, we illustrate how HTFA may be used to discover dynamic full-brain activity and connectivity patterns in real fMRI data, collected as participants listened to a story. In a third case study, we carried out a similar series of analyses on fMRI data collected as participants viewed an episode of a television show. In these latter case studies, we found that the HTFA-derived activity and connectivity patterns can be used to reliably decode which moments in the story or show the participants were experiencing. Further, we found that these two classes of patterns contained partially non-overlapping information, such that decoders trained on combinations of activity-based and dynamic connectivity-based features performed better than decoders trained on activity or connectivity patterns alone. We replicated this latter result with two additional (previously developed) methods for efficiently characterizing full-brain activity and connectivity patterns.

neuroscience

Using Single Nucleotide Variations in Cancer Single-Cell RNA-Seq Data for Subpopulation Identification and Genotype-phenotype Linkage Analysis

Despite its popularity, characterization of subpopulations with transcript abundance is subject to a significant amount of noise. We propose to use effective and expressed nucleotide variations (eeSNVs) from scRNA-seq as alternative features for tumor subpopulation identification. We developed a linear modeling framework, SSrGE, to link eeSNVs associated with gene expression. In all the datasets tested, eeSNVs achieve better accuracies than gene expression for identifying subpopulations. Previously validated cancer-relevant genes are also highly ranked, confirming the significance of the method. Moreover, SSrGE is capable of analyzing coupled DNA-seq and RNA-seq data from the same single cells, demonstrating its value in integrating multi-omics single cell techniques. In summary, SNV features from scRNA-seq data have merits for both subpopulation identification and linkage of genotype-phenotype relationship. The method SSrGE is available at https://github.com/lanagarmire/SSrGE.

bioinformatics

Cox-nnet: an artificial neural network Cox regression for prognosis prediction

Artificial neural networks (ANN) are computing architectures with massively parallel interconnections of simple neurons and has been applied to biomedical fields such as imaging analysis and diagnosis. We have developed a new ANN framework called Cox-nnet to predict patient prognosis from high throughput transcriptomics data. In over 10 TCGA RNA-Seq data sets, Cox-nnet achieves a statistically significant increase in predictive accuracy, compared to the other three methods including Cox-proportional hazards (Cox-PH), Random Forests Survival and CoxBoost. Cox-nnet also reveals richer biological information, from both pathway and gene levels. The outputs from the hidden layer node can provide a new approach for survival-sensitive dimension reduction. In summary, we have developed a new method for more accurate and efficient prognosis prediction on high throughput data, with functional biological insights. The source code is freely available at github.com/lanagarmire/cox-nnet.

bioinformatics

Genome-wide association analyses of sleep disturbance traits identify new loci and highlight shared genetics with neuropsychiatric and metabolic traits

Chronic sleep disturbances, associated with cardio-metabolic diseases, psychiatric disorders and all-cause mortality1,2, affect 25-30% of adults worldwide3. While environmental factors contribute importantly to self-reported habitual sleep duration and disruption, these traits are heritable4-9, and gene identification should improve our understanding of sleep function, mechanisms linking sleep to disease, and development of novel therapies. We report single and multi-trait genome-wide association analyses (GWAS) of self-reported sleep duration, insomnia symptoms including difficulty initiating and/or maintaining sleep, and excessive daytime sleepiness in the UK Biobank (n=112,586), with discovery of loci for insomnia symptoms (near MEIS1, TMEM132E, CYCL1, TGFBI in females and WDR27 in males), excessive daytime sleepiness (near AR/OPHN1) and a composite sleep trait (near INADL and HCRTR2), as well as replication of a locus for sleep duration (at PAX-8). Genetic correlation was observed between longer sleep duration and schizophrenia (rG=0.29, p=1.90x10-13) and between increased excessive daytime sleepiness and increased adiposity traits (BMI rG=0.20, p=3.12x10-09; waist circumference rG=0.20, p=2.12x10-07).

genetics