bioRxiv ScienceSearch

Biology subjects

Song, Y.

Publications and source records attributed to Song, Y..

At least 19 recordsLinked to original sources

Genome-wide association analysis of excessive daytime sleepiness identifies 42 loci that suggest phenotypic subgroups

Excessive daytime sleepiness (EDS) affects 10-20% of the population and is associated with substantial functional deficits. We identified 42 loci for self-reported EDS in GWAS of 452,071 individuals from the UK Biobank, with enrichment for genes expressed in brain tissues and in neuronal transmission pathways. We confirmed the aggregate effect of a genetic risk score of 42 SNPs on EDS in independent Scandinavian cohorts and on other sleep disorders (restless leg syndrome, insomnia) and sleep traits (duration, chronotype, accelerometer-derived sleep efficiency and daytime naps or inactivity). Strong genetic correlations were also seen with obesity, coronary heart disease, psychiatric diseases, cognitive traits and reproductive ageing. EDS variants clustered into two predominant composite phenotypes - sleep propensity and sleep fragmentation - with the former showing stronger evidence for enriched expression in central nervous system tissues, suggesting two unique mechanistic pathways. Mendelian randomization analysis indicated that higher BMI is causally associated with EDS risk, but EDS does not appear to causally influence BMI.

genomics

The genome of the plague-resistant great gerbil reveals species-specific duplication of an MHCII gene

The great gerbil (Rhombomys opimus) is a social rodent living in permanent, complex burrow systems distributed throughout Central Asia, where it serves as the main host of several important vector-borne infectious diseases and is defined as a key reservoir species for plague (Yersinia pestis). Studies from the wild have shown that the great gerbil is largely resistant to plague but the genetic basis for resistance is yet to be determined. Here, we present a highly contiguous annotated genome assembly of great gerbil, covering over 96 % of the estimated 2.47 Gb genome. Comparative genomic analyses focusing on the immune gene repertoire, reveal shared gene losses within TLR gene families (i.e. TLR8, TLR10 and all members of TLR11-subfamily) for the Gerbillinae lineage, accompanied with signs of diversifying selection of TLR7 and TLR9. Most notably, we find a great gerbil-specific duplication of the MHCII DRB locus. In silico analyses suggest that the duplicated gene provides high peptide binding affinity for Yersiniae epitopes. The great gerbil genome provides new insights into the genomic landscape that confers immunological resistance towards plague. The high affinity for Yersinia epitopes could be key in our understanding of the high resistance in great gerbils, putatively conferring a faster initiation of the adaptive immune response leading to survival of the infection. Our study demonstrates the power of studying zoonosis in natural hosts through the generation of a genome resource for further comparative and experimental work on plague survival and evolution of host-pathogen interactions.

genomics

Genome structure and evolution of Antirrhnum majus L.

Snapdragon (Antirrhinum majus L.), a member of Plantaginaceae, is an important model for plant genetics and molecular studies on plant growth and development, transposon biology and self-incompatibility. Here we report a high-quality genome assembly of A. majus cultivated JI7 (A. majus cv.JI7) of a 510 Mb with 37,714 annotated protein-coding genes. The scaffolds covering 97.12% of the assembled genome were anchored on 8 chromosomes. Comparative and evolutionary analyses revealed that Plantaginaceae and Solanaceae diverged from their most recent ancestor around 62 million years ago (MYA). We also revealed the genetic architectures associated with complex traits such as flower asymmetry and self-incompatibility including a unique TCP duplication around 46-49 MYA and a near complete{psi} S-locus of ca.2 Mb. The genome sequence obtained in this study not only provides the first genome sequenced from Plantaginaceae but also bring the popular plant model system of Antirrhinum into a genomic age.

genomics

The draft genome sequence of mandrill (Mandrillus sphinx)

BackgroundMandrill (Mandrillus sphinx) is a primate species which belong to Old World monkey (Cercopithecidae) family. It is closely related to human, serving as model for some human diseases researches. However, genetic researches and genomic resources of mandrill were limited, especially comparing to other primate species.\n\nFindingsHere we sequenced 284 Gb data, providing 96-fold coverage (considering the estimate genome size of 2.9 Gb), to construct a reference genome for mandrill. The assembled draft genome was 2.79 Gb with contig N50 of 20.48 Kb and scaffold N50 of 3.56 Mb. We annotated the mandrill genome to find 43.83% repeat elements, as well as 21,906 protein coding genes. We found good quality of the draft genome and gene annotation by BUSCO analysis which revealed 98% coverage of the BUSCOs.\n\nConclusionsWe established the first draft genome sequence of mandrill, which is valuable resource for future evolutionary and human diseases studies.

genomics

Extensive recoding of dengue virus type 2 specifically reduces replication in primate cells without gain-of-function in Aedes aegypti mosquitoes

Dengue virus (DENV), an arthropod-borne (\"arbovirus\") virus causing a range of human maladies ranging from self-limiting dengue fever to the life-threatening dengue shock syndrome, proliferates well in two different taxa of the Animal Kingdom, mosquitoes and primates. Unexpectedly, mosquitoes and primates have distinct preferences when expressing their genes by translation, e.g. members of these taxa show taxonomic group-specific intolerance to certain codon pairs. This is called \"codon pair bias\". By necessity, arboviruses evolved to delicately balance this fundamental difference in their ORFs. Using the mosquito-borne human pathogen DENV we have undone the evolutionarily conserved genomic balance in its ORF sequence and specifically shifted the encoding preference away from primates. However, this recoding of DENV raised concerns of gain-of-function, namely whether recoding could inadvertently increase fitness for replication in the arthropod vector. Using mosquito cell cultures and two strains of Aedes aegypti we did not observe any increase in fitness in DENV2 variants codon pair deoptimized for humans. This ability to disrupt and control an arboviruss host preference has great promise towards developing the next generation of synthetic vaccines not only for DENV but for other emerging arboviral pathogens such as chikungunya virus and Zika virus.

microbiology

Non-adhesive alginate hydrogels support growth of pluripotent stem cell-derived intestinal organoids

Human intestinal organoids (HIOs) represent a powerful system to study human development and are promising candidates for clinical translation as drug-screening tools or engineered tissue. Experimental control and clinical use of HIOs is limited by growth in expensive and poorly defined tumor-cell-derived extracellular matrices, prompting investigation of synthetic ECM-mimetics for HIO culture. Since HIOs possess an inner epithelium and outer mesenchyme, we hypothesized that adhesive cues provided by the matrix may be dispensable for HIO culture. Here, we demonstrate that alginate, a minimally supportive hydrogel with no inherent cell adhesion properties, supports HIO growth in vitro and leads to HIO epithelial differentiation that is virtually indistinguishable from Matrigel-grown HIOs. Additionally, alginate-grown HIOs mature to a similar degree as Matrigel-grown HIOs when transplanted in vivo, both resembling human fetal intestine. This work demonstrates that purely mechanical support from a simple-to-use and inexpensive hydrogel is sufficient to promote HIO survival and development.

developmental biology

Distinct isoforms of Nrf1 diversely regulate different subsets of its cognate target genes

The single Nrf1 gene has capability to be differentially transcripted alongside with alternative mRNA-splicing and subsequent translation through different initiation signals so as to yield distinct lengths of polypeptide isoforms. Amongst them, three of the most representatives are Nrf1, Nrf1{beta} and Nrf1{gamma}, but the putative specific contribution of each isoform to regulating ARE-driven target genes remains unknown. To address this, we have here established three cell lines on the base of the Flp-In T-REx system, which are allowed for tetracycline-inducibly stable expression of Nrf1, Nrf1{beta} and Nrf1{gamma}. The RNA-Sequencing results have demonstrated that a vast majority of differentially expressed genes (i.e. >90% DEGs detected) were dominantly up-regulated by Nrf1 and/or Nrf1{beta} following induction by tetracycline. By contrast, other DEGs regulated by Nrf1{gamma} were far less than those regulated by Nrf1/{beta} (i.e. ~11% of Nrf1 and 7% of Nrf1{beta}). Further transcriptomic analysis revealed that tetracycline-induced expression of Nrf1{gamma} significantly increased the percentage of down-regulated genes in total DEGs. These statistical data were further validated by quantitative real-time PCR. The experimental results indicate that distinct Nrf1 isoforms make diverse and even opposing contributions to regulating different subsets of target genes, such as those encoding 26S proteasomal subunits and others involved in various biological processes and functions. Collectively, Nrf1{gamma} acts as a major dominant-negative competitor against Nrf1/{beta} activity, such that a number of DEGs regulated by Nrf1/{beta} are counteracted by Nrf1{gamma}.

molecular biology

Recent mixing of Vibrio parahaemolyticus populations

BackgroundHumans have profoundly affected the ocean environment but little is known about anthropogenic effects on the distribution of microbes. Vibrio parahaemolyticus is found in warm coastal waters and causes gastroenteritis in humans and economically significant disease in shrimps.\n\nResultsBased on data from 1,103 genomes, we show that V. parahaemolyticus is divided into four diverse populations, VppUS1, VppUS2, VppX and VppAsia. The first two are largely restricted to the US and Northern Europe, while the others are found worldwide, with VppAsia making up the great majority of isolates in the seas around Asia. Patterns of diversity within and between the populations are consistent with them having arisen by progressive divergence via genetic drift during geographical isolation. However, we find that there is substantial overlap in their current distribution. These observations can be reconciled without requiring genetic barriers to exchange between populations if dispersal between oceans has increased dramatically in the recent past. We found that VppAsia isolates from the US have an average of 1.01% more shared ancestry with VppUS1 and VppUS2 isolates than VppAsia isolates from Asia itself. Based on time calibrated trees of divergence within epidemic lineages, we estimate that recombination affects about 0.017% of the genome per year, implying that the genetic mixture has taken place within the last few decades.\n\nConclusionsThese results suggest that human activity, such as shipping and aquatic products trade, are responsible for the change of distribution pattern of this marine species.

microbiology

First report and multilocus genotyping of Enterocytozoon bieneusi from Tibetan pigs in southwestern China

Enterocytozoon bieneusi is a common intestinal pathogen and a major cause of diarrhea and enteric diseases in a variety of animals. While the E. bieneusi genotype has become better-known, there are few reports on its prevalence in the Tibetan pig. This study investigated the prevalence, genetic diversity, and zoonotic potential of E. bieneusi in the Tibetan pig in southwestern China. Tibetan pig feces (266 samples) were collected from three sites in the southwest of China. Feces were subjected to PCR amplification of the internal transcribed spacer (ITS) region. E. bieneusi was detected in 83 (31.2%) of Tibetan pigs from the three different sites, with 25.4% in Kangding, 56% in Yaan and 26.7% in Qionglai. Age group demonstrated the prevalence of E. bieneusi range from 24.4%(aged 0 to 1 years) to 44.4%(aged 1 to 2 years). Four genotypes of E. bieneusi were identified: two known genotypes EbpC (n=58), Henan-IV (n=24) and two novel genotypes, SCT01 and SCT02 (one of each). Phylogenetic analysis showed these four genotypes clustered to group 1 with zoonotic potential. Multilocus sequence typing (MLST) analysis three microsatellites (MS1, MS3, MS7) and one minisatellite (MS4) revealed 47, 48, 23 and 47 positive specimens were successfully sequenced, and identified ten, ten, five and five genotypes at four loci, respectively. This study indicates the potential danger of E. bieneusi to Tibetan pigs in southwestern China, and offers basic data for preventing and controlling infections.

genetics

A genetically-encoded fluorescent acetylcholine indicator

Acetylcholine (ACh) regulates a diverse array of physiological processes throughout the body, yet cholinergic transmission in the majority of tissues/organs remains poorly understood due primarily to the limitations of available ACh-monitoring techniques. We developed a family of G-protein-coupled receptor activation-based ACh sensors (GACh) with sensitivity, specificity, signal-to-noise ratio, kinetics and photostability suitable for monitoring ACh signals in vitro and in vivo. GACh sensors were validated with transfection, viral and/or transgenic expression in a dozen types of neuronal and non-neuronal cells prepared from several animal species. In all preparations, GACh sensors selectively responded to exogenous and/or endogenous ACh with robust fluorescence signals that were captured by epifluorescent, confocal and/or two-photon microscopy. Moreover, analysis of endogenous ACh release revealed firing pattern-dependent release and restricted volume transmission, resolving two long-standing questions about central cholinergic transmission. Thus, GACh sensors provide a user-friendly, broadly applicable toolbox for monitoring cholinergic transmission underlying diverse biological processes.

neuroscience

GWAS in 446,118 European adults identifies 78 genetic loci for self-reported habitual sleep duration supported by accelerometer-derived estimates

Sleep is an essential homeostatically-regulated state of decreased activity and alertness conserved across animal species, and both short and long sleep duration associate with chronic disease and all-cause mortality1,2. Defining genetic contributions to sleep duration could point to regulatory mechanisms and clarify causal disease relationships. Through genome-wide association analyses in 446,118 participants of European ancestry from the UK Biobank, we discover 78 loci for self-reported sleep duration that further impact accelerometer-derived measures of sleep duration, daytime inactivity duration, sleep efficiency and number of sleep bouts in a subgroup (n=85,499) with up to 7-day accelerometry. Associations are enriched for genes expressed in several brain regions, and for pathways including striatum and subpallium development, mechanosensory response, dopamine binding, synaptic neurotransmission, catecholamine production, synaptic plasticity, and unsaturated fatty acid metabolism. Genetic correlation analysis indicates shared biological links between sleep duration and psychiatric, cognitive, anthropometric and metabolic traits and Mendelian randomization highlights a causal link of longer sleep with schizophrenia.

genetics

iTOP: Inferring the Topology of Omics Data

MotivationIn biology, we are often faced with multiple datasets recorded on the same set of objects, such as multi-omics and phenotypic data of the same tumors. These datasets are typically not independent from each other. For example, methylation may influence gene expression, which may, in turn, influence drug response. Such relationships can strongly affect analyses performed on the data, as we have previously shown for the identification of biomarkers of drug response. Therefore, it is important to be able to chart the relationships between datasets.\n\nResultsWe present iTOP, a methodology to infera topology of relationships between datasets. We base this methodology on the RV coefficient, a measure of matrix correlation, which can be used to determine how much information is shared between two datasets. We extended the RV coefficient for partial matrix correlations, which allows the use of graph reconstruction algorithms, such as the PC algorithm, to infer the topologies. In addition, since multi-omics data often contain binary data (e.g. mutations), we also extended the RV coefficient for binary data. Applying iTOP to pharmacogenomics data, we found that gene expression acts as a mediator between most other datasets and drug response: only proteomics clearly shares information with drug response that is not present in gene expression. Based on this result, we used TANDEM, a method for drug response prediction, to identify which variables predictive of drug response were distinct to either gene expression or proteomics.\n\nAvailabilityAn implementation of our methodology is available in the R package iTOP on CRAN. Additionally, an R Markdown document with code to reproduce all figures is provided as Supplementary Material.\n\nContacta.k.smilde@uva.nl and l.wessels@nki.nl\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Influence of elevated CO2 on development and food utilization of target armyworm Mythimna separata fed on transgenic Bt maize infected by nitrogen-fixing bacteria

Bt crops will face a new ecological risk of reduced effectiveness against target-insect pests owing to the general decrease in exogenous-toxin content in Bt crops grown under elevated CO2. How to deal with this issue may affect the sustainability of transgenic crops as an effective pest management tool especially under future CO2 raising. In this study, azotobacters, as being one potential biological regulator to enhance crops nitrogen utilization efficiency, were selected and the effects of Bt maize and non-Bt maize infected by Azospirillum brasilense and Azotobacter chroococcum on development and food utilization of target Mythimna separate were studied under ambient and elevated CO2. The results indicated that azotobacter infection significantly increased larval life-span, pupal duration, RCR and AD of M. separata, and significantly decreased RGR, ECD and ECI of M. separata fed on Bt maize; There were opposite trends in development and food utilization of M. separata fed on non-Bt maize infected with azotobacters compared with the buffer control regardless of CO2 level. Presumably, the application of azotobacter infection could make Bt maize facing lower field hazards from the target pest of M. separate, and finally improve the resistance of Bt maize against target lepidoptera pests especially under elevated CO2.\n\nSummary statementElevated CO2 effect on development and food utilization of target armyworm Mythimna separata fed on Bt maize infected by azotobacter, Azospirillum brasilense and Azotobacter chroococcum

ecology

GetOrganelle: a simple and fast pipeline for de novo assembly of a complete circular chloroplast genome using genome skimming data

GetOrganelle is a state-of-the-art toolkit to assemble accurate organelle genomes from NGS data. This toolkit recruit organelle-associated reads using a modified \"baiting and iterative mapping\" approach, conducts de novo assembly, filters and disentangles assembly graph, and produces all possible configurations of circular organelle genomes. For 50 published samples, we reassembled the circular plastome in 47 samples using GetOrganelle, but only in 12 samples using NOVOPlasty. In comparison with published/NOVOPlasty plastomes, we demonstrated that GetOrganelle assemblies are more accurate. Moreover, we assembled complete mitogenomes of fungi and animals using GetOrganelle. GetOrganelle is freely released under a GPL-3 license (https://github.com/Kinggerm/GetOrganelle).

bioinformatics

A Highly Efficient and Faithful MDS Patient-Derived Xenotransplantation Model for Pre-Clinical Studies

Comprehensive preclinical studies of Myelodysplastic Syndromes (MDS) have been elusive due to limited ability of MDS stem cells to engraft current immunodeficient murine hosts. We developed a novel MDS patient-derived xenotransplantation model in cytokine-humanized immunodeficient \"MISTRG\" mice that for the first time provides efficient and faithful disease representation across all MDS subtypes. MISTRG MDS patient-derived xenografts (PDX) reproduce patients' dysplastic morphology with multi-lineage representation, including erythro- and megakaryopoiesis. MISTRG MDS-PDX replicate the original sample's genetic complexity and can be propagated via serial transplantation. MISTRG MDS-PDX demonstrate the cytotoxic and differentiation potential of targeted therapeutics providing superior readouts of drug mechanism of action and therapeutic efficacy. Physiologic humanization of the hematopoietic stem cell niche proves critical to MDS stem cell propagation and function in vivo. The MISTRG MDS-PDX model opens novel avenues of research and long-awaited opportunities in MDS research.

cancer biology

Biological and clinical insights from genetics of insomnia symptoms

Insomnia is a common disorder linked with adverse long-term medical and psychiatric outcomes, but underlying pathophysiological processes and causal relationships with disease are poorly understood. Here we identify 57 loci for self-reported insomnia symptoms in the UK Biobank (n=453,379) and confirm their impact on self-reported insomnia symptoms in the HUNT study (n=14,923 cases, 47,610 controls), physician diagnosed insomnia in Partners Biobank (n=2,217 cases, 14,240 controls), and accelerometer-derived measures of sleep efficiency and sleep duration in the UK Biobank (n=83,726). Our results suggest enrichment of genes involved in ubiquitin-mediated proteolysis, phototransduction and muscle development pathways and of genes expressed in multiple brain regions, skeletal muscle and adrenal gland. Evidence of shared genetic factors is found between frequent insomnia symptoms and restless legs syndrome, aging, cardio-metabolic, behavioral, psychiatric and reproductive traits. Evidence is found for a possible causal link between insomnia symptoms and coronary heart disease, depressive symptoms and subjective well-being.\n\nOne Sentence SummaryWe identify 57 genomic regions associated with insomnia pointing to the involvement of phototransduction and ubiquitination and potential causal links to CAD and depression.

genomics

DNA 5-Hydroxymethylcytosines from Cell-free Circulating DNA as Diagnostic Biomarkers for Human Cancers

DNA modifications such as 5-methylcytosines (5mC) and 5-hydroxymethylcytosines (5hmC) are epigenetic marks known to affect global gene expression in mammals(1, 2). Given their prevalence in the human genome, close correlation with gene expression, and high chemical stability, these DNA epigenetic marks could serve as ideal biomarkers for cancer diagnosis. Taking advantage of a highly sensitive and selective chemical labeling technology(3), we report here genome-wide 5hmC profiling in circulating cell-free DNA (cfDNA) and in genomic DNA of paired tumor/adjacent tissues collected from a cohort of 90 healthy individuals and 260 patients recently diagnosed with colorectal, gastric, pancreatic, liver, or thyroid cancer. 5hmC was mainly distributed in transcriptionally active regions coincident with open chromatin and permissive histone modifications. Robust cancer-associated 5hmC signatures in cfDNA were identified with specificity for different cancers. 5hmC-based biomarkers of circulating cfDNA demonstrated highly accurate predictive value for patients with colorectal and gastric cancers versus healthy controls, superior to conventional biomarkers, and comparable to 5hmC biomarkers from tissue biopsies. This new strategy could lead to the development of effective blood-based, minimally-invasive cancer diagnosis and prognosis approaches.

cancer biology

Molecular evolution, diversity and adaptation of H7N9 influenza A viruses in China

A novel H7N9 avian influenza virus has caused five human epidemics in China since 2013. The substantial increase in prevalence and the emergence of antigenically divergent or highly pathogenic (HP) H7N9 strains during the current outbreak raises concerns about the epizootic-potential of these viruses. Here, we investigate the evolution and adaptation of H7N9 by combining publicly available data with newly generated virus sequences isolated in Guangdong between 2015-2017. Phylogenetic analyses show that currently-circulating H7N9 viruses belong to distinct lineages with differing spatial distributions. Using ancestral sequence reconstruction and structural modelling we have identified parallel amino-acid changes on multiple separate lineages. Furthermore, we infer mutations in HA primarily occur at sites involved in receptor-recognition and/or antigenicity. We also identify seven new HP strains, which likely emerged from viruses circulating in eastern Guangdong around March 2016 and is further associated with a high rate of adaptive molecular evolution.

evolutionary biology