bioRxiv ScienceSearch

Biology subjects

Wang, D.

Publications and source records attributed to Wang, D..

At least 19 recordsLinked to original sources

Cytologic, Genetic, and Proteomic Analysis of a Yellow Leaf Mutant of Sesame (Sesamum indicum L.), Siyl-1

Leaf color mutation in sesame always affects the growth and development of plantlets, and their yield. To clarify the mechanisms underlying leaf color regulation in sesame, we analyzed a yellow-green leaf mutant. Genetic analysis of the mutant selfing revealed 3 phenotypes--YY, light-yellow (lethal); Yy, yellow-green; and yy, normal green--controlled by an incompletely dominant nuclear gene, Siyl-1. In YY and Yy, the number and morphological structure of the chloroplast changed evidently, with disordered inner matter, and significantly decreased chlorophyll content. To explore the regulation mechanism of leaf color mutation, the proteins expressed among YY, Yy, and yy were analyzed. All 98 differentially expressed proteins (DEPs) were classified into 5 functional groups, in which photosynthesis and energy metabolism (82.7%) occupied a dominant position. Our findings provide the basis for further molecular mechanism and biochemical effect analysis of yellow leaf mutants in plants.

genetics

Genome structure and evolution of Antirrhnum majus L.

Snapdragon (Antirrhinum majus L.), a member of Plantaginaceae, is an important model for plant genetics and molecular studies on plant growth and development, transposon biology and self-incompatibility. Here we report a high-quality genome assembly of A. majus cultivated JI7 (A. majus cv.JI7) of a 510 Mb with 37,714 annotated protein-coding genes. The scaffolds covering 97.12% of the assembled genome were anchored on 8 chromosomes. Comparative and evolutionary analyses revealed that Plantaginaceae and Solanaceae diverged from their most recent ancestor around 62 million years ago (MYA). We also revealed the genetic architectures associated with complex traits such as flower asymmetry and self-incompatibility including a unique TCP duplication around 46-49 MYA and a near complete{psi} S-locus of ca.2 Mb. The genome sequence obtained in this study not only provides the first genome sequenced from Plantaginaceae but also bring the popular plant model system of Antirrhinum into a genomic age.

genomics

Dissecting the pharmacological landscape of cancer cells to reveal novel target populations

Personalised medicine has predominantly focused on genetically-altered cancer genes that stratify drug responses, but there is a need to objectively evaluate differential pharmacology patterns at a subpopulation level. Here, we introduce an approach based on unsupervised machine learning to compare the pharmacological response relationships between 327 pairs of cancer therapies. This approach integrated multiple measures of response to identify subpopulations that react differently to inhibitors of the same or different targets to understand mechanisms of resistance and pathway cross-talk. MEK, BRAF, and PI3K inhibitors were shown to be effective as combination therapies for particular BRAF mutant subpopulations. A systematic analysis of preclinical data for a failed phase III trial of selumetinib combined with docetaxel in lung cancer suggests potential indications in urogenital and colorectal cancers with KRAS mutation. This data-informed study exemplifies a method for stratified medicine to identify novel cancer subpopulations, their genetic biomarkers, and effective drug combinations.

bioinformatics

Localization of balanced chromosome translocation breakpoints by long-read sequencing on the Oxford Nanopore platform

Structural variants (SVs) in genomes, including translocations, inversions, insertions, deletions and duplications, remain difficult to be detected reliably by traditional genomic technologies. In particular, balanced translocations and inversions cannot be detected by microarrays since they do not alter chromosome copy numbers; they cannot be reliably detected by short-read sequencing either, since many breakpoints are located within repetitive regions of the genome that are unmappable by short reads. However, the detection and the precise localization of breakpoints at the nucleotide level are important to study the genetic causes in patients carrying balanced translocations or inversions. Long-read sequencing techniques, such as the Oxford Nanopore Technology (ONT), may detect these SVs in a more direct, efficient and accurate manner. In this study, we applied whole-genome long-read sequencing on the Oxford Nanopore GridION sequencer to detect the breakpoints from 6 carriers of balanced translocations and one carrier of inversion, where SVs had initially been detected by karyotyping at the chromosome level. The results showed that all the balanced translocations were detected with [~]10X coverage and were consistent with the karyotyping results. PCR and Sanger sequencing confirmed 8 of the 14 breakpoints to single base resolution, yet other breakpoints cannot be refined to single-base due to their localization at highly repetitive regions or pericentromeric regions, or due to the possible presence of local deletions/duplications. Our results indicate that low-coverage whole-genome sequencing is an ideal tool for the precise localization of most translocation breakpoints and may provide haplotype information on the breakpoint-linked SNPs, which may be widely applied in SV detection, therapeutic monitoring, assisted reproduction technology (ART) and preimplantation genetic diagnosis (PGD).

genetics

Non-driver somatic alteration burden confers good prognosis in non-small cell lung cancer

BackgroundGenomic profiling of patient tumors has linked somatic driver mutations to survival outcomes of non-small cell lung cancer (NSCLC) patients, especially for those receiving targeted therapies. However, it remains unclear whether specific non-driver mutations have any prognostic utility.\n\nMethodsWhole exomes and transcriptomes were measured from NSCLC xenograft models of patients with diverse clinical outcomes. Penalised regression analysis was performed to identify a set of 865 genes associated with patient survival. The number of somatic copy number aberrations, point mutations and associated expression changes within the 865 genes were used to stratify independent NSCLC patient populations, filtered for chemotherapy naive and early-stage. In-depth genomic analysis and functional testing was conducted on the genomic alterations to understand their effect on improving survival.\n\nResultsHigh burden of somatic alterations are associated with longer disease-free survival (HR=0.153, P=1.48x10-4) in NSCLC patients. When somatic alterations burden was integrated with gene expression, we were able to predict prognosis in three independent patient datasets. Patients with high alteration burden could be further stratified based on the presence of immunogenic mutations, revealing another subgroup of patients with even better prognosis (85% with >5 years survival), and associated with cytotoxic T-cell expression. In addition, 95% of these 865 genes lack documented activity relevant to cancer, but are in pathways regulating cell proliferation, motility and immune response were implicated.\n\nConclusionOur results demonstrate that non-driver somatic alterations may influence the outcome of cancer patients by increasing beneficial immune response and inhibiting processes associated to tumorigenesis.

cancer biology

Identification of genome-wide nucleotide sites associated with mammalian virulence in influenza A viruses

MotivationThe virulence of influenza viruses is a complex multigenic trait. Previous studies about the virulence determinants of influenza viruses mainly focused on amino acid sites, ignoring the influence of nucleotide mutations.\n\nResultsWe collected more than 200 viral strains from 21 subtypes of influenza A viruses with virulence in mammals and obtained over 100 mammalian virulence-related nucleotide sites across the genome by computational analysis. Interestingly, 50 of these nucleotide sites only experienced synonymous mutations. Further experiments showed that synonymous mutations in the top two of these nucleotide sites, i.e., PB1-2031 and PB1-633, enhanced the pathogenicity of the viruses in mice. Finally, machine-learning models with accepted accuracy for predicting mammalian virulence of influenza A viruses were built. Overall, this study highlighted the importance of nucleotide mutations, especially synonymous mutations in viral virulence, and provided rapid methods for evaluating the virulence of influenza A viruses. It could be helpful for early warning of newly emerging influenza A viruses.

microbiology

Genome-wide identification and expression specificity analysis of the DNA methyltransferase gene family under adversity stresses in cotton

DNA methylation is an important epigenetic mode of genomic DNA modification that is an important part of maintaining epigenetic content and regulating gene expression. DNA methyltransferases (MTases) are the key enzymes in the process of DNA methylation. Thus far, there has been no systematic analysis the DNA MTases found in cotton. In this study, the whole genome of cotton C5-Mtase coding genes was identified and analyzed using a bioinformatics method based on information from the cotton genome. In this study, 51 DNA MTase genes were identified, of which 8 belonged to G. raimondii (group D), 9 belonged to G. arboretum L. (group A), 16 belonged to G. hirsutum L. (group AD1) and 18 belonged to G. barbadebse L. (group AD2). Systematic evolutionary analysis divided the 51 genes into four subfamilies, including 7 MET homologous proteins, 25 CMT homologous proteins, 14 DRM homologous proteins and 5 DNMT2 homologous proteins. Further studies showed that the DNA MTases in cotton were more phylogenetically conserved. The comparison of their protein domains showed that the C-terminal functional domain of the 51 proteins had six conserved motifs involved in methylation modification, indicating that the protein has a basic catalytic methylation function and the difference in the N-terminal regulatory domains of the 51 proteins divided the proteins into four classes, MET, CMT, DRM and DNMT2, in which DNMT2 lacks an N-terminal regulatory domain. Gene expression in cotton is not the same under different stress treatments. Different expression patterns of DNA MTases show the functional diversity of the cotton DNA methyltransferase gene family. VIGS silenced Gossypium hirsutum l. in the cotton seedling of DNMT2 family gene GhDMT6, after stress treatment the growth condition was better than the control. The distribution of DNA MTases varies among cotton species. Different DNA MTase family members have different genetic structures, and the expression level changes with different stresses, showing tissue specificity. Under salt and drought stress, G. hirsutum L. TM-1 increased the number of genes more than G. raimondii and G. arboreum L. Shixiya 1. The resistance of Gossypium hirsutum L.TM-1 to cold, drought and salt stress was increased after the plants were silenced with GhDMT6 gene.

genomics

Intensive and Specific Feedback Self-control of MicroRNA Targeting Activity

The miRNA pathway consists of three segments - biogenesis, targeting and downstream regulatory effectors. How the cells control their activities remains incompletely understood. This study explored the intrinsically complex miRNA-mRNA targeting relationships, and suggested differential mechanistic control of the three segments. We first analyzed evolutionarily conserved sites for conserved miRNAs in the human transcriptome. Strikingly, AGO1, AGO2 and AGO3 are all among the top 14 mRNAs with highest numbers of unique conserved miRNA sites, and so is ANKRD52, the phosphatase regulatory subunit of the recently identified AGO phosphorylation cycle (AGOs, CSNK1A1, ANKRD52 and PPP6C). The mRNAs for TNRC6, which acts together with loaded AGO to channel miRNA-mediated regulation actions onto specific mRNAs, are also heavily miRNA-targeted. Moreover, mRNAs of the AGO phosphorylation cycle share much more than expected miRNA binding sites. In contrast, upstream miRNA biogenesis mRNAs do not display these characteristics, and neither do the downstream regulatory effector mRNAs. In a word, miRNAs heavily and directly feedback-regulate their targeting machinery mRNAs, but neither upstream biogenesis nor downstream regulatory effector mRNAs. The observation was then confirmed with experimentally determined miRNA-mRNA target relationships. In summary, our exploration of the miRNA-mRNA target relationship uncovers intensive, and specific, feedback auto-regulation of miRNA targeting activity directly by miRNAs themselves, i.e., segment-specific feedback auto-regulation of miRNA pathway. Our results also suggest that the complexity of miRNA-mRNA targeting relationship - a defining feature of miRNA biology - should be a rich source for further functional exploration.

cell biology

Efficient Multivariate Analysis Algorithms for Longitudinal Genome-wide Association Studies

MotivationCurrent dynamic phenotyping system introduces time as an extra dimension to genome-wide association studies (GWAS), which helps to explore the mechanism of dynamical genetic control for complex longitudinal traits. However, existing methods for longitudinal GWAS either ignore the covariance among observations of different time points or encounter computational efficiency issues.\n\nResultsWe herein developed efficient genome-wide multivariate association algorithms (GMA) for longitudinal data. In contrast to existing univariate linear mixed model analyses, the proposed new method has improved statistic power for association detection and computational speed. In addition, the new method can analyze unbalanced longitudinal data with thousands of individuals and more than ten thousand records within a few hours. The corresponding time for balanced longitudinal data is just a few minutes.\n\nAvailability and ImplementationWe wrote a software package to implement the efficient algorithm named GMA (https://github.com/chaoning/GMA), which is available freely for interested users in relevant fields.

bioinformatics

Multilayer network analysis of miRNA and protein expression profiles in breast cancer patients

MiRNAs and proteins play important roles in different stages of tumor development and serve as biomarkers for the early diagnosis of cancer. A new algorithm that combines machine learning algorithms and multilayer complex network analysis is hereby proposed to explore the potential diagnostic values of miRNAs and proteins. XGBoost and random forest algorithms were employed to exclude unrelated miRNAs and proteins, and the most significant candidates were retained for the further analysis. Given these candidates possible functional relationships to one other, a multilayer complex network was constructed to identify miRNAs and proteins that could serve as biomarkers for breast cancer. Proteins and miRNAs that are nodes in the network were subsequently categorized into two network layers considering their distinct functions. Maximal information coefficient (MIC) was applied to assess intralayer and interlayer connection. The betweenness centrality was used as the first measurement of the importance of the nodes within each single layer. To further characterize the interlayer interaction between miRNAs and proteins, the degree of the nodes was chosen as the second measurement to map their signalling pathways. By combining these two measurements into one score and comparing the difference of the same candidate between normal tissue and cancer tissue, this novel multilayer network analysis could be applied to successfully identify molecules associated with breast cancer.

bioinformatics

Unraveling the genetic architecture of grain size in einkorn wheat through linkage and homology mapping, and transcriptomic profiling

HighlightGenome-wide linkage and homology mapping revealed 17 genomic regions through a high-density einkorn wheat genetic map constructed using RAD-seq, and transcription levels of 20 candidate genes were explored using RNA-seq.\n\nAbstractUnderstanding the genetic architecture of grain size is a prerequisite to manipulate the grain development and improve the yield potential in crops. In this study, we conducted a whole genome-wide QTL mapping of grain size related traits in einkorn wheat by constructing a high-density genetic map, and explored the candidate genes underlying QTL through homologous analysis and RNA sequencing. The high-density genetic map spanned 1873 cM and contained 9937 SNP markers assigned to 1551 bins in seven chromosomes. Strong collinearity and high genome coverage of this map were revealed with the physical maps of wheat and barley. Six grain size related traits were surveyed in five agro-climatic environments with 80% or more broad-sense heritability. In total, 42 QTL were identified and assigned to 17 genomic regions on six chromosomes and accounted for 52.3-66.7% of the phenotypic variations. Thirty homologous genes involved in grain development were located in 12 regions. RNA sequencing provided 4959 genes differentially expressed between the two parents. Twenty differentially expressed genes involved in grain size development and starch biosynthesis were mapped to nine regions that contained 26 QTL, indicating that the starch biosynthesis pathway played a vital role on grain development in einkorn wheat. This study provides new insights into the genetic architecture of grain size in einkorn wheat, the underlying genes enables the understanding of grain development and wheat genetic improvement, and the map facilitates the mapping of quantitative traits, map-based cloning, genome assembling and comparative genomics in wheat taxa.

genetics

Quantification and discovery of sequence determinants of protein per mRNA amount in 29 human tissues

Despite their importance in determining protein abundance, a comprehensive catalogue of sequence features controlling protein-to-mRNA (PTR) ratios and a quantification of their effects is still lacking. Here we quantified PTR ratios for 11,575 proteins across 29 human tissues using matched transcriptomes and proteomes. We analyzed the contribution of known sequence determinants of protein synthesis and degradation and 15 novel mRNA and protein sequence motifs that we found by association testing. While the dynamic range of PTR ratios spans more than 2 orders of magnitude, our integrative model predicts PTR ratios at a median precision of 3.2-fold. A reporter assay provided significant functional support for two novel UTR motifs and a proteome-wide competition-binding assay identified motif-specific bound proteins for one motif. Moreover, our direct comparison of protein to RNA levels led to a new metrics of codon optimality. Altogether, this study shows that a large fraction of PTR ratio variance across genes can be predicted from sequence and identified many new candidate post-transcriptional regulatory elements in the human genome.

systems biology

A deep proteome and transcriptome abundance atlas of 29 healthy human tissues

Genome-, transcriptome- and proteome-wide measurements provide valuable insights into how biological systems are regulated. However, even fundamental aspects relating to which human proteins exist, where they are expressed and in which quantities are not fully understood. Therefore, we have generated a systematic, quantitative and deep proteome and transcriptome abundance atlas from 29 paired healthy human tissues from the Human Protein Atlas Project and representing human genes by 17,615 transcripts and 13,664 proteins. The analysis revealed that few proteins show truly tissue-specific expression, that vast differences between mRNA and protein quantities within and across tissues exist and that the expression levels of proteins are often more stable across tissues than those of transcripts. In addition, only ~2% of all exome and ~7% of all mRNA variants could be confidently detected at the protein level showing that proteogenomics remains challenging, requires rigorous validation using synthetic peptides and needs more sophisticated computational methods. Many uses of this resource can be envisaged ranging from the study of gene/protein expression regulation to protein biomarker specificity evaluation to name a few.

systems biology

Effect of Live Attenuated Influenza Vaccine on Pneumococcal Carriage

The widely used nasally-administered Live Attenuated Influenza Vaccine (LAIV) alters the dynamics of naturally occurring nasopharyngeal carriage of Streptococcus pneumoniae in animal models. Using a human experimental model (serotype 6B) we tested two hypotheses: 1) LAIV increased the density of S. pneumoniae in those already colonised; 2) LAIV administration promoted colonisation. Randomised, blinded administration of LAIV or nasal placebo either preceded bacterial inoculation or followed it, separated by a 3-day interval. The presence and density of S. pneumoniae was determined from nasal washes by bacterial culture and PCR. Overall acquisition for bacterial carriage were not altered by prior LAIV administration vs. controls (25/55 [45.5%] vs 24/62 [38.7%] respectively, p=0.46). Transient increase in acquisition was detected in LAIV recipients at day 2 (33/55 [60.0%] vs 25/62 [40.3%] in controls, p=0.03). Bacterial carriage densities were increased approximately 10-fold by day 9 in the LAIV recipients (2.82 vs 1.81 log10 titers, p=0.03). When immunisation followed bacterial acquisition (n=163), LAIV did not change area under the bacterial density-time curve (AUC) at day 14 by conventional microbiology (primary endpoint), but significantly reduced AUC to day 27 by PCR (p=0.03). These studies suggest that LAIV may transiently increase nasopharyngeal density of S. pneumoniae. Transmission effects should therefore be considered in the timing design of vaccine schedules.\n\nTrial registrationThe study was registered on EudraCT (2014-004634-26)\n\nFundingThe study was funded by the Bill and Melinda Gates Foundation and the UK Medical Research Council.

immunology

Co-outbreak of ST37 and a novel ST3006 Klebsiella pneumoniae from multi-site infection in a neonatal intensive care unit, Fuzhou, China

BackgroundThe outbreak of carbapenems resistant Klebsiella pneumoniae (K. pneumoniae) is a serious public health problem, especially in the neonatal intensive care unit (NICU).\n\nMethodsFifteen strains of K. pneumoniae were isolated from seven neonates during June 3-28, 2017 in a NICU. Antimicrobial susceptibility was determined by the Vitek 2 system and micro-broth dilution method. Multi-locus sequence typing (MLST) and pulsed-field gel electrophoresis (PFGE) were used to analyse the genetic relatedness of isolates. Genome sequencing and gene function analyses were performed for investigating pathogenicity and drug resistance and screening genomic islands.\n\nFindingsTwo K. pneumoniae clones were identified from seven neonates, one ST37 strain and another new sequence type ST3006. The ST37 strain exhibited multi-drug resistance genes and resistance to carbapenem. MLST and PFGE showed that 15 strains were divided into three groups, with a high level of homology. Gene sequencing and analysis indicated that KPN1343 harboured 12 resistance genes, 15 genomic islands and 205 reduced virulence genes. KPN1344 harboured four resistance genes, 19 genomic islands and 209 reduced virulence genes.\n\nConclusionCo-outbreak of K. pneumoniae involved two clones, ST36 and ST3006, causing multi-site infection. Genome sequencing and analysis is an effective method for studying bacterial resistance genes and their functions.

microbiology

Long-read sequencing identified a causal structural variant in an exome-negative case and enabled preimplantation genetic diagnosis

For a proportion of individuals judged clinically to have a recessive Mendelian disease, only one pathogenic variant can be found from clinical whole exome sequencing (WES), posing a challenge to genetic diagnosis and genetic counseling. Here we describe a case study, where WES identified only one pathogenic variant for an individual suspected to have glycogen storage disease type Ia (GSD-Ia), which is an autosomal recessive disease caused by bi-allelic mutations in the G6PC gene. Through Nanopore long-read whole-genome sequencing, we identified a 7kb deletion covering two exons on the other allele, suggesting that complex structural variants (SVs) may explain a fraction of cases when the second pathogenic allele is missing from WES on recessive diseases. Both breakpoints of the deletion are within Alu elements, and we designed Sanger sequencing and quantitative PCR assays based on the breakpoints for preimplantation genetic diagnosis (PGD) for the family planning on another child. Four embryos were obtained after in vitro fertilization (IVF), and an embryo without deletion in G6PC was transplanted after PGD and was confirmed by prenatal diagnosis, postnatal diagnosis, and subsequent lack of disease symptoms after birth. In summary, we present one of the first examples of using long-read sequencing to identify causal yet complex SVs in exome-negative patients, which subsequently enabled successful personalized PGD.

genetics

Molecular Detection of H.pylori Antibiotic-Resistant Genes and Bioinformatics Predictive Analysis

To explore the mutation characteristics of H.pylori resistance-related genes to antibiotics of clarithromycin, levofloxacin and metronidazole. 23S rRNA, gyrA, gyrB, rdxA and frxA genes were amplified and sequenced, respectively. Their structural alteration after mutation was predicted using bioinformatics software. In the clarithromycin-resistant strains, the mutation rate in site A2143G was 74.2% (n=23). The mutations in sites C1883T, C2131T and T2179G might cause structural alteration. In the levofloxacin-resistant strains, the mutation rates in 87 (N to K/I) and 91 (D to N/Y/G) of gyrA were 28.6% (n=16) and 12.5% (n =7), respectively. Meanwhile, one of the mutation strains in site 91 was accompanied by D99N variation. Additionally, a D143E mutation was found in one drug-resistant strain. Some changes of tertiary structure occurred after these mutations. The mutation types of RdxA protein consisted of protein truncation caused by premature stop codons (n=26, 33.3%), frameshift mutations (n=8, 10.3%), FMN-binding sites (n=16, 20.5%) and the others (n=11, 14.1%). Predictive analysis showed that mutations in the first three groups and the A118S of the last group could lead to structural alteration. Our study suggested the clarithromycin-resistant sites of H.pylori were mainly located in A2143G of 23S rRNA. C1883T, C2131T and T2179G might also be related to resistance. Levofloxacin resistance was mainly based on the amino acid changes in 87 and 91 sites of gyrA. The new sites D99N and D143E might also be associated with resistance. Metronidazole resistance was related to RdxA protein truncation, frameshift, and FMN binding. The new site A118S might also be linked to drug resistance.

microbiology

Anti-BP180 Autoantibodies Are Present in Stroke and Recognize Human Cutaneous BP180 and BP180-NC16A

BackgroundCurrent evidence has revealed a significant association between bullous pemphigoid (BP) and neurological diseases (ND), including stroke, but the incidence of BP autoantibodies in patients with stroke has not previously been investigated.\n\nObjectiveOur study aims to assess BP antigen-specific antibodies in stroke patients.\n\nMethods100 patients with stroke and 100 healthy controls were randomly selected to measure anti-BP180/230 IgG autoantibodies by enzyme-linked immunosorbent assay (ELISA), salt split indirect immunofluorescence (IIF) and immunoblotting against human cutaneous BP180 and BP180-NC16A.\n\nResultsAnti-BP180 autoantibodies were found in 14(14.0%) patients with stroke and 5(5.0 %) of controls by ELISA (p<0.05). Sera from 13(13.0%) patients with stroke and 3(3.0 %) controls reacted with 180-kDa proteins from human cutis extract (p<0.05). 11(11.0%) of stroke and 2(2.0 %) of control sera recognized the human recombinant full length BP180 and NC16A (p<0.05). The anti-BP180-positive patients were significantly younger than the negative patients in stroke (p<0.001).\n\nLimitationsLongitudinal changes in antibody titers and long-term clinical outcome for a long duration were not fully investigated.\n\nConclusionDevelopment of anti-BP180 autoantibodies occur at a higher frequency after stroke, suggesting BP180 as a shared autoantigen in stroke with BP and providing novel insights into BP pathogenesis in aging.

neuroscience