bioRxiv Science⌕ Search

Biology subjects

Lai, B.

Publications and source records attributed to Lai, B..

5 recordsLinked to original sources

Genome-wide H3K9 Acetylation Level Increases with Age-Dependent Senescence of Flag Leaf in Rice (Oryza sativa)

Flag leaf senescence is an important biological process that drives the remobilization of nutrients to the growing organs of rice. Leaf senescence is controlled by genetic information via gene expression and epigenetic modification, but the precise mechanism is as of yet unclear. Here, we analyzed genome-wide acetylated lysine residue 9 of histone H3 (H3K9ac) enrichment by chromatin immunoprecipitation-sequencing (ChIP-seq) and examined its association with transcriptomes by RNA-seq during flag leaf aging in rice (Oryza sativa). We found that genome-wide H3K9 acetylation levels increased with age-dependent senescence in rice flag leaf, and there was a positive correlation between the density and breadth of H3K9ac and gene expression and transcript elongation. A set of 1,249 up-regulated, differentially expressed genes (DEGs) and 996 down-regulated DEGs showing a strong relationship between temporal changes in gene expression and gain/loss of H3K9ac was observed during rice flag leaf aging. We produced a landscape of H3K9 acetylation-modified gene expression targets that includes known senescence-associated genes, metabolism-related genes, as well as miRNA biosynthesis-related genes. Our findings reveal a complex regulatory network of metabolism- and senescence-related pathways mediated by H3K9ac and also elucidate patterns of H3K9ac-mediated regulation of gene expression during flag leaf aging in rice. Significance statementGenome-wide H3K9 acetylation levels increased with age-dependent senescence in rice flag leaf, and positively correlation the density and breadth of H3K9ac with transcript elongation and expression. Identified numerous H3K9 acetylation-modified gene expression targets reveal a complex regulatory network and metabolism-mediated senescence network that are associated with H3K9ac during leaf aging in rice.

plant biology↗

Accurate Protein Function Prediction via Graph Attention Networks with Predicted Structure Information

Experimental protein function annotation does not scale with the fast-growing sequence databases. Only a tiny fraction (<0.1%) of protein sequences in UniProtKB has experimentally determined functional annotations. Computational methods may predict protein function in a high-throughput way, but its accuracy is not very satisfactory. Based upon recent breakthroughs in protein structure prediction and protein language models, we develop GAT-GO, a graph attention network (GAT) method that may substantially improve protein function prediction by leveraging predicted inter-residue contact graphs and protein sequence embedding. Our experimental results show that GAT-GO greatly outperforms the latest sequence- and structure-based deep learning methods. On the PDB-mmseqs testset where the train and test proteins share <15% sequence identity, GAT-GO yields Fmax(maximum F-score) 0.508, 0.416, 0.501, and AUPRC(area under the precision-recall curve) 0.427, 0.253, 0.411 for the MFO, BPO, CCO ontology domains, respectively, much better than homology-based method BLAST (Fmax 0.117,0.121,0.207 and AUPRC 0.120, 0.120, 0.163). On the PDB-cdhit testset where the training and test proteins share higher sequence identity, GAT-GO obtains Fmax 0.637, 0.501, 0.542 for the MFO, BPO, CCO ontology domains, respectively, and AUPRC 0.662, 0.384, 0.481, significantly exceeding the just-published graph convolution method DeepFRI, which has Fmax 0.542, 0.425, 0.424 and AUPRC 0.313, 0.159, 0.193.

bioinformatics↗

Heterogeneity and molecular programming of progenitors for motor neurons and oligodendrocytes

The pMN domain is a restricted domain in the ventral spinal cords, defined by the expression of olig2 gene. The fate determination of pMN progenitors is highly temporally and spatially regulated, with motor neurons and oligodendrocyte progenitor cells (OPCs) developing sequentially. Insight into the heterogeneity and molecular programs of pMN progenitors is currently lacking. With the zebrafish model, we identified multiple states of neural progenitors using single-cell sequencing: proliferating progenitors, common progenitors for both motor neurons and OPCs, and restricted precursors for either motor neurons or OPCs. We found specific molecular programs for neural progenitor fate transition, and manipulations of representative genes in the motor neuron or OPC lineage confirmed their critical role in cell fate determination. The transcription factor NPAS3 is necessary for the development of the OPC lineage and can interact with many known genes associated with schizophrenia. Deciphering progenitor heterogeneity and molecular mechanisms for these transitions will elucidate the formation of complex neuron-glia networks in the central nervous system during development, and understand the basis of neurodevelopmental disorders.

neuroscience↗

Predicting Epigenomic Functions of Genetic Variants in the Context of Neurodevelopment via Deep Transfer Learning

Decoding the regulatory effects of non-coding variants is a key challenge in understanding the mechanisms of gene regulation as well as the genetics of common diseases. Recently, deep learning models have been introduced to predict genome-wide epigenomic profiles and effects of DNA variants, in various cellular contexts, but they were often trained in cell lines or bulk tissues that may not be related to phenotypes of interest. This is particularly a challenge for neuropsychiatric disorders, since the most relevant cell and tissue types are often missing in the training data of such models. To address this issue, we introduce a deep transfer learning framework termed MetaChrom that takes advantage of both a reference dataset - an extensive compendium of publicly available epigenomic data, and epigenomic profiles of cell types related to specific phenotypes of interest. We trained and evaluated our model on a comprehensive set of epigenomic profiles from fetal and adult brain, and cellular models representing early neurodevelopment. MetaChrom predicts these epigenomic features with much higher accuracy than previous methods, and than models without the use of reference epigenomic data for transfer learning. Using experimentally determined regulatory variants from iPS cell-derived neurons, we show that MetaChrom predicts functional variants more accurately than existing non-coding variant scoring tools. By combining genome-wide association study (GWAS) data with MetaChrom predictions, we prioritized 31 SNPs for Schizophrenia (SCZ). These candidate SNPs suggest potential risk genes of SCZ and the biological contexts where they act. In summary, MetaChrom is a general transfer learning framework that can be applied to the study of regulatory functions of DNA sequences and variants in any disease-related cell or tissue types. The software tool is available at https://github.com/bl-2633/MetaChrom and a prediction web server is accessible at https://metachrom.ttic.edu/.

bioinformatics↗

Reversing Transcriptome-Wide Association Studies to improve expression Quantitative Trait Loci associations

1Transcriptome-Wide Association Studies discover SNP effects mediated by gene expression through a two-stage process: a typically small reference panel is used to infer SNP-expression effects, and then these are applied to discover associations between imputed expression and phenotypes. We investigate whether the accuracy of SNP-expression and expression-phenotype associations can be increased by performing inference on both the reference panel and independent GWAS cohorts simultaneously. We develop EMBER (Estimation of Mediated Binary Effects in Regression) to re-estimate these effects using a liability threshold model with an adjustment to variance components accounting for imputed expression from GWAS data. In simulated data with only gene-mediated effects, EMBER more than doubles the performance of SNP-expression linear regression, increasing mean r2 from 0.3 to 0.65 with a gene-mediated variance explained of 0.01. EMBER also improves estimation accuracy when the fraction of cis-SNP variance mediated by genes is as low as 30%. We apply EMBER to genotype and gene expression data in schizophrenia by combining 512 samples from the CommonMind Consortium and 56,081 samples from the Psychiatric Genomic Consortium. We evaluate performance of EMBER in 36 genes suggested by TWAS by concordance of inferred effects with effects reported independently for frontal cortex expression. Applying the EMBER framework to a baseline linear regression model increases performance in 26 out of 36 genes (sign test p-value .0020) with an increase in mean r2 from 0.200 to 0.235.

genomics↗