bioRxiv Science⌕ Search

Biology subjects

Sielemann, J.

Publications and source records attributed to Sielemann, J..

3 recordsLinked to original sources

Transcription factors mediating regulation of photosynthesis

Photosynthesis by which plants convert carbon dioxide to sugars using the energy of light is fundamental to life as it forms the basis of nearly all food chains. Surprisingly, our knowledge about its transcriptional regulation remains incomplete. Effort for its agricultural optimization have mostly focused on post-translational regulatory processes1-3 but photosynthesis is regulated at the post-transcriptional4 and the transcriptional level5. Stacked transcription factor mutations remain photosynthetically active5,6 and additional transcription factors have been difficult to identify possibly due to redundancy6 or lethality. Using a random forest decision tree-based machine learning approach for gene regulatory network calculation7 we determined ranked candidate transcription factors and validated five out of five tested transcription factors as controlling photosynthesis in vivo. The detailed analyses of previously published and newly identified transcription factors suggest that photosynthesis is transcriptionally regulated in a partitioned, non-hierarchical, interlooped network.

plant biology↗

plASgraph - using graph neural networks to detect plasmid contigs from an assembly graph

Identification of plasmids from sequencing data is an important and challenging problem related to antimicrobial resistance spread and other One-Health issues. In our work, we provide a new architecture for identifying plasmid contigs in fragmented genome assemblies built from short-read data. Unlike previous machine-learning approaches for this problem, which classify individual contigs separately, we employ graph neural networks (GNNs) to include information from the assembly graph. Propagation of information from nearby nodes in the graph allows accurate classification of even short contigs that are difficult to classify based on sequence features or database searches alone. Our new species-agnostic software tool plASgraph outperforms recently developed PlasForest, which uses database searches to supplement sequence-based features. Since our tool does not rely on existing plasmid databases, it is more suitable for classification of contigs in novel species and discovery of previously unknown plasmid sequences. Our tool can also be trained on a specific species, and in that scenario it outperforms mlplasmids trained on the same species. On one hand, our work provides a new, accurate, and easy to use tool for plasmid classification; on the other hand, it serves as a motivation for more widespread use of GNNs in bioinformatics, such as in pangenome sequence analysis, where sequence graphs serve as a fundamental data structure. Availabilityhttps://github.com/cchauve/plASgraph

bioinformatics↗

Local DNA shape is a general principle of transcription factor binding specificity in Arabidopsis thaliana

A genome encodes two types of information, the "what can be made" and the "when and where". The "what" are mostly proteins which perform the majority of functions within living organisms and the "when and where" is the regulatory information that encodes when and where DNA is transcribed. Currently, it is possible to efficiently predict the majority of the protein content of a genome but nearly impossible to predict the transcriptional regulation. This regulation is based upon the interaction between transcription factors and genomic sequences at the site of binding motifs1,2,3. Information contained within the motif is necessary to predict transcription factor binding, however, it is not sufficient4, as experimentally verified binding sites are substantially scarcer than the corresponding binding motif. Thus, it remains challenging to derive regulational information from binding motifs. Here we show that a random forest machine learning approach, which incorporates the 3D-shape of DNA, enhances binding prediction for all 216 tested Arabidopsis thaliana transcription factors and improves the resolution of differential binding by transcription factor family members which share the same binding motif. Our results contribute to the understanding of protein-DNA recognition and demonstrate the extraction of binding site features beyond the binding sequence. We observed that those features were individually weighted for each transcription factor, even if they shared the same binding sequence. We show that the gained insights enable a more robust prediction of binding behavior regarding novel, not-in-genome motif sequences. Understanding transcription factor binding as a combination of motif sequence and motif shape brings us closer to predicting gene expression from promoter sequence.

plant biology↗