bioRxiv Science⌕ Search

Biology subjects

BOTTINI, S.

Publications and source records attributed to BOTTINI, S..

4 recordsLinked to original sources

A knowledge graph and topological data analysis framework to disentangle the tomato-multi pathogens complex gene regulatory network

Global population is rapidly increasing, representing a major challenge for food supply, exacerbated by climate change and environmental degradation. Despite the pivotal role of agriculture, plant health and survival are threatened by various biotic stressors. Although how plants respond to each of these individual stresses is well studied, little is known about how they respond to a combination of many of these bio-aggressors occurring together. To tackle this question, first, we built TomTom, a knowledge graph gathering molecular interactions from nine publicly available databases, including transcription factors- or microRNAs-targets, protein-protein interactions, and functional terms. Then, we selected transcriptomics data of tomato subjected to six distinct pathogens and performed an integrative analysis. We found 5561 candidate genes involved in the multi-stress response of tomato. To study how the response is orchestrated, we mapped those genes in TomTom and extracted a comprehensive gene regulatory network (GRN) composed of 71 transcription factors (TF) and 1786 target genes. By estimating the TF activity, we identified 43 TFs responding either specifically to one or multiple bio-aggressors. GRN analyses with a topological data analysis approach allowed to identify 18 clusters of TFs with similar properties, yielding four main configurations localized in specific regions of the GRN. Finally, we found one NAC and four ERF hubs which cooperatively coordinate the tomato response to multiple pathogens. Our findings allowed to study the complex molecular reprogramming in tomato upon interaction with different biotic agents, providing tools scalable to other questions involving tomato molecular interactions and beyond.

systems biology↗

Disentangling plant response to biotic and abiotic stress using HIVE, a novel tool to perform unpaired multi-omics integration

All organisms are subjected to multiple stresses usually occurring at the same time, requiring the activation of the appropriate signalling pathways to respond to all or by prioritizing the response to one stress factor. Plants, as sessile organisms, are particularly impacted by the constantly changing environment that is often unfavourable or even hostile. Because of the experimental complexity of studying the response of one organism to multiple stressors simultaneously, usually experiments are conducted considering one individual stress factor at the time. An alternative consists in performing in silico integration of those data on single stress response. Currently used methods to integrate unpaired experiments consist of performing meta-analysis or finding differentially expressed genes for each condition separately and then selecting the commonly regulated ones. Although these approaches allowed to find valuable results, they mainly identify specific signatures in response to one stress and very few signature responding to multiple stresses and lack those modulated differently in each condition. For this purpose, we developed HIVE (Horizontal Integration analysis using Variational AutoEncoders) to integrate multiple single-stress transcriptomics datasets composed of unpaired experiments. Briefly, we coupled a variational autoencoder, that alleviates batch effects, with a random forest regression and the SHAP explainer to select relevant genes modulated specifically in response to one or multiple stresses. We illustrate the functionality of HIVE to study the transcriptional changes of several different plants namely Arabidopsis thaliana, rice, maize, wheat, grapevine and peanut by collecting publicly available experiments on single stress, either biotic and/or abiotic, and jointly analyse them. HIVE performed better than the differential expression analysis, meta-analysis and the state-of-the-art tool for horizontal integration allowing to identify novel promising candidates responsible for triggering effective defence responses to multiple stresses.

plant biology↗

Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data.

ObjectiveClassification tasks are an open challenge in the field of biomedicine. While several machine-learning techniques exist to accomplish this objective, several peculiarities associated with biomedical data, especially when it comes to omics measurements, prevent their use or good performance achievements. Omics approaches aim to understand a complex biological system through systematic analysis of its content at the molecular level. On the other hand, omics data are heterogeneous, sparse and affected by the classical "curse of dimensionality" problem, i.e. having much fewer observation samples (n) than omics features (p). Furthermore, a major problem with multi- omics data is the imbalance either at the class or feature level. The objective of this work is to study whether feature extraction and/or feature selection techniques can improve the performances of classification machine-learning algorithms on omics measurements. MethodsAmong all omics, metabolomics has emerged as a powerful tool in cancer research, facilitating a deeper understanding of the complex metabolic landscape associated with tumorigenesis and tumor progression. Thus, we selected three publicly available metabolomics datasets, and we applied several feature extraction techniques both linear and non-linear, coupled or not with feature selection methods, and evaluated the performances regarding patient classification in the different configurations for the three datasets. ResultsWe provide general workflow and guidelines on when to use those techniques depending on the characteristics of the data available. For the three datasets, we showed that applying feature selection based on biological previous knowledge improves the performances of the classifiers. Notebook used to perform all analysis are available at: https://github.com/Plant-Net/Metabolomic_project/.

bioinformatics↗

Definition of the effector landscape across 13 Phytoplasma proteomes with LEAPH and EffectorComb

BackgroundCrop pathogens are a major threat to plants health, reducing the yield and quality of agricultural production. Among them, the Candidatus Phytoplasma genus, a group of fastidious phloem-restricted bacteria, can parasite a wide variety of both ornamental and agro-economically important plants. Several aspects of the interaction with the plant host are still unclear but it was discovered that phytoplasmas secrete certain proteins (effectors) responsible for the symptoms associated with the disease. Identifying and characterizing these proteins is of prime importance for globally improving plant health in an environmentally friendly context. ResultsWe challenged the identification of phytoplasmas effectors by developing LEAPH, a novel machine-learning ensemble predictor for phytoplasmas pathogenicity proteins. The prediction core is composed of four models: Random Forest, XGBoost, Gaussian, and Multinomial Naive Bayes. The consensus prediction is achieved by a novel consensus prediction score. LEAPH was trained on 479 proteins from 53 phytoplasmas species, described by 30 features accounting for the biological complexity of these protein sequences. LEAPH achieved 97.49% accuracy, 95.26% precision, and 98.37% recall, ensuring a low false-positive rate and outperforming available state-of-the-art methods for putative effector prediction. The application of LEAPH to 13 phytoplasma proteomes yields a comprehensive landscape of 2089 putative pathogenicity proteins. We identified three classes of these proteins according to different secretion models: "classical", presenting a signal peptide, "classically-like" and "non-classical", lacking the canonical secretion signal. Importantly, LEAPH was able to identify 15 out of 17 known experimentally validated effectors belonging to the three classes. Furthermore, to help the selection of novel candidates for biological validation, we applied the Self-Organizing Maps algorithm and developed a shiny app called EffectorComb. Both tools would be a valuable resource to improve our understanding of effectors in plant-phytoplasmas interactions. ConclusionsLEAPH and EffectorComb app can be used to boost the characterization of putative effectors at both computational and experimental levels and can be employed in other phytopathological models. Both tools are available at https://github.com/Plant-Net/LEAPH-EffectorComb.git.

bioinformatics↗