bioRxiv Science⌕ Search

Biology subjects

Di Gioacchino, A.

Publications and source records attributed to Di Gioacchino, A..

5 recordsLinked to original sources

Deciphering the code of viral-host adaptation through maximum entropy models

Understanding how the genome of a virus evolves depending on the host it infects is an important question that challenges our knowledge about several mechanisms of host-pathogen interactions, including mutational signatures, innate immunity, and codon optimization. A key facet of this general topic is the study of viral genome evolution after a host-jumping event, a topic which has experienced a surge in interest due to the fight against emerging pathogens such as SARS-CoV-2. In this work, we tackle this question by introducing a new method to learn Maximum Entropy Nucleotide Bias models (MENB) reflecting single, di- and tri-nucleotide usage, which can be trained from viral sequences that infect a given host. We show that both the viral family and the host leave a fingerprint in nucleotide usages which MENB models decode. When the task is to classify both the host and the viral family for a sequence of unknown viral origin MENB models outperform state of the art methods based on deep neural networks. We further demonstrate the generative properties of the proposed framework, presenting an example where we change the nucleotide composition of the 1918 H1N1 Influenza A sequence without changing its protein sequence, while manipulating the nucleotide usage, by diminishing its CpG content. Finally we consider two well-known cases of zoonotic jumps, for the H1N1 Influenza A and for the SARS-CoV-2 viruses, and show that our method can be used to track the adaptation to the new host and to shed light on the more relevant selective pressures which have acted on motif usage during this process. Our work has wide-ranging applications, including integration into metagenomic studies to identify hosts for diverse viruses, surveillance of emerging pathogens, prediction of synonymous mutations that effect immunogenicity during viral evolution in a new host, and the estimation of putative evolutionary ages for viral sequences in similar scenarios. Additionally, the computational frame-work introduced here can be used to assist vaccine design by tuning motif usage with fine-grained control. Author summaryIn our research, we delved into the fascinating world of viruses and their genetic changes when they jump from one host to another, a critical topic in the study of emerging pathogens. We developed a novel computational method to capture how viruses change the nucleotide usage of their genes when they infect different hosts. We found that viruses from various families have unique strategies for tuning their nucleotide usage when they infect the same host. Our model could accurately pinpoint which host a viral sequence came from, even when the sequence was vastly different from the ones we trained on. We demonstrated the power of our method by altering the nucleotide usage of an RNA sequence without affecting the protein it encodes, providing a proof-of-concept of a method that can be used to design better RNA vaccines or to fine-tune other nucleic acid-based therapies. Moreover the framework we introduce can help tracking emerging pathogens, predicting synonymous mutations in the adaptation to a new host and estimating how long viral sequences have been evolving in it. Overall, our work sheds light on the intricate interactions between viruses and their hosts.

evolutionary biology↗

Designing molecular RNA switches with Restricted Boltzmann machines

Riboswitches are structured allosteric RNA molecules that change conformation upon metabolite binding, triggering a regulatory response. Here we focus on the de novo design of riboswitch-like aptamers, the core part of the riboswitch undergoing structural changes. We use Restricted Boltzmann machines (RBM) to learn generative models from homologous sequence data. We first verify, on four different riboswitch families, that RBM-generated sequences correctly capture the conservation, covariation and diversity of natural aptamers. The RBM model is then used to design new SAM-I riboswitch aptamers. To experimentally validate the properties of the structural switch in designed molecules, we resort to chemical probing (SHAPE and DMS), and develop a tailored analysis pipeline adequate for high-throughput tests of diverse sequences. We probe a total of 476 RBM-designed and 201 natural sequences. Designed molecules with high RBM scores, with 20% to 40% divergence from any natural sequence, display{approx} 30% success rate of responding to SAM with a structural switch similar to their natural counterparts. We show how the capability of the designed molecules to switch conformation is connected to fine energetic features of their structural components.

bioinformatics↗

Learning the differences: a transfer-learning approach to predict antigen immunogenicity and T-cell receptor specificity

Antigen immunogenicity and the specificity of binding of T-cell receptors to antigens are key properties underlying effective immune responses. Here we propose diffRBM, an approach based on transfer learning and Restricted Boltzmann Machines, to build sequence-based predictive models of these properties. DiffRBM is designed to learn the distinctive patterns in amino acid composition that, one the one hand, underlie the antigens probability of triggering a response, and on the other hand the T-cell receptors ability to bind to a given antigen. We show that the patterns learnt by diffRBM allow us to predict putative contact sites of the antigen-receptor complex. We also discriminate immunogenic and non-immunogenic antigens, antigen-specific and generic receptors, reaching performances that compare favorably to existing sequence-based predictors of antigen immunogenicity and T-cell receptor specificity. More broadly, diffRBM provides a general framework to detect, interpret and leverage selected features in biological data.

bioinformatics↗

Generative and interpretable machine learning for aptamer design and analysis of in vitro sequence selection

Selection protocols such as SELEX, where molecules are selected over multiple rounds for their ability to bind to a target molecule of interest, are popular methods for obtaining binders for diagnostic and therapeutic purposes. With the increasing amount of such high-throughput experimental data available, machine learning techniques have become increasingly popular for molecular datasets analysis. Here, we show that Restricted Boltzmann Machines (RBMs), a two-layer neural network architecture, can successfully be trained on sequence ensembles from SELEX experiments for thrombin aptamers, and used to estimate the fitness of the sequences obtained through the experimental protocol. As a direct consequence, we show that trained RBMs can be exploited to classify as well as generate novel molecules. To confirm our findings, we experimentally verify the generated sequences from RBM.

biophysics↗

sgDI-tector: defective interfering viral genome bioinformatics for detection of coronavirus subgenomic RNAs

Coronavirus RNA-dependent RNA polymerases produce subgenomic RNAs (sgRNAs) that encode viral structural and accessory proteins. User-friendly bioinformatic tools to detect and quantify sgRNA production are urgently needed to study the growing number of next-generation sequencing (NGS) data of SARS-CoV-2. We introduced sgDI-tector to identify and quantify sgRNA in SARS-CoV-2 NGS data. sgDI-tector allowed detection of sgRNA without initial knowledge of the transcription-regulatory sequences. We produced NGS data and successfully detected the nested set of sgRNAs with the ranking M>ORF3a>N>ORF6>ORF7a>ORF8>S>E>ORF7b. We also compared the level of sgRNA production with other types of viral RNA products such as defective interfering viral genomes.

bioinformatics↗