bioRxiv Science⌕ Search

Biology subjects

Lopez-Chavarrias, V.

Publications and source records attributed to Lopez-Chavarrias, V..

2 recordsLinked to original sources

Population analysis and host-disease associations of Shiga toxin-producing Escherichia coli from various sources across eleven European countries using whole genome sequencing

Shiga toxin-producing Escherichia coli (STEC) are important foodborne pathogens, able to cause severe disease in humans. In the DiSCoVeR project (https://onehealthejp.eu/jrp-discover/) a STEC inventory from human and non-human sources from 11 European countries was set up and [≥] 3500 strains were sequenced to perform comparative genomics analysis. We used this dataset to assess STEC population structure and to investigate potential associations between genomic features, host reservoirs and symptoms. Most STEC isolates analysed by Whole Genome Sequencing (WGS) in this study were collected between years 2010-2020. An ad hoc pipeline was deployed for a harmonised characterization of the STEC in the database, allowing the determination of serotyping, stx gene subtyping, 7-loci MLST, virulotyping and cgMLST. The results were analysed with Principal Component Analysis (PCoA) in relation with isolation source to assess clustering of STEC subpopulations. When human STEC data were analysed, the PCoA revealed three distinct human STEC subpopulations (STEC_1, STEC_2 and STEC_3), which were further analysed for associations between genomic features, symptoms and variance. The non-human STEC showed a more dispersed distribution, except for one subpopulation with genes linked to specific host species, and some virulence profiles overlapping with the STEC_1 population. In conclusion, our analysis identified distinct STEC subpopulations from human cases, each characterized by specific genetic features and associated with varying proportions of severe disease outcomes. These findings provide novel insights supporting the risk assessment of STEC. Impact statement[This lay summary of your article should be no more than 200 words, and should a) provide a perspective of how this article adds to the literature in the field; b) identify breadth of interest/utility; and c) state the significance of output (incremental or step), in terms of relevance.] This study is based on the establishment of a One Health STEC genomes database, including sequences from isolates of different sources. Most of the isolates had been isolated in the ten-years time span 2010-2020, in 11 different countries, for surveillance and monitoring activities or specific surveys and research purposes. The final dataset included the whole genome sequencing of 3,418 STEC isolates, mainly from human cases of infections. The metadata included the host symptoms, where available, for human STEC strains and the animal source the strains had been isolated from. We set up a pipeline for the harmonized analysis of STEC WGS, called Discover, made available though ARIES webserver or GitHub. The analysis allowed a deep characterization of STEC strains circulating in Europe. We used this resource to assess STEC population structure and to investigate potential associations between genomic features, host reservoirs, and various symptoms associated with STEC infection by PCoA. This analysis highlighted the presence of subpopulation of human STEC associated with specific features. We provide new information useful for risk characterization, as well as a large dataset genome database and associated metadata compiled from STEC strains, representing a valuable resource for the scientific community, enabling further investigations into STEC diversity, evolution, source attribution and public health relevance. Data summaryThe authors confirm all supporting data, including sequence data accession numbers, code and protocols have been provided within the article or through supplementary data files. One supplementary method and five supplementary tables are available with the online version of this article

genomics↗

'PePApipe': a complete bioinformatics analysis pipeline for African Swine Fever Virus genome

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. Viral genome is large and complex, but genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed PePApip, a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, denovo genome assembly and variant calling. Programmed in Phyton, it can be executed locally through bash scripts, or using a SLURM protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structures folders and produces intermediate files that can be used as inputs for further or parallel analyses; also users can enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error and enhances reproducibility and efficiency. This easy-to-use pipeline will facilitate the transition from sequencing to assembly and analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance. Author summaryAfrican swine fever is a devastating viral disease that threatens pig production worldwide, causing severe economic losses and limiting international trade. Tracing and understanding how the virus spreads and evolves are key for preparedness and control of the disease, currently being a major sanitary challenge. However, viral genome analysis is often technically demanding and difficult to standardize, especially for laboratories without specialized bioinformatics expertise. In this study, we present PePApipe, an automated and user-friendly computational pipeline designed to simplify and standardize the complete genomic analysis of African swine fever virus. Starting directly from raw sequencing data, PePApipe guides users through all essential steps, from quality control to genome assembly and variant detection, producing reliable and reproducible results in a short time. By integrating multiple established tools into a single workflow and providing clear intermediate outputs, our approach reduces the risk of user error and increases transparency. Importantly, we designed PePApipe with accessibility in mind, enabling laboratory scientists to perform advanced genomic analyses with minimal computational background. While developed for African swine fever virus, the pipeline can be adapted to other viruses, making it a flexible resource for viral genomics, outbreak investigation, and future data-driven research.

bioinformatics↗