bioRxiv ScienceSearch

Biology subjects

Siren, K.

Publications and source records attributed to Siren, K..

4 recordsLinked to original sources

The pangenome of the fungal pathogen Neonectria neomacrospora

The fungal plant pathogen Neonectria neomacrospora (C. Booth & Samuels) Mantiri & Samuels (Ascomycota, Hypocreales) is a bark parasite causing twig blight, canker, and in severe cases, dieback in fir (Abies spp.). Although often described as a mild pathogen, foresty and phytosanitary agencies have expressed their concern for potential economic impact. Two epidemics caused by this species are known: one from eastern Canada and one current within Northern Europe. We present key genome features of N. neomacrospora, to facilitate the research into the biology of this pathogen. We present the first genome assembly of N. neomacrospora as well as the first pangenome within this genus. The reference genome for N. neomacrospora is a long-read sequenced Danish isolate, while the pangenome is pieced together using additional 60 short-read sequenced strains covering the known geographical distribution of the species, including Europe, North America, and China. The gapless reference genome consist of twelve chromosomes sequenced telomere to telomere to a total length of 37.1 Mb. The mitochondrial genome was assembled and circularised with a length of 22 Kb. The gapless nuclear genome contains a total of 11,291 annotated genes, where 642 only have a hypothetical function, and a 4.3 % repeat content. Two minor chromosomes are enriched in transposable elements, AT content, and effector candidates. Chromosome 12 segregates within the population, indicating an accessory nature. The pangenome compile 15,101 genes, 34% more genes than present in the single isolate reference genome of N. neomacrospora. These genes organise into 13,069 homologous clusters, of which 8,316 clusters are present in all analysed strains, 985 are private to single strains. The British Columbian population branched out before the other populations and are characterized by comparatively larger genomes. The increased genome size can be explained by an expansion of repetitive elements. The comparative analysis finds a higher number of genes with a signal peptide within N. neomacrospora and species within the genus compared to the closely related genera. A species-specific pattern is observed in the carbohydrate-active enzyme repertoire, with a reduced number of polysaccharide lyases, compared to other species within the genus. The CAZymes battery responsible for plant cell wall degradation is similar to that observed in necrotrophic and hemibiotrophic plant pathogenic fungi. The genome size of N. neomacrospora is close to the median size for Ascomycota but is the smallest genome within the Neonectria genus. Comparative analysis revealed significant intraspecies genome size differences between populations explained by a difference in repeat content. Isolates with the smallest genomes formed a monophyletic group consisting of all strains from Europe and Quebec. Based on the field observations, we assume that N. neomacrospora is a hemibiotroph. Our analysis revealed a secretome consistent with a hemibiotrophic lifestyle.

genomics

Population genomics of the emerging forest pathogen Neonectria neomacrospora

The fungal pathogen Neonectria neomacrospora is of increasing concern in Europe where, within the last decade, it has caused substantial damage to forest stands and ornamental trees of the genus Abies (Mill.). Using whole-genome sequencing of a comprehensive collection of isolates, we show the extent of three major clades within N. neomacrospora, which most likely diverged around the end of the last Ice Age. We find it likely that the current European epidemic of N. neomacrospora was founded from a population belonging to the east North American clade. All European isolates (1957-2019) had a common evolutionary history, but substantial and asymmetrical gene flow from the larger American source population could be detected. The European population shows multiple signs of having gone through a bottleneck and subsequent population expansion.

genomics

Rapid discovery of novel prophages using biological feature engineering and machine learning

Prophages are phages that are integrated into bacterial genomes and which are key to understanding many aspects of bacterial biology. Their extreme diversity means they are challenging to detect using sequence similarity, yet this remains the paradigm and thus many phages remain unidentified. We present a novel, fast and generalizing machine learning method based on feature space to facilitate novel prophage discovery. To validate the approach, we reanalyzed publicly available marine viromes and single-cell genomes using our feature-based approaches and found consistently more phages than were detected using current state-of-the-art tools while being notably faster. This demonstrates that our approach significantly enhances bacteriophage discovery and thus provides a new starting point for exploring new biologies.

bioinformatics

Light into the darkness: Unifying the known and unknown coding sequence space in microbiome analyses

Genes of unknown function are among the biggest challenges in molecular biology, especially in microbial systems, where 40%-60% of the predicted genes are unknown. Despite previous attempts, systematic approaches to include the unknown fraction into analytical workflows are still lacking. Here, we propose a conceptual framework and a computational workflow that bridge the known-unknown gap in genomes and metagenomes. We showcase our approach by exploring 415,971,742 genes predicted from 1,749 metagenomes and 28,941 bacterial and archaeal genomes. We quantify the extent of the unknown fraction, its diversity, and its relevance across multiple biomes. Furthermore, we provide a collection of 283,874 lineage-specific genes of unknown function for Cand. Patescibacteria, being a significant resource to expand our understanding of their unusual biology. Finally, by identifying a target gene of unknown function for antibiotic resistance, we demonstrate how we can enable the generation of hypotheses that can be used to augment experimental data.

microbiology