bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

Complete genome of the Medicago anthracnose fungus, Colletotrichum destructivum, reveals a mini-chromosome-like region within a core chromosome.

Colletotrichum destructivum (Cd) is a phytopathogenic fungus causing significant economic losses on forage legume crops (Medicago and Trifolium species) worldwide. To gain insights into the genetic basis of fungal virulence and host specificity, we sequenced the genome of an isolate from M. sativa using long-read (PacBio) technology. The resulting genome assembly has a total length of 51.7 Mb and comprises 10 core chromosomes and two accessory chromosomes, all of which were sequenced from telomere to telomere. A total of 15,631 gene models were predicted, including genes encoding potentially pathogenicity-related proteins such as candidate secreted effectors (484), secondary metabolism key enzymes (110) and carbohydrate-active enzymes (619). Synteny analysis revealed extensive structural rearrangements in the genome of Cd relative to the closely-related Brassicaceae pathogen, C. higginsianum. In addition, a 1.2 Mb species-specific region was detected within the largest core chromosome of Cd that has all the characteristics of fungal accessory chromosomes (transposon-rich, gene-poor, distinct codon usage), providing evidence for exchange between these two genomic compartments. This region was also unique in having undergone extensive intra-chromosomal segmental duplications. Our findings provide insights into the evolution of accessory regions and possible mechanisms for generating genetic diversity in this asexual fungal pathogen. Impact statementColletotrichum is a large genus of fungal phytopathogens that cause major economic losses on a wide range of crop plants throughout the world. These pathogens vary widely in their host specificity and may have either broad or narrow host ranges. Here, we report the first complete genome of the alfalfa (Medicago sativa) pathogen, Colletotrichum destructivum, which will facilitate the genomic analysis of host adaptation and comparison with other members of the Destructivum species complex. We identified a species-specific 1.2 Mb region within chromosome 1 displaying all the hallmarks of fungal accessory chromosomes, which may have arisen through the integration of a mini-chromosome into a core chromosome and could be linked to the pathogenicity of this fungus. We show this region is also a focus for segmental duplications, which may contribute to generating genetic diversity for adaptive evolution. Finally, we report infection by this fungus of the model legume, Medicago truncatula, providing a novel pathosystem for studying fungal-plant interactions. Data summaryAll RNA-seq data were submitted to the NCBI GEO portal under the GEO accession GSE246592. C. destructivum genome assembly and annotation are available under the NCBI BioProject PRJNA1029933 with sequence accessions CP137305-CP137317. Supplementary data (genomic and annotation files, genome browser) are available from the INRAE BIOGER Bioinformatics platform (https://bioinfo.bioger.inrae.fr/). Transposable Elements consensus sequences are also available from the French national data repository, research.data.gouv.fr with doi 10.57745/TOO1JS.

genomics↗

The genome of the rayed Mediterranean limpet Patella caerulea (Linnaeus, 1758)

Patella caerulea (Linnaeus, 1758) is a molluscan limpet species of the class Gastropoda. Endemic to the Mediterranean Sea, it is considered to be a keystone species in tidal and subtidal habitats due to its primary role in structuring and regulating the ecological balance of these habitats. It is currently being used as a bioindicator to assess the environmental quality of coastal marine waters and as a model species to understand adaptation to ocean acidification. Here we provide a high-quality reference genome assembly and annotation for Patella caerulea. We used a single specimen collected in the field to generate [~]30 Gb of PacBio HiFi data. The final assembly is 749.8 Mb large and contains 62 contigs, including the mitochondrial genome (14,938 bp). With an N50 of 48.8 Mb and 98% of the assembly contained in the 18 largest contigs, this assembly is near chromosome-scale. BUSCO scores were high (Mollusca: 87.8% complete; Metazoa: 97.2% complete) and similar to metrics observed for other chromosome-level Patella genomes, highlighting a possible bias in the Mollusca database for Patellids. We generated transcriptomic Illumina data from a second individual collected at the same locality, and used it together with protein evidence to annotate the genome. 23,938 protein coding gene models were found. By comparing this annotation with other published Patella annotations, we found that the distribution and median values of exon and gene lengths was comparable to other Patella species despite different annotation approaches. The present high-quality Patella caerulea reference genome is an important resource for future ecological and evolutionary studies. SignificanceReference genomes are essential resources for biodiversity conservation and management. Patella caerulea (Linnaeus, 1758) is a gastropod species occurring in the Mediterranean Sea that is currently used as a model to understand the impact of pollution and ocean acidification on marine biodiversity. Here we present a high-quality reference genome of P. caerulea, that almost reaches chromosome-level contiguity. We further provide a high-quality genome annotation supported by transcriptomic evidence. This reference genome will be of interest for researchers working on the ecology and evolution of marine biodiversity. All data is available on the public database NCBI for future use by researchers.

genomics↗

A Phased, Chromosome-scale Genome for Malus domestica 'WA 38'

Genome sequencing for agriculturally important Rosaceous crops has made rapid progress both in completeness and annotation quality. Whole genome sequence and annotation gives breeders, researchers, and growers information about cultivar specific traits such as fruit quality, disease resistance, and informs strategies to enhance postharvest storage. Here we present a haplotype-phased, chromosomal level genome of Malus domestica, WA 38, a new apple cultivar released to market in 2017 as Cosmic Crisp (R). Using both short and long read sequencing data with a k-mer based approach, chromosomes originating from each parent were assembled and segregated. This is the first pome fruit genome fully phased into parental haplotypes in which chromosomes from each parent are identified and separated into their unique, respective haplomes. The two haplome assemblies, Honeycrisp originated HapA and Enterprise originated HapB, are about 650 Megabases each, and both have a BUSCO score of 98.7% complete. A total of 53,028 and 54,235 genes were annotated from HapA and HapB, respectively. Additionally, we provide genome-scale comparisons to Gala, Honeycrisp, and other relevant cultivars highlighting major differences in genome structure and gene family circumscription. This assembly and annotation was done in collaboration with the American Campus Tree Genomes project that includes WA 38 (Washington State University), dAnjou pear (Auburn University), and many more. To ensure transparency, reproducibility, and applicability for any genome project, our genome assembly and annotation workflow is recorded in detail and shared under a public GitLab repository. All software is containerized, offering a simple implementation of the workflow.

genomics↗

Multiple Displacement Amplification Facilitates SMRT Sequencing of Microscopic Animals and the Genome of the Gastrotrich Lepidodermella squamata (Dujardin, 1841)

BackgroundObtaining adequate DNA for long-read genome sequencing remains a roadblock to producing contiguous genomes from small-bodied organisms. Multiple displacement amplification (MDA) leverages Phi29 DNA polymerase to produce micrograms of DNA from picograms of input. Few genomes have been generated using this approach, due to concerns over biases in amplification related to GC and repeat content and chimera production. Here, we explored the utility of MDA for generating template DNA for PacBio HiFi sequencing using Caenorhabditis elegans (Nematoda) and Lepidodermella squamata (Gastrotricha). ResultsHiFi sequencing of libraries prepared from MDA DNA produced highly contiguous and complete genomes for both C. elegans (102 Mbp assembly; 336 contigs; N50 = 868 Kbp; L50 = 39; BUSCO_nematoda: S:92.2%, D:2.7%) and L. squamata (122 Mbp assembly; 157 contigs; N50 = 3.9 Mb; L50 = 13; BUSCO_metazoa: S: 78.0%, D: 2.8%). Amplified C. elegans reads mapped to the reference genome with a rate of 99.92% and coverage of 99.75% with just one read (of 708,811) inferred to be chimeric. Coverage uniformity was nearly identical for reads from MDA DNA and reads from pooled worm DNA when mapped to the reference genome. The genome of Lepidodermella squamata, the first of its phylum, was leveraged to infer the phylogenetic position of Gastrotricha, which has long been debated, as the sister taxon of Platyhelminthes. ConclusionsThis methodology will help generate contiguous genomes of microscopic taxa whose body size precludes standard long-read sequencing. L. squamata is an emerging model in evolutionary developmental biology and this genome will facilitate further work on this species.

genomics↗

Differential Conservation and Loss of CR1 Retrotransposons in Squamates Reveals Lineage-Specific Genome Dynamics across Reptiles

Transposable elements (TEs) are repetitive DNA sequences which create mutations and generate genetic diversity across the tree of life. In amniotic vertebrates, TEs have been mainly studied in mammals and birds, whose genomes generally display low TE diversity. Squamates (Order Squamata; [~]11,000 extant species of lizards and snakes) show as much variation in TE abundance and activity as they do in species and phenotypes. Despite this high TE activity, squamate genomes are remarkably uniform in size. We hypothesize that novel, lineage-specific dynamics have evolved over the course of squamate evolution to constrain genome size across the order. Thus, squamates may represent a prime model for investigations into TE diversity and evolution. To understand the interplay between TEs and host genomes, we analyzed the evolutionary history of the CR1 retrotransposon, a TE family found in most tetrapod genomes. We compared 113 squamate genomes to the genomes of turtles, crocodilians, and birds, and used ancestral state reconstruction to identify shifts in the rate of CR1 copy number evolution across reptiles. We analyzed the repeat landscapes of CR1 in squamate genomes and determined that shifts in the rate of CR1 copy number evolution are associated with lineage-specific variation in CR1 activity. We then used phylogenetic reconstruction of CR1 subfamilies across amniotes to reveal both recent and ancient CR1 subclades across the squamate tree of life. The patterns of CR1 evolution in squamates contrast other amniotes, suggesting key differences in how TEs interact with different host genomes and at different points across evolutionary history.

genomics↗

A haplotype-resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

O_LIWhite oak (Quercus alba) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. C_LIO_LIWe report the first haplotype-resolved chromosome-scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. In addition, we probe the genetic diversity of this widespread species and investigate its phylogenetic relationships with other oaks using whole-genome data. C_LIO_LIOur genome assembly comprises two haplotypes each consisting of 12 chromosomes. We found that the species has high genetic diversity, much of which predates the divergence of Q. alba from other oak species and likely impacts divergence time estimation in Quercus. Our phylogenetic results highlight phylogenetic discordance across the genus and suggest different relationships among North American oaks than have been reported previously. Despite a high preservation of chromosome synteny and genome size across the Quercus phylogeny, certain gene families have undergone rapid changes in size including resistance genes (R genes). C_LIO_LIThe white oak genome represents a major new resource for studying genome diversity and evolution in Quercus and forest trees more generally. Future research will continue to reveal the full scope of genomic diversity across the white oak clade. C_LI

genomics↗

The dynamic genomes of Hydra and the anciently active repeat complement of animal chromosomes

Many animal genomes are characterized by highly conserved chromosomal homologies that pre-date the ancient origin of this clade. Despite such conservation, the evolutionary forces behind the retention, expansion, and contraction of chromosomal elements, and the resulting macro-evolutionary implications, are unknown. Here we present a comprehensive stem-cell resolved genomic and transcriptomic study of the fresh-water cnidarian Hydra, an animal characterized by its high regenerative capacity, the ability to propagate clonally, and an apparent lack of aging. Using single-haplotype telomere-to-telomere genome assemblies of two recently diverged hydra strains, we show how the macro-evolutionary history of chromosomal elements is shaped by both old and recent transposable element (TE) expansions. Unique features of hydra biology allowed us to compare the individual genomes of hydras three stem cell lineages. We show that distinct TE families are active at both transcriptional and genomic levels via non-random insertions in the genomes of each of these lineages. In transcriptomes, over 14,000 transcripts were composed of nearly complete TE sequences, and further classification into families, subfamilies, and individual loci reveals cell type-specific TE expression. The active TEs include elements that differentially contribute to changes in the genome size as well as persistent structural variants around loci associated with cell proliferation. Our study reveals 14 active TE families that primarily act in this role and are predominantly composed of DNA elements. Evolutionary analysis revealed that these families constitute a highly conserved TE core in eukaryotic and metazoan genomes. Our results suggest an ancient role for these core TEs as self-renewing genomic components that persist beyond ancient chromosomal homologies.

genomics↗

Wampee chromosome-level reference genome elucidates fruit sugar-acid metabolism

Wampee (Clausena lansium) is an economically significant subtropical fruit tree widely cultivated in Southern China. High-quality genomic resources are unavailable, but they are essential for functional genomics and germplasm enhancement of wampee. Here, we provide a chromosome-level genome sequence for the wampee cultivar JinFeng and a population genomic analysis of 266 accessions. The 297.1 Mb wampee genome, containing nine chromosomes with a scaffold N50 of 29.2 Mb and encoding 23,468 protein-coding genes, showed a significant improvement over the previous version. We dissected the wampee population structure and genetic differentiation in China using population genomic analysis, which detected 110 and 671 genes under a selective sweep associated with sour and sweet wampee evolution in domesticated clones, respectively. Homozygous non-synonymous single nucleotide polymorphisms are likely associated with fruit flavor differentiation. A genome-wide association study identified 220 remarkable marker-trait associations for total acid content, harboring 289 genes encoding transcription factors, transporters, and enzymes involved in sugar and acid metabolism, which are potentially useful for sour and sweet taste development in wampee fruit. Furthermore, the ethylene response factor family gene ClERF061 and the SWEET family gene ClSWEET7 were identified. Linkage assessment between the relative expression levels of ClERF061 or ClSWEET7 and the total acid/total sugar contents implied their potential involvement in sugar-acid metabolism in wampee fruits. High-quality genome resources are valuable for expediting wampee research and genome-assisted breeding.

genomics↗

A chromosome-level reference-quality genome of Punica granatum L.

Pomegranate (Punica granatum L.) is one of the most ancient edible fruit tree species. Here we reported a new chromosome-level genome assembly and annotation of sour pomegranate. We assembled the genome with a size of 331.47 Mb and used BUSCO to estimate the completeness of the assembly as 98.8%. More than 97.40% of sequences in the final assembly were anchored to 8 pseudochromosomes, higher than the corresponding percentages for the existing reference genomes Tunisia (92.62%). Using a combination of de novo prediction, protein homology and RNA-seq annotation, 29,326 protein-coding genes were predicted. We re-annotated the protein-coding genes of five other published pomegranate genomes using the same annotation method. We constructed the pan-genome of pomegranate using protein-coding genes, integrating data from our newly assembled genome and five other published genomes. The pan-genome was composed of 28,314 gene families, of which 68.96% were core genes, 30.00% were dispensable genes, and 1.04% were private genes. The chromosome-level reference genome of sour pomegranate would be valuable resource for research and molecular breeding of pomegranate.

genomics↗

Reference genome bias in light of species-specific chromosomal reorganization and translocations

Whole-genome sequencing efforts has during the past decade unveiled the central role of genomic rearrangements--such as chromosomal inversions--in evolutionary processes, including local adaptation in a wide range of taxa. However, employment of reference genomes from distantly or even closely related species for mapping and the subsequent variant calling, can lead to errors and/or biases in the datasets generated for downstream analyses. Here, we capitalize on the recently generated chromosome-anchored genome assemblies for Arctic cod (Arctogadus glacialis), polar cod (Boreogadus saida), and Atlantic cod (Gadus morhua) to evaluate the extent and consequences of reference bias on population sequencing datasets (approx. 15-20x coverage) for both Arctic cod and polar cod. Our findings demonstrate that the choice of reference genome impacts population genetic statistics, including individual mapping depth, heterozygosity levels, and cross-species comparisons of nucleotide diversity ({pi}) and genetic divergence (DXY). Further, it became evident that using a more distantly related reference genome can lead to inaccurate detection and characterization of chromosomal inversions, i.e., in terms of size (length) and location (position), due to inter-chromosomal reorganizations between species. Additionally, we observe that several of the detected species-specific inversions were split into multiple genomic regions when mapped towards a heterospecific reference. Inaccurate identification of chromosomal rearrangements as well as biased population genetic measures could potentially lead to erroneous interpretation of species-specific genomic diversity, impede the resolution of local adaptation, and thus, impact predictions of their genomic potential to respond to climatic and other environmental perturbations.

genomics↗

Silver chimaera genome assembly and identification of the holocephalan sex chromosome sequence

Cartilaginous fishes are divided into holocephalans and elasmobranchs, and comparative studies involving them are expected to elucidate how variable phenotypes and distinctive genomic properties were established in those ancient vertebrate lineages. To date, molecular-level studies on holocephalans have concentrated on the family Callorhinchidae, with a chromosome-scale genome assembly of Callorhinchus milii available. In this study, we focused on the most species-rich holocephalan family Chimaeridae and sequenced the genome of its member, silver chimaera (Chimaera phantasma). We report the first chromosome-scale genome assembly of the Chimaeridae, with high continuity and completeness, which exhibited a large intragenomic variation of chromosome lengths, which is correlated with intron size. This pattern is observed more widely in vertebrates and at least partly accounts for cross-species genome size variation. A male-female comparison identified a silver chimaera genomic scaffold with a double sequence depth for females, which we identify as an X chromosome fragment. This is the first DNA sequence-based evidence of a holocephalan sex chromosome, suggesting a male heterogametic sex determination system. This study, allowing the first chromosome-level comparison among holocephalan genomes, will trigger in-depth understanding of the genomic diversity among vertebrates as well as species population genetic structures based on the genome assembly of high completeness.

genomics↗

A workflow for practical training in ecological genomics using Oxford Nanopore long-read sequencing

Long-read single molecule sequencing technologies continue to grow in popularity for genome assembly and provide an effective way to resolve large and complex genomic variants. However, uptake of these technologies for teaching and training is hampered by the complexity of high molecular weight DNA extraction protocols, the time required for library preparation and the costs for sequencing, as well as challenges with downstream data analyses. Here, we present a full long-read workflow optimised for teaching, that covers each stage from DNA extraction, to library preparation and sequencing, to data QC and genome assembly and characterisation, that can be completed in under two weeks. We use a specific case study of plant identification, where students identify an anonymous plant sample by sequencing and assembling the genome and comparing it to other samples and to reference databases. In testing, long-read genome skimming of nine wild-collected plant species extracted with a modified kit-based approach produced an average of 8Gb of Oxford Nanopore data, enabling the complete assembly of plastid genomes, and partial assembly of nuclear genomes. In the classroom, all students were able to complete the protocols, and to correctly identify their plant samples based on BOLD searches of barcoding loci extracted from the plastid genome, coupled with phylogenetic analyses of whole plastid genomes. We supply all the learning material and raw data allowing this to be adapted to a range of teaching settings.

genomics↗

The Nuclear and Mitochondrial Genomes of Amoebophrya sp. ex Karlodinium veneficum

Dinoflagellates are a diverse group of microplankton that include free-living, symbiotic, and parasitic species. Amoebophrya, a basal lineage of parasitic dinoflagellates, infects a variety of marine microorganisms, including harmful-bloom-forming algae. Although there are currently three published Amoebophrya genomes, this genus has considerable genomic diversity. We add to the growing genomic data for Amoebophrya with an annotated genome assembly for Amoebophrya sp. ex Karlodinium veneficum. This species appears to translate all three canonical stop codons contextually. Stop codons are present in the open reading frames of about half of the predicted gene models, including genes essential for cellular function. The in-frame stop codons are likely translated by suppressor tRNAs that were identified in the assembly. We also assembled the mitochondrial genome, which has remained elusive in the previous Amoebophrya genome assemblies. The mitochondrial genome assembly consists of many fragments with high sequence identity in the genes but low sequence identity in intergenic regions. Nuclear and mitochondrially-encoded proteins indicate that Amoebophrya sp. ex K. veneficum does not have a bipartite electron transport chain, unlike previously analyzed Amoebophrya species. This study highlights the importance of analyzing multiple genomes from highly diverse genera such as Amoebophrya. SummaryThis new long-read assembly demonstrates the remarkable diversity found within Amoebophrya. Despite being assigned the rank of genus, the available genome assemblies indicate significant variation in gene content, AT content, genetic codes, and potentially mitochondrial biology. Furthermore, this study contributes to the expanding list of organisms that contextually translate all three canonical stop codons. Although the mechanisms underlying such a genetic code remain elusive, the relative ease of culturing Amoebophrya suggests it may be useful as a model organism for future research on this subject.

genomics↗

A blended genome and exome sequencing method captures genetic variation in an unbiased, high-quality, and cost-effective manner

We deployed the Blended Genome Exome (BGE), a DNA library blending approach that generates low pass whole genome (1-4x mean depth) and deep whole exome (30-40x mean depth) data in a single sequencing run. This technology is cost-effective, empowers most genomic discoveries possible with deep whole genome sequencing, and provides an unbiased method to capture the diversity of common SNP variation across the globe. To evaluate this new technology at scale, we applied BGE to sequence >53,000 samples from the Populations Underrepresented in Mental Illness Associations Studies (PUMAS) Project, which included participants across African, African American, and Latin American populations. We evaluated the accuracy of BGE imputed genotypes against raw genotype calls from the Illumina Global Screening Array. All PUMAS cohorts had R2 concordance [&ge;]95% among SNPs with MAF[&ge;]1%, and never fell below [&ge;]90% R2 for SNPs with MAF<1%. Furthermore, concordance rates among local ancestries within two recently admixed cohorts were consistent among SNPs with MAF[&ge;]1%, with only minor deviations in SNPs with MAF<1%. We also benchmarked the discovery capacity of BGE to access protein-coding copy number variants (CNVs) against deep whole genome data, finding that deletions and duplications spanning at least 3 exons had a positive predicted value of [~]90%. Our results demonstrate BGE scalability and efficacy in capturing SNPs, indels, and CNVs in the human genome at 28% of the cost of deep whole-genome sequencing. BGE is poised to enhance access to genomic testing and empower genomic discoveries, particularly in underrepresented populations.

genomics↗

Coordinated control of genome-nuclear lamina interactions by Topoisomerase 2B and Lamin B receptor

Lamina-associated domains (LADs) are megabase-sized genomic regions anchored to the nuclear lamina (NL). Factors controlling the interactions of the genome with the NL have largely remained elusive. Here, we identified DNA topoisomerase 2 beta (TOP2B) as a regulator of these interactions. TOP2B binds predominantly to inter-LAD (iLAD) chromatin and its depletion results in a partial loss of genomic partitioning between LADs and iLADs, suggesting that its activity might protect specific iLADs from interacting with the NL. TOP2B depletion affects LAD interactions with lamin B receptor (LBR) more than with lamins. LBR depletion phenocopies the effects of TOP2B depletion, despite the different positioning of the two proteins in the genome. This suggests a complementary mechanism for organising the genome at the NL. Indeed, co-depletion of TOP2B and LBR causes partial LAD/iLAD inversion, reflecting changes typical of oncogene-induced senescence. We propose that a coordinated axis controlled by TOP2B in iLADs and LBR in LADs maintains the partitioning of the genome between the NL and the nuclear interior. HighlightsO_LILADs and iLADs differ in supercoiling state C_LIO_LITOP2B controls genome partitioning between nuclear lamina and nuclear interior C_LIO_LITOP2B depletion preferentially affects genome interactions with LBR C_LIO_LISimilar impact of TOP2B depletion and LBR depletion on genome-NL interactions C_LIO_LICo-depletion of TOP2B and LBR recapitulates LAD reshaping typical of oncogene-induced senescence. C_LI

genomics↗

StableLift: Optimized Germline and Somatic Variant Detection Across Genome Builds

Reference genomes are foundational to modern genomics. Our growing understanding of genome structure leads to continual improvements in reference genomes and new genome "builds" with incompatible coordinate systems. We quantified the impact of genome build on germline and somatic variant calling by analyzing tumour-normal whole-genome pairs against the two most widely used human genome builds. The average individual had a build-discordance of 3.8% for germline SNPs, 8.6% for germline SVs, 25.9% for somatic SNVs and 49.6% for somatic SVs. Build-discordant variants are not simply false-positives: 47% were verified by targeted resequencing. Build-discordant variants were associated with specific genomic and technical features in variant- and algorithm-specific patterns. We leveraged these patterns to create StableLift, an algorithm that predicts cross-build stability with AUROCs of 0.934 {+/-} 0.029. These results call for significant caution in cross-build analyses and for use of StableLift as a computationally efficient solution to mitigate inter-build artifacts.

genomics↗

Kudoa genomes from contaminated hosts reveal extensive gene order conservation and rapid sequence evolution

Myxozoans are obligate endoparasites that belong to the phylum Cnidaria. Compared to their closest free-living relatives, they have evolved highly simplified body plans and reduced genomes. Kudoa iwatai, for example, has lost upwards of two thirds of genes thought to have been present in its ancestors. However, little is known about myxozoan genome architecture because of a lack of sufficiently contiguous genome assemblies. This work presents two new, near-chromosomal Kudoa genomes, built entirely from low-coverage long reads from infected fish samples. The results illustrate the potential of using unsupervised learning methods to disentangle sequences from different sources, and facilitate producing genomes from undersampled taxa. Extracting distinct components of chromatin interaction networks allows scaffolds from mixed samples to be assigned to their source genomes. Meanwhile, low-dimensional embeddings of read composition permit targeted assembly of potential parasite reads. Despite drastic changes in genome architecture in the lineage leading to Kudoa and considerable sequence divergence between the two genomes, gene order is highly conserved. Although parasitic cnidarians show rapid protein evolution compared to their free-living relatives, there is limited evidence of less efficient selection. While deleterious substitutions may become fixed at a higher rate, large evolutionary distances between species make robustly analysing patterns of molecular evolution challenging. These observations highlight the importance of filling in taxonomic gaps, to allow a comprehensive assessment of the impacts of parasitism on genome evolution.

genomics↗

A Portable and Scalable Genomic Analysis Pipeline for Streptococcus pneumoniae Surveillance: GPS Pipeline

Ever increasing global sequencing capacity provides an unprecedented opportunity in utilising genomic information captured from whole-genome sequencing to enhance pathogen surveillance. However, there is a growing need for developing user-friendly tools to effectively analyse the increasing volume of data. To meet this need, we have developed a genomic analysis pipeline, GPS Pipeline, which is portable and scalable to analyse genomes of Streptococcus pneumoniae, a major bacterial pathogen that is estimated to cause 317,000 child deaths worldwide every year. The GPS Pipeline is based on Nextflow and containerisation technology, and designed to enable researchers generating public health relevant output, including in silico serotypes, pneumococcal lineages (i.e. GPSCs), multilocus sequence types, and antimicrobial susceptibilities against 20 commonly used antibiotics,with minimal software setup requirements and bioinformatic expertise, in order to analyse genomic data at scale with ease. The GPS Pipeline provides a streamlined workflow that improves responsiveness in genomic surveillance on pneumococci. Data SummaryThe GPS Pipeline is available on GitHub at github.com/GlobalPneumoSeq/gps-pipeline. Published data from the GPS Database is available on Monocle Data Viewer at data.monocle.sanger.ac.uk and associated sequence read files are searchable and downloadable in the European Nucleotide Archive at ebi.ac.uk/ena via their ERR accession numbers. Impact StatementThe GPS Pipeline advances global genomic surveillance of Streptococcus pneumoniae by providing a scalable, portable, and user-friendly tool for analysing whole-genome sequencing data. Leveraging Nextflow and containerisation technology, it minimises bioinformatics expertise requirements and infrastructure needs, making it particularly valuable in low- and middle-income countries where pneumococcal disease burden is high. This pipeline ensures reproducibility and stability across platforms, facilitating rapid and accurate pneumococci genomic analysis. By streamlining data processing, the GPS Pipeline enhances pathogen surveillance, generates evidence to support vaccine strategy development, and empowers researchers worldwide, ultimately contributing to improved public health outcomes.

genomics↗