bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

Pangolin genomes offer key insights and resources for the world's most trafficked wild mammals

Pangolins form a group of scaly mammals that are trafficked at record numbers for their meat and medicinal properties. Despite their great conservation concern, knowledge of their evolution is limited by a paucity of genomic data. We aim to produce exhaustive genomic resources that include 3 238 orthologous genes and whole-genome polymorphisms to assess the evolution of all eight pangolin species. Robust orthologous gene-based phylogenies recovered the monophyly of the three genera of pangolins, and highlighted the existence of an undescribed species closely related to South-East Asian pangolins. Signatures of middle Miocene admixture between an extinct, possibly European, lineage and the ancestor of South-East Asian pangolins, provides new insights into the early evolutionary history of the group. Demographic trajectories and genome-wide heterozygosity estimates revealed contrasts between continental vs. island populations and species lineages, suggesting that conservation planning should consider intra-specific patterns. With the expected loss of genomic diversity from recent, extensive trafficking not yet been realized in pangolins, we recommend that populations are genetically surveyed to anticipate any deleterious impact of the illegal trade. Finally, we produce a complete set of genomic resources that will be integral for future conservation management and forensic endeavors required for conserving pangolins, including tracing their illegal trade. These include the completion of whole-genomes for pangolins through the first reference genome with long reads for the giant pangolin (Smutsia gigantea) and new draft genomes (~43x-77x) for four additional species, as well as a database of orthologous genes with over 3.4 million polymorphic sites.

genomics↗

Distinct genomic contexts predict gene presence-absence variation in different pathotypes of a fungal plant pathogen

BackgroundFungi use the accessory segments of their pan-genomes to adapt to their environments. While gene presence-absence variation (PAV) contributes to shaping these accessory gene reservoirs, whether these events happen in specific genomic contexts remains unclear. Additionally, since pan-genome studies often group together all members of the same species, it is uncertain whether genomic or epigenomic features shaping pan-genome evolution are consistent across populations within the same species. Fungal plant pathogens are useful models for answering these questions because members of the same species often infect distinct hosts, and they frequently rely on gene PAV to adapt to these hosts. ResultsWe analyzed gene PAV in the rice and wheat blast fungus, Magnaporthe oryzae, and found that PAV of disease-causing effectors, antibiotic production, and non-self-recognition genes may drive the adaptation of the fungus to its environment. We then analyzed genomic and epigenomic features and data from available datasets for patterns that might help explain these PAV events. We observed that proximity to transposable elements (TEs), gene GC content, gene length, expression level in the host, and histone H3K27me3 marks were different between PAV genes and conserved genes, among other features. We used these features to construct a random forest classifier that was able to predict whether a gene is likely to experience PAV with high precision (86.06%) and recall (92.88%) in rice-infecting M. oryzae. Finally, we found that PAV in wheat- and rice-infecting pathotypes of M. oryzae differed in their number and their genomic context. ConclusionsOur results suggest that genomic and epigenomic features of gene PAV can be used to better understand and even predict fungal pan-genome evolution. We also show that substantial intra-species variation can exist in these features.

genomics↗

Evolutionary history of the transposon-invaded Pithoviridae genomes

Pithoviridae are amoeba-infecting giant viruses possessing the largest viral particles known so far. Since the discovery of Pithovirus sibericum, recovered from a 30,000-y-old permafrost sample, other pithoviruses, and related cedratviruses, were isolated from various terrestrial and aquatic samples. Here we report the isolation and genome sequencing of two Pithoviridae from soil samples, in addition to three other recent isolates. Using the 12 available genome sequences, we conducted a thorough comparative genomics study of the Pithoviridae family to decipher the organization and evolution of their genomes. Our study reveals a non-uniform genome organization in two main regions: one concentrating core genes, and another gene duplications. We also found that Pithoviridae genomes are more conservative than other families of giant viruses, with a low and stable proportion (5% to 7%) of genes originating from horizontal transfers. Genome size variation within the family is mainly due to variations in gene duplication rates (from 14% to 28%) and massive invasion by inverted repeats. While these repeated elements are absent from cedratviruses, repeat-rich regions cover as much as a quarter of the pithoviruses genomes. These regions, identified using a dedicated pipeline, are hotspots of mutations, gene capture events and genomic rearrangements, that contribute to their evolution.

genomics↗

Endophyte genomes support greater metabolic gene cluster diversity compared with non-endophytes in Trichoderma

Trichoderma is a cosmopolitan genus with diverse lifestyles and nutritional modes, including mycotrophy, saprophytism, and endophytism. Previous research has reported greater metabolic gene repertoires in endophytic fungal species compared to closely-related non-endophytes. However, the extent of this ecological trend and its underlying mechanisms are unclear. Some endophytic fungi may also be mycotrophs and have one or more mycoparasitism mechanisms. Mycotrophic endophytes are prominent in certain genera like Trichoderma, therefore, the mechanisms that enable these fungi to colonize both living plants and fungi may be the result of expanded metabolic gene repertoires. Our objective was to determine what, if any, genomic features are overrepresented in endophytic fungi genomes in order to undercover the genomic underpinning of the fungal endophytic lifestyle. Here we compared metabolic gene cluster and mycoparasitism gene diversity across a dataset of thirty-eight Trichoderma genomes representing the full breadth of environmental Trichodermas diverse lifestyles and nutritional modes. We generated four new Trichoderma endophyticum genomes to improve the sampling of endophytic isolates from this genus. As predicted, endophytic Trichoderma genomes contained, on average, more total biosynthetic and degradative gene clusters than non-endophytic isolates, suggesting that the ability to create/modify a diversity of metabolites potential is beneficial or necessary to the endophytic fungi. Still, once the phylogenetic signal was taken in consideration, no particular class of metabolic gene cluster was independently associated with the Trichoderma endophytic lifestyle. Several mycoparasitism genes, but no chitinase genes, were associated with endophytic Trichoderma genomes. Most genomic differences between Trichoderma lifestyles and nutritional modes are difficult to disentangle from phylogenetic divergences among species, suggesting that Trichoderma genomes maybe particularly well-equipped for lifestyle plasticity. We also consider the role of endophytism in diversifying secondary metabolism after identifying the horizontal transfer of the ergot alkaloid gene cluster to Trichoderma.

genomics↗

Intraspecies genomic divergence of coral algal symbionts shaped by gene duplication

Dinoflagellates of Order Suessiales include the diverse Family Symbiodiniaceae known for their role as essential coral reef symbionts, and the cold-adapted Polarella glacialis. These taxa inhabit a broad range of ecological niches and exhibit extensive genomic divergence, although their genomes are in the smaller size ranges (haploid size < 3 Gbp) compared to most other dinoflagellates. Different isolates of a species are known to form symbiosis with distinct hosts and exhibit different regimes of gene expression, but intraspecies whole-genome divergence remains little known. Focusing on three Symbiodiniaceae species (the free-living Effrenium voratum, and the symbiotic Symbiodinium microadriaticum and Durusdinium trenchii) and the free-living outgroup P. glacialis, all for which whole-genome data from multiple isolates are available, we assessed intraspecies genomic divergence at sequence and structural levels. Our analysis based on alignment and alignment-free methods revealed greater extent of intraspecies sequence divergence in symbiodiniacean species than in P. glacialis. Our results also reveal the implications of gene duplication in generating functional innovation and diversification of Symbiodiniaceae, particularly in D. trenchii for which whole-genome duplication was involved. Interestingly, tandem duplication of single-exon genes was found to be more prevalent in genomes of free-living species than in those of symbiotic species. These results in combination demonstrate the remarkable intraspecies genomic divergence in dinoflagellates under the constraint of reduced genome sizes, shaped by genetic duplications and symbiogenesis events during diversification of Symbiodiniaceae.

genomics↗

Unravelling genome organization of neopolyploid flatworm Macrostomum lignano

Whole genome duplication (WGD) is an evolutionary event resulting in a redundancy of genetic material. Different mechanisms of genome doubling through allo- or autopolyploidization could lead to distinct evolutionary trajectories of newly formed polyploids. Genome studies on such species are undoubtedly important for understanding one of the crucial stages of genome evolution. However, assembling neopolyploid appears to be a challenging task because its genome consists of two homologous (or homeologous) chromosome sets and therefore contains the extended paralogous regions with a high homology level. Post-WGD evolution of polyploids includes rediploidization, first part of which is cytogenetic diploidization led to the formation of species, whose polyploid origin might be hidden by disomic inheritance and diploid-like meiosis. Earlier we uncovered the hidden polyploid origin of free-living flatworms of the genus Macrostomum (Macrostomum lignano, M. janickei, and M. mirumnovem). Despite the different mechanisms for their genome doubling, cytogenetic diploidization in these species accompanied by intensive chromosomal rearrangements including chromosomes fusions. In this study, we reported unusual subgenomic organization of M. lignano through generation and sequencing of two new laboratory sublines of DV1 that differ only by a copy number of the large chromosome MLI1. Using non-trivial assembly-free comparative analysis of their genomes, including adapted multivariate k-mer analysis, and self-homology within the published genome assembly of M. lignano, we deciphered DNA sequences belonging to MLI1 and validated them by sequencing the pool of microdissected MLI1. Here we presented the uncommon mechanism of genome rediplodization of M. lignano, which consists in (1) presence of three subgenomes, emerged via formation of large fused chromosome and its variants, and (2) sustaining their heterozygosity through inter- and intrachromosomal rearrangements.

genomics↗

A chromosome scale genome assembly and evaluation of mtDNA variation in the willow leaf beetle Chrysomela aeneicollis

The leaf beetle Chrysomela aeneicollis has a broad geographic range across Western North America, but is restricted to cool habitats at high elevations along the west coast. Central California populations occur only at high altitudes (2900-3450 m) where they are limited by reduced oxygen supply and recent drought conditions that are associated with climate change. Here we report a chromosome-scale genome assembly alongside a complete mitochondrial genome, and characterize differences among mitochondrial genomes along a latitudinal gradient over which beetles show substantial population structure and adaptation to fluctuating temperatures. Our scaffolded genome assembly consists of 21 linkage groups; one of which we identified as the X chromosome based on female/male whole genome sequencing coverage and orthology with Tribolium castaneum. We identified repetitive sequences in the genome and found them to be broadly distributed across all linkage groups. Using a reference transcriptome, we annotated a total of 12,586 protein coding genes. We also describe differences in putative secondary structures of mitochondrial RNA molecules, which may generate functional differences important in adaptation to harsh abiotic conditions. We document substitutions at mitochondrial tRNA molecules and substitutions and insertions in the 16S rRNA region that could affect intermolecular interactions with products from the nuclear genome. This first chromosome-level reference genome will enable genomic research in this important model organism for understanding the biological impacts of climate change on montane insects.

genomics↗

Three-dimensional genome architecture persists in a 52,000-year-old woolly mammoth skin sample

Ancient DNA (aDNA) sequencing analysis typically involves alignment to a modern reference genome assembly from a related species. Since aDNA molecules are fragmentary, these alignments yield information about small-scale differences, but provide no information about larger features such as the chromosome structure of ancient species. We report the genome assembly of a female Late Pleistocene woolly mammoth (Mammuthus primigenius) with twenty-eight chromosome-length scaffolds, generated using mammoth skin preserved in permafrost for roughly 52,000 years. We began by creating a modified Hi-C protocol, dubbed PaleoHi-C, optimized for ancient samples, and using it to map chromatin contacts in a woolly mammoth. Next, we developed "reference-assisted 3D genome assembly," which begins with a reference genome assembly from a related species, and uses Hi-C and DNA-Seq data from a target species to split, order, orient, and correct sequences on the basis of their 3D proximity, yielding accurate chromosome-length scaffolds for the target species. By means of this reference-assisted 3D genome assembly, PaleoHi-C data reveals the 3D architecture of a woolly mammoth genome, including chromosome territories, compartments, domains, and loops. The active (A) and inactive (B) genome compartments in mammoth skin more closely resemble those observed in Asian elephant skin than the compartmentalization patterns seen in other Asian elephant tissues. Differences in compartmentalization between these skin samples reveal sequences whose transcription was potentially altered in mammoth. We observe a tetradic structure for the inactive X chromosome in mammoth, distinct from the bipartite architecture seen in human and mouse. Generating chromosome-length genome assemblies for two other elephantids (Asian and African elephant), we find that the overall karyotype, and this tetradic Xi structure, are conserved throughout the clade. These results illustrate that cell-type specific epigenetic information can be preserved in ancient samples, in the form of DNA geometry, and that it may be feasible to perform de novo genome assembly of some extinct species.

genomics↗

Histone deacetylation and cytosine methylation compartmentalize heterochromatic regions in the genome organization of Neurospora crassa

Chromosomes must correctly fold in eukaryotic nuclei for proper genome function. Eukaryotic organisms hierarchically organize their genomes, including in the fungus Neurospora crassa, where chromatin fiber loops compact into Topologically Associated Domain (TAD)-like structures formed by heterochromatic region aggregation. However, insufficient data exists on how histone post-translational modifications, including acetylation, affect genome organization. In Neurospora, the HCHC complex (comprised of the proteins HDA-1, CDP-2, HP1, and CHAP) deacetylates heterochromatic nucleosomes, as loss of individual HCHC members increases centromeric acetylation and alters the methylation of cytosines in DNA. Here, we assess if the HCHC complex affects genome organization by performing Hi-C in strains deleted of the cdp-2 or chap genes. CDP-2 loss increases intra- and inter-chromosomal heterochromatic region interactions, while loss of CHAP decreases heterochromatic region compaction. Individual HCHC mutants exhibit different patterns of histone post-translational modifications genome-wide: without CDP-2, heterochromatic H4K16 acetylation is increased, yet smaller heterochromatic regions lose H3K9 trimethylation and gain inter-heterochromatic region interactions; CHAP loss produces minimal acetylation changes but increases heterochromatic H3K9me3 enrichment. Loss of both CDP-2 and the DIM-2 DNA methyltransferase causes extensive genome disorder, as heterochromatic-euchromatic contacts increase despite additional H3K9me3 enrichment. Our results highlight how the increased cytosine methylation in HCHC mutants ensures genome compartmentalization when heterochromatic regions become hyperacetylated without HDAC activity. Significance StatementThe mechanisms driving chromosome organization in eukaryotic nuclei, including in the filamentous fungus Neurospora crassa, are currently unknown, but histone post-translational modifications may be involved. Histone proteins can be acetylated to form active euchromatin while histone deacetylases (HDACs) remove acetyl marks to form silent heterochromatin; these heterochromatic regions cluster, forming strong interactions, in Neurospora genome organization. Here, we show that mutants of a heterochromatin-specific HDAC, HCHC, increase heterochromatic histone acetylation genome-wide and contact probability between distant heterochromatic loci. HCHC loss also impacts cytosine methylation, and in strains lacking both the HCHC and cytosine methylation, heterochromatic regions interact more with euchromatin. Our results suggest cytosine methylation normally functions to segregate silent and active loci when heterochromatic acetylation increases.

genomics↗

A high-quality reference genome for the common creek chub, Semotilus atromaculatus

Creek chub (Semotilus atromaculatus) are a leuciscid minnow species commonly found in anthropogenically disturbed environments, making them an excellent model organism to study human impacts on aquatic systems. Genomic resources for creek chub and other leuciscid species are currently limited. However, advancements in DNA sequencing now allow us to create genomic resources at a historically low cost. Here, we present a high quality 239 contig reference genome for the common creek chub, created with PacBio HiFi sequencing. We compared the assembly quality of two pipelines: Pacific Biosciences Improved Phase Assembly (IPA; 873 contigs) and Hifiasm (239 contigs). Quality and completeness of this genome is comparable to the zebrafish (Danioninae) and fathead minnow (Leuciscidae) genomes. The creek chub genome is highly syntenic to the zebrafish and fathead minnow genomes, and while our assembly does not resolve into the expected 25 chromosomes, synteny with zebrafish suggests that each creek chub chromosome is likely represented by 1-4 large contigs in our assembly. This reference genome is a valuable resource that will enhance genomic bio-diversity studies of creek chub and other non-model leuciscid species common to disturbed environments.

genomics↗

A laboratory framework for ongoing optimisation of amplification based genomic surveillance programs

Constantly evolving viral populations affect the specificity of primers and quality of genomic surveillance. This study presents a framework for continuous optimisation of sequencing efficiency for public health surveillance based on the ongoing evolution of the COVID-19 pandemic. SARS-CoV-2 genomic clustering capacity based on three amplification based whole genome sequencing schemes was assessed using decreasing thresholds of genome coverage and measured against epidemiologically linked cases. Overall genome coverage depth and individual amplicon depth were used to calculate an amplification efficiency metric. Significant loss of genome coverage over time was documented which was recovered by optimisation of primer pooling or implementation of new primer sets. A minimum of 95% genome coverage was required to cluster 94% of epidemiologically defined SARS-CoV-2 transmission events. Clustering resolution fell to 70% when only 85% of genome coverage was achieved. The framework presented in this study can provide public health genomic surveillance programs a systematic process to ensure an agile and effective laboratory response during rapidly evolving viral outbreaks.

genomics↗

Concurrent profiling of multiscale 3D genome organization and gene expression in single mammalian cells

The organization of mammalian genomes within the nucleus features a complex, multiscale three-dimensional (3D) architecture. The functional significance of these 3D genome features, however, remains largely elusive due to limited single-cell technologies that can concurrently profile genome organization and transcriptional activities. Here, we report GAGE-seq, a highly scalable, robust single-cell co-assay that simultaneously measures 3D genome structure and transcriptome within the same cell. Employing GAGE-seq on mouse brain cortex and human bone marrow CD34+ cells, we comprehensively characterized the intricate relationships between 3D genome and gene expression. We found that these multiscale 3D genome features collectively inform cell type-specific gene expressions, hence contributing to defining cell identity at the single-cell level. Integration of GAGE-seq data with spatial transcriptomic data revealed in situ variations of the 3D genome in mouse cortex. Moreover, our observations of lineage commitment in normal human hematopoiesis unveiled notable discordant changes between 3D genome organization and gene expression, underscoring a complex, temporal interplay at the single-cell level that is more nuanced than previously appreciated. Together, GAGE-seq provides a powerful, cost-effective approach for interrogating genome structure and gene expression relationships at the single-cell level across diverse biological contexts.

genomics↗

Whole genome assembly of a hybrid Trypanosoma cruzi strain assembled with nanopore sequencing alone

Trypanosoma cruzi is the causative agent of Chagas disease, which causes 10,000 deaths per year. Despite the high mortality caused by the pathogen, relatively few parasite genomes have been assembled to date; even some commonly used laboratory strains do not have publicly available genome assemblies. This is at least partially due to T. cruzis highly complex and highly repetitive genome: while describing the variation in genome content and structure is critical to better understanding T. cruzi biology and the mechanisms that underlie Chagas disease, the complexity of the genome defies investigation using traditional short read sequencing methods. Here, we have generated a high-quality whole genome assembly of the hybrid Tulahuen strain, a commercially available Type VI strain, using long read Nanopore sequencing without short read scaffolding. Using automated tools and manual curation for annotation, we report a genome with 25% repeat regions, 17% variable multigene family members, and 27% transposable elements. Notably, we find that regions with transposable elements are significantly enriched for surface proteins, and that on average surface proteins are closer to transposable elements compared to other coding regions. This finding supports a possible mechanism for diversification of surface proteins in which mobile genetic elements such as transposons facilitate recombination within the gene family. This work demonstrates the feasibility of nanopore sequencing to resolve complex regions of T. cruzi genomes, and with these resolved regions, provides support for a possible mechanism for genomic diversification.

genomics↗

High-quality genome of the zoophytophagous stinkbug, Nesidiocoris tenuis, informs their food habitadaptation

The zoophytophagous stink bug, Nesidiocoris tenuis, is a promising natural enemy of micropests such as whiteflies and thrips. This bug possesses both phytophagous and entomophagous food habits, enabling it to obtain nutrition from both plants and insects. This trait allows us to maintain its population density in agricultural fields by introducing insectary plants, even when the pest prey density is extremely low. However, if the bugs population becomes too dense, they can sometimes damage crop plants. This dual character seems to arise from the food preferences and chemosensation of this predator. To understand the genomic landscape of N. tenuis, we examined the whole genome sequence of a commercially available Japanese strain. We used long-read sequencing and Hi-C analysis to assemble the genome at the chromosomal level. We then conducted a comparative analysis of the genome with previously reported genomes of phytophagous and hematophagous stink bugs to focus on the genetic factors contributing to this species herbivorous and carnivorous tendencies. Our findings suggest that the gustatory gene set plays a pivotal role in adapting to food habits, making it a promising target for selective breeding. Furthermore, we identified the whole genomes of microorganisms symbiotic with this species through genomic analysis. We believe that our results shed light on the food habit adaptations of N. tenuis and will accelerate breeding efforts based on new breeding techniques for natural enemy insects, including genomics and genome editing.

genomics↗

Chromosome-level genome assembly for the angiosperm Silene conica

The angiosperm genus Silene has been the subject of extensive study in the field of ecology and evolution, but the availability of high-quality reference genome sequences has been limited for this group. Here, we report a chromosome-level assembly for the genome of Silene conica based on PacBio HiFi, Hi-C and Bionano technologies. The assembly produced 10 scaffolds (one per chromosome) with a total length of 862 Mb and only [~]1% gap content. These results confirm previous observations that S. conica and its relatives have a reduced base chromosome number relative to the genuss ancestral state of 12. Silene conica has an exceptionally large mitochondrial genome (>11 Mb), predominantly consisting of sequence of unknown origins. Analysis of shared sequence content suggests that it is unlikely that transfer of nuclear DNA is the primary driver of this mitochondrial genome expansion. More generally, this assembly should provide a valuable resource for future genomic studies in Silene, including comparative analyses with related species that recently evolved sex chromosomes. SignificanceWhole-genome sequences have been largely lacking for species in the genus Silene even though these flowering plants have been used for studying ecology, evolution, and genetics for over a century. Here, we address this gap by providing a high-quality nuclear genome assembly for S. conica, a species known to have greatly accelerated rates of sequence and structural divergence in its mitochondrial and plastid genomes. This resource will be valuable in understanding the coevolutionary interactions between nuclear and cytoplasmic genomes and in comparative analyses across this highly diverse genus.

genomics↗

The landscape of genomic structural variation in Indigenous Australians

Indigenous Australians harbour rich and unique genomic diversity. However, Aboriginal and Torres Strait Islander ancestries are historically under-represented in genomics research and almost completely missing from reference databases. Addressing this representation gap is critical, both to advance our understanding of global human genomic diversity and as a prerequisite for ensuring equitable outcomes in genomic medicine. Here, we apply population-scale whole genome long-read sequencing to profile genomic structural variation across four remote Indigenous communities. We uncover an abundance of large indels (20-49bp; n=136,797) and structural variants (SVs; [&ge;]50bp; n=159,912), the majority of which are composed of tandem repeat or interspersed mobile element sequences (90%) and have not been previously annotated (73%). A large fraction of SVs appear to be exclusive to Indigenous Australians (>30%) and the majority of these are found in only a single community, underscoring the need for broad and deep sampling to achieve a comprehensive catalogue of genomic structural variation across the Australian continent. Finally, we explore short-tandem repeats (STRs) throughout the genome to characterise allelic diversity at 50 known disease loci, uncover hundreds of novel repeat expansion sites within protein-coding genes, and identify unique patterns of diversity and constraint among STR sequences. Our study sheds new light on the dimensions, diversity and evolutionary trajectories of genomic structural variation within and beyond Australia.

genomics↗

The genomes of the Macadamia genus

Macadamia, a genus native to Eastern Australia, comprises four species, Macadamia integrifolia, M. tetraphylla, M. ternifolia, and M. jansenii. Macadamia was recently domesticated largely from a limited gene pool of Hawaiian germplasm and has become a commercially significant nut crop. Disease susceptibility and climate adaptability challenges, highlight the need for use of a wider range of genetic resources for macadamia production. High quality haploid resolved genome assemblies were generated using HiFiasm to allow comparison of the genomes of the four species. Assembly sizes ranged from 735 Mb to 795 Mb and N50 from 53.7 Mb to 56 Mb, indicating high assembly continuity with most of the chromosomes covered telomere to telomere. Repeat analysis revealed that approximately 61% of the genomes were repetitive sequence. The BUSCO completeness scores ranged from 95.0% to 98.9%, confirming good coverage of the genomes. Gene prediction identified 37198 to 40534 genes. The ks distribution plot of Macadamia and Telopea suggests Macadamia has undergone a whole genome duplication event prior to divergence of the four species and that Telopea genome was duplicated more recently. Synteny analysis revealed a high conservation and similarity of the genome structure in all four species. Differences in the content of genes of fatty acid and cyanogenic glycoside biosynthesis were found between the species. An antimicrobial gene with a conserved cysteine motif was found in all four species. The four genomes provide reference genomes for exploring genetic variation across the genus in wild and domesticated germplasm to support plant breeding.

genomics↗

Shared features underlying compact genomes and extreme habitat use in chironomid midges

Non-biting midges (family Chironomidae) are found throughout the world in a diverse array of aquatic and terrestrial habitats, can often tolerate harsh conditions such as hypoxia or desiccation, and have consistently compact genomes. Yet we know little about the shared molecular basis for these attributes and how they have evolved across the family. Here, we address these questions by first creating high-quality, annotated reference assemblies for Tanytarsus gracilentus (subfamily Chironominae, tribe Tanytarsini) and Parochlus steinenii (subfamily Podonominae). Using these and other publicly available assemblies, we created a time-calibrated phylogenomic tree for family Chironomidae with outgroups from order Diptera. We used this phylogeny to test for features associated with compact genomes, as well as examining patterns of gene family evolution and positive selection that may underlie chironomid habitat tolerances. Our results suggest that compact genomes evolved in the most recent common ancestor of Chironomidae and Ceratopogonidae, and that this occurred mainly through reductions in non-coding regions (introns, intergenic sequences, and repeat elements). Gene families that significantly expanded in Chironomidae included biological processes that may relate to tolerance of stressful environments, such as temperature homeostasis, inflammatory response, melanization defense response, and trehalose transport. We identified a number of genes with evidence for positive selection in Chironomidae, notably sulfonylurea receptor, peroxiredoxin-1, and protein kinase D. Our results help to understand the genomic basis for the small genomes and extreme habitat use in this widely distributed group. Significance StatementChironomid midges are known for having small genomes and being able to tolerate many forms of environmental stress, yet little is known of the shared features of their genomes that may underlie these traits. We found that reductions in non-coding regions coincide with small chironomid genomes, and we identified duplicated and/or selected genes that may equip chironomids to tolerate harsh conditions. These results describe the key genomic changes in chironomid midges that may explain their ability to inhabit a range of extreme habitats across the world.

genomics↗