bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Conservative route to extreme genome compaction in a miniature annelid

The causes and consequences of genome reduction in animals are unclear, because our understanding of this process mostly relies on lineages with often exceptionally high rates of evolution. Here, we decode the compact 73.8 Mb genome of Dimorphilus gyrociliatus, a meiobenthic segmented worm. The D. gyrociliatus genome retains traits classically associated with larger and slower-evolving genomes, such as an ordered, intact Hox cluster, a generally conserved developmental toolkit, and traces of ancestral bilaterian linkage. Unlike some other animals with small genomes, the analysis of the D. gyrociliatus epigenome revealed canonical features of genome regulation, excluding the presence of operons and trans-splicing. Instead, the gene dense D. gyrociliatus genome presents a divergent Myc pathway, a key physiological regulator of growth, proliferation, and genome stability in animals. Altogether, our results uncover a conservative route to genome compaction in annelids, reminiscent of that observed in the vertebrate Takifugu rubripes.

genomics↗

A chromosome-scale genome assembly for the Fusarium oxysporum strain Fo5176 to establish a model Arabidopsis-fungal pathosystem

Plant pathogens cause widespread yield losses in agriculture. Understanding the drivers of plant-pathogen interactions requires decoding the molecular dialog leading to either resistance or disease. However, progress in deciphering pathogenicity genes has been severely hampered by suitable model systems and incomplete fungal genome assemblies. Here, we report a significant improvement of the assembly and annotation of the genome of the Fusarium oxysporum (Fo) strain Fo5176. Fo comprises a large number of serious plant pathogens on dozens of plant species with largely unresolved pathogenicity factors. The strain Fo5176 infects Arabidopsis thaliana and, hence, constitutes a highly promising model system. We use high-coverage Pacific Biosciences Sequel long-read and Hi-C sequencing data to assemble the genome into 19 chromosomes and a total genome size of 67.98 Mb. The genome has a N50 of 4 Mb and a 99.1% complete BUSCO score. Phylogenomic analyses based on single-copy orthologs clearly place the Fo5176 strain in the Fo f sp. conglutinans clade as expected. We generated RNAseq data from culture medium and plant infections to train gene predictions and identified [~]18,000 genes including ten effector genes known from other Fo clades. We show that Fo5176 is able to infect cabbage and Brussel sprouts of the Brassica oleracea, expanding the usefulness of the Fo5176 model pathosystem. Finally, we performed large-scale comparative genomics analyses comparing the Fo5176 to 103 additional Fo genomes to define core and accessory genomic regions. In conjunction with the molecular tool sets available for A. thaliana, the Fo5176 genome and annotation provides a crucial step towards the establishment of a highly promising pathosystem.

genomics↗

A chromosome-level assembly of the black tiger shrimp (Penaeus monodon) genome facilitates the identification of novel growth-associated genes

The black tiger shrimp (Penaeus monodon) is one of the most prominent farmed crustacean species with an average annual global production of 0.5 million tons in the last decade. To ensure sustainable and profitable production through genetic selective breeding programs, several research groups have attempted to generate a reference genome using short-read sequencing technology. However, the currently available assemblies lack the contiguity and completeness required for accurate genome annotation due to the highly repetitive nature of the genome and technical difficulty in extracting high-quality, high-molecular weight DNA in this species. Here, we report the first chromosome-level whole-genome assembly of P. monodon. The combination of long-read Pacific Biosciences (PacBio) and long-range Chicago and Hi-C technologies enabled a successful assembly of this first high-quality genome sequence. The final assembly covered 2.39 Gb (92.3% of the estimated genome size) and contained 44 pseudomolecules, corresponding to the haploid chromosome number. Repetitive elements occupied a substantial portion of the assembly (62.5%), highest of the figures reported among crustacean species. The availability of this high-quality genome assembly enabled the identification of novel genes associated with rapid growth in the black tiger shrimp through the comparison of hepatopancreas transcriptome of slow-growing and fast-growing shrimps. The results highlighted several gene groups involved in nutrient metabolism pathways and revealed 67 newly identified growth-associated genes. Our high-quality genome assembly provides an invaluable resource for accelerating the development of improved shrimp strain in breeding programs and future studies on gene regulations and comparative genomics.

genomics↗

Chromosome-level reference genome of the European wasp spider Argiope bruennichi: a resource for studies on range expansion and evolutionary adaptation

BackgroundArgiope bruennichi, the European wasp spider, has been studied intensively as to sexual selection, chemical communication, and the dynamics of rapid range expansion at a behavioral and genetic level. However, the lack of a reference genome has limited insights into the genetic basis for these phenomena. Therefore, we assembled a high-quality chromosome-level reference genome of the European wasp spider as a tool for more in-depth future studies. FindingsWe generated, de novo, a 1.67Gb genome assembly of A. bruennichi using 21.5X PacBio sequencing, polished with 30X Illumina paired-end sequencing data, and proximity ligation (Hi-C) based scaffolding. This resulted in an N50 scaffold size of 124Mb and an N50 contig size of 288kb. We found 98.4% of the genome to be contained in 13 scaffolds, fitting the expected number of chromosomes (n = 13). Analyses showed the presence of 91.1% of complete arthropod BUSCOs, indicating a high quality of the assembly. ConclusionsWe present the first chromosome-level genome assembly in the class Arachnida. With this genomic resource, we open the door for more precise and informative studies on evolution and adaptation in A. bruennichi, as well as on several interesting topics in Arachnids, such as the genomic architecture of traits, whole genome duplication and the genomic mechanisms behind silk and venom evolution.

genomics↗

Stable unmethylated DNA demarcates expressed genes and their cis-regulatory space in plant genomes

The genomic sequences of crops continue to be produced at a frenetic pace. However, it remains challenging to develop complete annotations of functional genes and regulatory elements in these genomes. Here, we explore the potential to use DNA methylation profiles to develop more complete annotations. Using leaf tissue in maize, we define [~]100,000 unmethylated regions (UMRs) that account for 5.8% of the genome; 33,375 UMRs are found greater than 2 kilobase pairs from genes. UMRs are highly stable in multiple vegetative tissues and they capture the vast majority of accessible chromatin regions from leaf tissue. However, many UMRs are not accessible in leaf (leaf-iUMRs) and these represent a set of genomic regions with potential to become accessible in specific cell types or developmental stages. Leaf-iUMRs often occur near genes that are expressed in other tissues and are enriched for transcription factor (TF) binding sites of TFs that are also not expressed in leaf tissue. The leaf-iUMRs exhibit unique chromatin modification patterns and are enriched for chromatin interactions with nearby genes. The total UMRs space in four additional monocots ranges from 80-120 megabases, which is remarkably similar considering the range in genome size of 271 megabases to 4.8 gigabases. In summary, based on the profile from a single tissue, DNA methylation signatures pinpoint both accessible regions and regions poised to become accessible or expressed in other tissues. UMRs provide powerful filters to distill large genomes down to the small fraction of putative functional genes and regulatory elements. Significance StatementCrop genomes can be very large with many repetitive elements and pseudogenes. Distilling a genome down to the relatively small fraction of regions that are functionally valuable for trait variation can be like looking for needles in a haystack. The unmethylated regions in a genome are highly stable during vegetative development and can reveal the locations of potentially expressed genes or cis-regulatory elements. This approach provides a framework towards complete annotation of genes and discovery of cis-regulatory elements using methylation profiles from only a single tissue.

genomics↗

Synteny-based genome assembly for 16 species of Heliconius butterflies, and an assessment of structural variation across the genus

Heliconius butterflies (Lepidoptera: Nymphalidae) are a group of 48 neotropical species widely studied in evolutionary research. Despite the wealth of genomic data generated in past years, chromosomal level genome assemblies currently exist for only two species, Heliconius melpomene and H. erato, each a representative of one of the two major clades of the genus. Here, we use these reference genomes to improve the contiguity of previously published draft genome assemblies of 16 Heliconius species. Using a reference-assisted scaffolding approach, we place and order the scaffolds of these genomes onto chromosomes, resulting in 95.7-99.9% of their genomes anchored to chromosomes. Genome sizes are somewhat variable among species (270-422 Mb) and in one small group of species (H. hecale, H. elevatus and H. pardalinus) differences in genome size are mainly driven by a few restricted repetitive regions. Genes within these repeat regions show an increase in exon copy number, an absence of internal stop codons, evidence of constraint on non-synonymous changes, and increased expression, all of which suggest that the extra copies are functional. Finally, we conducted a systematic search for inversions and identified five moderately large inversions fixed between the two major Heliconius clades. We infer that one of these inversions was transferred by introgression between the lineages leading to the erato/sara and burneyi/doris clades. These reference-guided assemblies represent a major improvement in Heliconius genomic resources that should aid further genetic and evolutionary studies in this genus.

genomics↗

A high-quality, chromosome-level genome assembly of the Black Soldier Fly (Hermetia Illucens L.)

BackgroundHermetia illucens L. (Diptera: Stratiomyidae), the Black Soldier Fly (BSF) is an increasingly important mass reared entomological resource for bioconversion of organic material into animal feed. ResultsWe generated a high-quality chromosome-scale genome assembly of the BSF using Pacific Bioscience, 10X Genomics linked read and high-throughput chromosome conformation capture sequencing technology. Scaffolding the final assembly with Hi-C data produced a highly contiguous 1.01 Gb genome with 99.75% of scaffolds assembled into pseudo-chromosomes representing seven chromosomes with 16.01 Mb contig and 180.46 Mb scaffold N50 values. The highly complete genome obtained a BUSCO completeness of 98.6%. We masked 67.32% of the genome as repetitive sequences and annotated a total of 17,664 protein-coding genes using the BRAKER2 pipeline. We analysed an established lab population to investigate the genomic variation and architecture of the BSF revealing six autosomes and the identification of an X chromosome. Additionally, we estimated the inbreeding coefficient (1.9%) of a lab population by assessing runs of homozygosity. This revealed a plethora of inbreeding events including recent long runs of homozygosity on chromosome five. ConclusionsRelease of this novel chromosome-scale BSF genome assembly will provide an improved platform for further genomic studies and functional characterisation of candidate regions of artificial selection. This reference sequence will provide an essential tool for future genetic modifications, functional and population genomics.

genomics↗

The de novo genome of the "Spanish" slug Arion vulgaris Moquin-Tandon, 1855 (Gastropoda: Panpulmonata): massive expansion of transposable elements in a major pest species

BackgroundThe "Spanish" slug, Arion vulgaris Moquin-Tandon, 1855, is considered to be among the 100 worst pest species in Europe. It is common and invasive to at least northern and eastern parts of Europe, probably benefitting from climate change and the modern human lifestyle. The origin and expansion of this species, the mechanisms behind its outstanding adaptive success and ability to outcompete other land slugs are worth to be explored on a genomic level. However, a high-quality chromosome-level genome is still lacking. FindingsThe final assembly of A. vulgaris was obtained by combining short reads, linked reads, Nanopore long reads, and Hi-C data. The genome assembly size is 1.54 Gb with a contig N50 length of 8.6 Mb. We found a recent expansion of transposable elements (TEs) which results in repetitive sequences accounting for more than 75% of the A. vulgaris genome, which is the highest among all known gastropod species. We identified 32,518 protein coding genes, and 2,763 species specific genes were functionally enriched in response to stimuli, nervous system and reproduction. With 1,237 single-copy orthologs from A. vulgaris and other related mollusks with whole-genome data available, we reconstructed the phylogenetic relationships of gastropods and estimated the divergence time of stylommatophoran land snails (Achatina) and Arion slugs at around 126 million years ago, and confirmed the whole genome duplication event shared by them. ConclusionsTo our knowledge, the A. vulgaris genome is the first land slug genome assembly published to date. The high-quality genomic data will provide valuable genetic resources for further phylogeographic studies of A. vulgaris origin and expansion, invasiveness, as well as molluscan aquatic-land transition and shell formation.

genomics↗

PopAmaranth: A population genetic genome browser for grain amaranths and their wild relatives

The last decades of genomic, physiological, and population genetic research have accelerated the understanding and improvement of a numerous crops. The transfer of methods to minor crops could accelerate their improvement if knowledge is effectively shared between disciplines. Grain amaranth is an ancient nutritious pseudocereal from the Americas that is regaining importance due to its high protein content and favorable amino acid and micronutrient composition. To effectively combine genomic and population genetic information with molecular genetics, plant physiology, and use it for interdisciplinary research and crop improvement, an intuitive interaction for scientists across disciplines is essential. Here, we present PopAmaranth, a population genetic genome browser, which provides an accessible representation of the genetic variation of the three grain amaranth species (A. hypochondriacus, A. cruentus, and A. caudatus) and two wild relatives (A. hybridus and A. quitensis) along the A. hypochondriacus reference sequence. We performed population-scale diversity and selection analysis from whole-genome sequencing data of 88 curated genetically and taxonomically unambiguously classified accessions. We incorporate the domestication history of the three grain amaranths to make an evolutionary perspective for candidate genes and regions available. We employ the platform to show that genetic diversity in the water stress-related MIF1 gene declined during amaranth domestication and provide evidence for convergent saponin reduction between amaranth and quinoa. These examples show that our tool enables the detailed study of individual genes, provides target regions for breeding efforts and can enhance the interdisciplinary integration of population genomic findings across species. PopAmaranth is available through amaranthGDB at amaranthgdb.org/popamaranth.html SignificanceSharing population genetic results between disciplines can facilitate interdisciplinary research and accelerate the improvement of crops. Since the onset of genome sequencing online genome browser platforms have provide access to features of an organisms genetic information. Rarely this has been extended to population-wide summary statistics for evolutionary hypothesis testing. We implemented a population genetic genome browser PopAmaranth for three grain amaranth species and their two wild relatives. The intuitive and user-friendly interface of PopA-maranth makes the genetic diversity of the species complex available to broad audience of biologists across disciplines. We show how our tool can be used to study convergence across distant genera and find signals of past selection in domestication and stress related genes. Community platforms and genome browsers are an integrative element of numerous study systems. PopAmaranth can serve as template for other research communities to integrate and share their results.

genomics↗

Population genomics of the maize pathogen Ustilago maydis: demographic history and role of virulence clusters in adaptation

The tight interaction between pathogens and their hosts results in reciprocal selective forces that impact the genetic diversity of the interacting species. The footprints of this selection differ between pathosystems because of distinct life-history traits, demographic histories, or genome architectures. Here, we studied the genome-wide patterns of genetic diversity of 22 isolates of the causative agent of the corn smut disease, Ustilago maydis, originating from five locations in Mexico, the presumed center of origin of this species. In this species, many genes encoding secreted effector proteins reside in so-called virulence clusters in the genome, an arrangement that is so far not found in other filamentous plant pathogens. Using a combination of population genomic statistical analyses, we assessed the geographical, historical and genome-wide variation of genetic diversity in this fungal pathogen. We report evidence of two partially admixed subpopulations that are only loosely associated with geographic origin. Using the multiple sequentially Markov coalescent model, we inferred the demographic history of the two pathogen subpopulations over the last 0.5 million years. We show that both populations experienced a recent strong bottleneck starting around 10,000 years ago, coinciding with the assumed time of maize domestication. While the genome average genetic diversity is low compared to other fungal pathogens, we estimated that the rate of non-synonymous adaptive substitutions is three times higher in genes located within virulence clusters compared to non-clustered genes, including non-clustered effector genes. These results highlight the role that these singular genomic regions play in the evolution of this pathogen. Significance statementThe maize pathogen Ustilago maydis is a model species to study fungal cell biology and biotrophic host-pathogen interactions. Population genetic studies of this species, however, were so far restricted to using a few molecular markers, and genome-wide comparisons involved species that diverged more than 20 million years ago. Here, we sequenced the genomes of 22 Mexican U. maydis isolates to study the recent evolutionary history of this species. We identified two co-existing populations that went through a recent bottleneck and whose divergence date overlaps with the time of maize domestication. Contrasting the patterns of genetic diversity in different categories of genes, we further showed that effector genes in virulence clusters display a high rate of adaptive mutations, highlighting the importance of these effector arrangements for the adaptation of U. maydis to its host.

genomics↗

Comparative analysis of whole genome sequences of Leptospira spp. from RefSeq database provide interspecific divergence and repertoire of virulence factors

Leptospirosis is an emerging zoonotic and neglected disease across the world causing huge loss of life and economy. The disease is caused by Leptospira of which 605 sequenced genomes representing 72 species are available in RefSeq database. A comparative genomics approach based on Average Amino acid Identity (AAI), Average Nucleotide Identity (ANI), and Insilco DNA-DNA hybridization provide insight that taxonomic and evolutionary position of few genomes needs to be changed and reclassified. Clustering on the basis of AAI of core and pan-genome contradict clustering pattern on basis of ANI into 4 clusters. Amino acid identity based hierarchical clustering clearly established 3 clusters of Leptospira correlating with level of virulence. Whole genome tree supported three cluster classifications and grouped Leptospira into three clades termed as pathogenic, intermediate and saprophytic. Leptospira genus consist of diverse species and exist in heterogeneous environment, it contains relatively large and closed core genome of 1038 genes. Analysis provided pan genome remains open with 20822 genes. COG analysis revealed that mobilome related genes were found mainly in pan-genome of pathogenic clade. Clade specific genes mined in the study can be used as marker for determining clade and associating level of virulence of any new Leptospira species. Many known Leptospira virulent genes were absent in set of 78 virulent factors mined using Virulence Factor database. A deep search approach provided a repertoire of 496 virulent genes in pan-genome. Further validation of virulent genes will help in accurately targeting pathogenic Leptospira and controlling leptospirosis. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=120 SRC="FIGDIR/small/426470v2_ufig1.gif" ALT="Figure 1"> View larger version (31K): org.highwire.dtl.DTLVardef@104babborg.highwire.dtl.DTLVardef@17f727dorg.highwire.dtl.DTLVardef@35a6c1org.highwire.dtl.DTLVardef@56fd53_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗

APOBEC Mutagenesis is Concordant Between Tumor and Viral Genomes in HPV Positive Head and Neck Squamous Cell Carcinoma

APOBEC (apolipoprotein B mRNA-editing enzyme, catalytic polypeptide-like) is a major mutagenic source in human papillomavirus positive oropharyngeal squamous cell carcinoma (HPV+ OPSCC). Why APOBEC mutations predominate in HPV+OPSCC remains an area of active investigation. Prevailing theories focus on APOBECs role as a viral restriction agent. APOBEC-induced mutations have been identified in both human cancers and HPV genomes, but whether they are directly linked in HPV+OPSCCs remains unknown. We performed sequencing of host somatic exomes, transcriptomes and HPV16 genomes from 79 HPV+ OPSCC samples, quantifying APOBEC mutational burden and activity in both the host and virus. APOBEC was the dominant mutational signature in somatic exomes. APOBEC vulnerable PIK3CA hotspot mutations were exclusively present in APOBEC enriched samples. In viral genomes, there was a mean (range) of 5 (0-29) mutations per genome. Mean (range) of APOBEC mutations in the viral genomes was 1 (0-5). Viral APOBEC mutations, compared to non-APOBEC mutations, were more likely to be low-variant allele frequency mutations, suggesting that APOBEC mutagenesis is actively occurring in viral genomes during infection. Paired host and viral analyses revealed that APOBEC-enriched tumor samples had higher viral APOBEC mutation rates (p=0.028), and APOBEC-associated RNA editing (p=0.008) suggesting that APOBEC mutagenesis in host and viral genomes are directly linked. Using paired sequencing of host somatic exomes, transcriptomes, and viral genomes from HPV+OPSCC samples, here, we show concordance between tumor and viral APOBEC mutagenesis, suggesting that APOBEC-mediated viral restriction results in off-target host-genome mutations. These data provide a missing link connecting APOBEC mutagenesis in host and virus and support a common mechanism driving APOBEC dysregulation.

genomics↗

A phased genome assembly for allele-specific analysis in Trypanosoma brucei

Many eukaryotic organisms are diploid or even polyploid, i.e. they harbour two or more independent copies of each chromosome. Yet, to date most reference genome assemblies represent a mosaic consensus sequence in which the homologous chromosomes have been collapsed into one sequence. This procedure generates sequence artefacts and impedes analyses of allele-specific mechanisms. Here, we report the allele-specific genome assembly of the diploid unicellular protozoan parasite Trypanosoma brucei. As a first step, we called variants on the allele-collapsed assembly of the T. brucei Lister 427 isolate using short-read error-corrected PacBio reads. We identified 96 thousand heterozygote variants across the genome (average of 4.2 variants / kb), and observed that the variant density along the chromosomes was highly uneven. Several long (>100 kb) regions of loss-of-heterozigosity (LOH) were identified, suggesting recent recombination events between the alleles. By analysing available genomic sequencing data of multiple Lister 427 derived clones, we found that most LOH regions were conserved, except for some that were specific to clones adapted to the insect lifecycle stage. Surprisingly, we also found that some Lister 427 clones were aneuploid. We found evidence of trisomy in chromosome five (chr 5), chr 2, chr 6 and chr 7. Moreover, by analysing RNA-seq data, we showed that the transcript level is proportional to the ploidy, evidencing the lack of a general expression control at the transcript level in T. brucei. As a second step, to generate an allele-specific genome assembly, we used two powerful datatypes for haplotype reconstruction: raw long reads (PacBio) and chromosome conformation (Hi-C) data. With this approach, we were able to assign 99.5% of all heterozygote variants to a specific homologous chromosome, building a 66 Mb long T. brucei Lister 427 allele-specific genome assembly. Hereby, we identified genes with allele-specific premature termination codons and showed that differences in allele-specific expression at the level of transcription and translation can be accurately monitored with the fully phased genome assembly. The obtained reference-grade allele-specific genome assembly of T. brucei will enable the analysis of allele-specific phenomena, as well as the better understanding of recombination and evolutionary processes. Furthermore, it will serve as a standard to benchmark much needed automatic genome assembly pipelines for highly heterozygous wild species isolates.

genomics↗

Nucleosome patterns in four plant pathogenic fungi with contrasted genome structures

AO_SCPLOWBSTRACTC_SCPLOWFungal pathogens represent a serious threat towards agriculture, health, and environment. Control of fungal diseases on crops necessitates a global understanding of fungal pathogenicity determinants and their expression during infection. Genomes of phytopathogenic fungi are often compartmentalized: the core genome contains housekeeping genes whereas the fast-evolving genome mainly contains transposable elements and species-specific genes. In this study, we analysed nucleosome landscapes of four phytopathogenic fungi with contrasted genome organizations to describe and compare nucleosome repartition patterns in relation with genome structure and gene expression level. We combined MNase-seq and RNA-seq analyses to concomitantly map nucleosome-rich and transcriptionally active regions during fungal growth in axenic culture; we developed the MNase-seq Tool Suite (MSTS) to analyse and visualise data obtained from MNase-seq experiments in combination with other genomic data and notably RNA-seq expression data. We observed different characteristics of nucleosome profiles between species, as well as between genomic regions within the same species. We further linked nucleosome repartition and gene expression. Our findings support that nucleosome positioning and occupancies are subjected to evolution, in relation with underlying genome sequence modifications. Understanding genomic organization and its role in expression regulation is the next gear to understand complex cellular mechanisms and their evolution.

genomics↗

The draft genome sequence of Eucalyptus polybractea based on hybrid assembly with short- and long-reads reads

Eucalyptus polybractea is a small, multi-stemmed tree, which is widely cultivated in Australia for the production of Eucalyptus oil. We report the hybrid assembly of the E. polybractea genome utilizing both short- and long-read technology. We generated 44 Gb of Illumina HiSeq short reads and 8 Gb of Nanopore long reads, representing approximately 83x and 15x genome coverage, respectively. The hybrid-assembled genome, after polishing, contained 24,864 scaffolds with an accumulated length of 523 Mb (N50 = 40.3 kb; BUSCO-calculated genome completeness of 94.3%). The genome contained 35,385 predicted protein-coding genes detected by combining homology-based and de novo approaches. We have provided the first assembled genome based on hybrid sequences from the highly diverse Eucalyptus subgenus Symphyomyrtus, and revealed the value of including long-reads from Nanopore technology for enhancing the contiguity of the assembled genome, as well as for improving its completeness. We anticipate that the E. polybractea genome will be an invaluable resource supporting a range of studies in genetics, population genomics and evolution of related species in Eucalyptus.

genomics↗

The complete sequence of a human genome

In 2001, Celera Genomics and the International Human Genome Sequencing Consortium published their initial drafts of the human genome, which revolutionized the field of genomics. While these drafts and the updates that followed effectively covered the euchromatic fraction of the genome, the heterochromatin and many other complex regions were left unfinished or erroneous. Addressing this remaining 8% of the genome, the Telomere-to-Telomere (T2T) Consortium has finished the first truly complete 3.055 billion base pair (bp) sequence of a human genome, representing the largest improvement to the human reference genome since its initial release. The new T2T-CHM13 reference includes gapless assemblies for all 22 autosomes plus Chromosome X, corrects numerous errors, and introduces nearly 200 million bp of novel sequence containing 2,226 paralogous gene copies, 115 of which are predicted to be protein coding. The newly completed regions include all centromeric satellite arrays and the short arms of all five acrocentric chromosomes, unlocking these complex regions of the genome to variational and functional studies for the first time.

genomics↗

Genome-wide reconstruction of rediploidization following autopolyploidizationacross one hundred million years of salmonid evolution

The long-term evolutionary impacts of whole genome duplication (WGD) are strongly influenced by the ensuing rediploidization process. Following autopolyploidization, rediploidization involves a transition from tetraploid to diploid meiotic pairing, allowing duplicated genes (ohnologues) to diverge genetically and functionally. Our understanding of autopolyploid rediploidization has been informed by a WGD event ancestral to salmonid fishes, where large genomic regions are characterized by temporally delayed rediploidization, allowing lineage-specific ohnologue sequence divergence in the major salmonid clades. Here, we investigate the long-term outcomes of autopolyploid rediploidization at genome-wide resolution, exploiting a recent explosion of salmonid genome assemblies, including a new genome sequence for the huchen (Hucho hucho). We developed a genome alignment approach to capture duplicated regions across multiple species, allowing us to create 121,864 phylogenetic trees describing ohnologue divergence across salmonid evolution. Using molecular clock analysis, we show that 61% of the ancestral salmonid genome experienced an initial wave of rediploidization in the late Cretaceous (85-106 Mya). This was followed by a period of relative genomic stasis lasting 17-39 My, where much of the genome remained in a tetraploid state. A second rediploidization wave began in the early Eocene and proceeded alongside species diversification, generating predictable patterns of lineage-specific ohnologue divergence, scaling in complexity with the number of speciation events. Finally, using gene set enrichment, gene expression, and codon-based selection analyses, we provide insights into potential functional outcomes of delayed rediploidization. Overall, this study enhances our understanding of delayed autopolyploid rediploidization and has broad implications for future studies of WGD events.

genomics↗

Whole-genome sequencing and analysis of two azaleas, Rhododendron ripense and Rhododendron kiyosumense

To enhance the genomics and genetics of azalea, the whole-genome sequences of two species of Rhododendron were determined and analyzed in this study: Rhododendron ripense, the cytoplasmic donor and ancestral species of large-flowered and evergreen azalea cultivars, respectively; and Rhododendron kiyosumense, a native of Chiba prefecture (Japan) seldomly bred and cultivated. A chromosome-level genome sequence assembly of R. ripense was constructed by single-molecule real-time (SMRT) sequencing and genetic mapping, while the genome sequence of R. kiyosumense was assembled using the single-tube long fragment read (stLFR) sequencing technology. The R. ripense genome assembly contained 319 contigs (506.7 Mb; N50 length: 2.5 Mb) and was assigned to the genetic map to establish 13 pseudomolecule sequences. On the other hand, the genome of R. kiyosumense was assembled into 32,308 contigs (601.9 Mb; N50 length: 245.7 kb). A total of 34,606 genes were predicted in the R. ripense genome, while 35,785 flower and 48,041 leaf transcript isoforms were identified in R. kiyosumense through Iso-Seq analysis. Overall, the genome sequence information generated in this study enhances our understanding of genome evolution in the Ericales and reveals the phylogenetic relationship of closely-related species. This information will also facilitate the development of phenotypically attractive azalea cultivars.

genomics↗