bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

Hybrid assembly using ultra-long reads resolves repeats and completes the genome sequence of a laboratory strain of Staphylococcus aureus subsp. aureus RN4220

Staphylococcus aureus RN4220 has been extensively used by staphylococcal researchers as an intermediate strain for genetic manipulation due to its ability to accept foreign DNA. Despite its wide use in laboratories, its complete genome is not available. In this study, we used the hybrid genome assembly approach using the minION long reads and Illumina short reads to sequence the complete genome of S. aureus RN4220. The comparative analysis of the annotated complete genome showed the presence of 39 genes fragmented in the previous assembly, many of which were located near the repeat regions. Using RNA-Seq reads, we showed that a higher number of reads could be mapped to the complete genome than the draft genome and the gene expression profile obtained using the complete genome also differs from that obtained from the draft genome. Furthermore, by comparative transcriptomic analysis, we showed the correlation between expression levels of staphyloxanthin biosynthetic genes and the production of yellow pigment. This study highlighted the importance of long reads in completing the microbial genomes, especially those possessing repetitive elements.

genomics↗

Chromosome-level genome assembly of Japanese chestnut (Castanea crenata Sieb. et Zucc.) reveals conserved chromosomal segments in woody rosids

Japanese chestnut (Castanea crenata Sieb. et Zucc.), unlike other Castanea species, is resistant to most diseases and wasps. However, genomic data of Japanese chestnut that could be used to determine its biotic stress resistance mechanisms have not been reported to date. In this study, we employed long-read sequencing and genetic mapping to generate genome sequences of Japanese chestnut at the chromosome level. Long reads (47.7 Gb; 71.6x genome coverage) were assembled into 781 contigs, with a total length of 721.2 Mb and a contig N50 length of 1.6 Mb. Genome sequences were anchored to the chestnut genetic map, comprising 14,973 single nucleotide polymorphisms (SNPs) and covering 1,807.8 cM map distance, to establish a chromosome-level genome assembly (683.8 Mb), with 69,980 potential protein-encoding genes and 425.5 Mb repetitive sequences. Furthermore, comparative genome structure analysis revealed that Japanese chestnut shares conserved chromosomal segments with woody plants, but not with herbaceous plants, of rosids. Overall, the genome sequence data of Japanese chestnut generated in this study is expected to enhance not only its genetics and genomics but also the evolutionary genomics of woody rosids.

genomics↗

High-quality genome assembly of Cinnamomum burmami (chvar. Borneol) provides insights into the natural borneol biosynthesis

Cinnamomum burmannii (chvar. Borneol) is a well-known medicinal and industrial plant cultivated in the Lingnan region of China. It is the key source from organism of natural borneol (D-borneol), one of the precious and widely used Chinese herbal medicines with a variety of medicinal effects. Here, we report a high-quality chromosome-scale genome assembly of C. burmannii (chvar. Borneol) using Pacbio single-molecule sequencing and Hi-C technology. The assembled genome size was 1.14 GB with a scaffold N50 of 94.30 Mb, while 98.77% of the assembled sequences were anchored on 12 pseudochromosomes including 41549 protein-coding genes. Genomic evolution analysis revealed C. burmannii and C. micranthum shared two Lauraceae unique ancestral whole-genome duplication (WGD) events. Likewise, comparative genomic analysis showed strong collinearity between these two species. Besides, the analysis for Long Terminal Repeat Retrotransposons (LTR-RTs) indicated the outbreak of LTR-RTs insertion made a great contribution to the size difference of genomes between C. burmannii and C. micranthum. Furthermore, the candidate genes in pathway associated with natural borneol synthesis were identified on the genome and their differential expressions were analyzed in various biological tissues. We considered that several of genes in Mevalonate (MVA) Methylerythritol Phosphate (MEP) pathways or in downstream pathway have the potential to be the key factors in the biosynthesis of D-borneol. We also constructed the genome database (CAMD; http://www.cinnamomumdatabase.com/) of Cinnamomum species for a better data utilization in the future. All these results will enrich the genomic data of Lauraceae plants and facilitate genetic improvement of this commercially important plant.

genomics↗

Genome-wide association, prediction and heritability in bacteria

Advances in whole-genome genotyping and sequencing have allowed genome-wide analyses of association, prediction and heritability in many organisms. However, the application of such analyses to bacteria is still in its infancy, being limited by difficulties including the plasticity of bacterial genomes and their strong population structure. Here we propose, and validate using simulations, a suite of genome-wide analyses for bacteria. We combine methods from human genetics and previous bacterial studies, including linear mixed models, elastic net and LD-score regression, and introduce innovations such as frequency-based allele coding, testing for both insertion/deletion and nucleotide effects and partitioning heritability by genome region. We then analyse three phenotypes of a major human pathogen Streptococcus pneumoniae, including the first analyses of minimum inhibitory concentrations (MIC) for each of two antibiotics, penicillin and ceftriaxone. We show that these are highly heritable leading to high prediction accuracy, which is explained by many genetic associations identified under good control of population structure effects. In the case of ceftriaxone MIC, these results are surprising because none of the isolates was resistant according to the inhibition zone diameter threshold. We estimate that just over half of the heritability of penicillin MIC is explained by a known drug-resistance region, which also contributes around a quarter of the heritability of ceftriaxone MIC. For the within-host survival phenotype carriage duration, no reliable associations were found but we observed moderate heritability and prediction accuracy, indicating a polygenic trait. While generating important new results for S. pneumoniae, we have critically assessed existing methods and introduced innovations that will be useful for future large-scale population genomics studies to help decipher the genetic architecture of bacterial traits. Author summaryGenome-wide association, prediction and heritability analyses in bacteria are beginning to help unravel the genetic underpinnings of traits such as antimicrobial resistance, virulence, within-host survival and transmissibility. Progress to date is limited by challenges including the effects of strong population structure and variable recombination, and the many gaps in sequence alignments including the absence of entire genes in many isolates. More work is required to critically asses and develop methods for bacterial genomics. We address this task here, using a range of existing methods from bacterial and human genetics, such as linear mixed models, elastic net and LD-score regression. Using simulations, we first validate and then adapt these methods to introduce new analyses, including separate assessment of gap and nucleotide effects, a new allele coding for association analyses and a method to partition heritability into genome regions. We analyse within-host survival and two antimicrobial response traits of Streptococcus pneumoniae, identifying many novel associations while demonstrating good control of population structure and accurate prediction. We present both new results for an important pathogen and methodological advances that will be useful in guiding future studies in bacterial population genomics.

genomics↗

Genome assembly of the Australian black tiger shrimp (Penaeus monodon) reveals a fragmented IHHNV EVE sequence

Shrimp are a valuable aquaculture species globally; however, disease remains a major hindrance to shrimp aquaculture sustainability and growth. Mechanisms mediated by endogenous viral elements (EVEs) have been proposed as a means by which shrimp that encounter a new virus start to accommodate rather than succumb to infection over time. However, evidence on the nature of such EVEs and how they mediate viral accommodation is limited. More extensive genomic data on Penaeid shrimp from different geographical locations should assist in exposing the diversity of EVEs. In this context, reported here is a PacBio Sequel-based draft genome assembly of an Australian black tiger shrimp (Penaeus monodon) inbred for one generation. The 1.89 Gbp draft genome is comprised of 31,922 scaffolds (N50: 496,398 bp) covering 85.9% of the projected genome size. The genome repeat content (61.8% with 30% representing simple sequence repeats) is almost the highest identified for any species. The functional annotation identified 35,517 gene models, of which 25,809 were protein-coding and 17,158 were annotated using interproscan. Scaffold scanning for specific EVEs identified an element comprised of a 9,045 bp stretch of repeated, inverted and jumbled genome fragments of Infectious hypodermal and hematopoietic necrosis virus (IHHNV) bounded by a repeated 591/590 bp host sequence. As only near complete linear ~4 kb IHHNV genomes have been found integrated in the genome of P. monodon previously, its discovery has implications regarding the validity of PCR tests designed to specifically detect such linear EVE types. The existence of joined inverted IHHNV genome fragments also provides a means by which hairpin dsRNAs could be expressed and processed by the shrimp RNA interference (RNAi) machinery.

genomics↗

Genomic architecture controls spatial structuring in Amazonian birds

Large rivers are ubiquitously invoked to explain the distributional limits and speciation of the Amazon Basins mega-diversity. However, inferences on the spatial and temporal origins of Amazonian species have narrowly focused on evolutionary neutral models, ignoring the potential role of natural selection and intrinsic genomic processes known to produce heterogeneity in differentiation across the genome. To test how genomic architecture impacts our ability to reconstruct patterns of spatial diversification across multiple taxa, we sequenced whole genomes for populations of bird species that co-occur in southeastern Amazonian. We found that phylogenetic relationships within species and demographic parameters varied across the genome in predictable ways. Genetic diversity was positively associated with recombination rate and negatively associated with the species tree topology weight. Gene flow was less pervasive in regions of low recombination, making these windows more likely to retain patterns of population structuring that matched the species tree. We further found that approximately a third of the genome showed evidence of selective sweeps and linked selection skewing genome-wide estimates of effective population sizes and gene flow between populations towards lower values. In sum, we showed that the effects of intrinsic genomic characteristics and selection can be disentangled from the neutral processes to elucidate how speciation hypotheses and biogeographic patterns are sensitive to genomic architecture.

genomics↗

The 3D architecture of the pepper (Capsicum annum) genome and its relationship to function and evolution

The architecture of topologically associating domains (TADs) varies across plant genomes. Understanding the functional consequences of this diversity requires insights into the pattern, structure, and function of TADs. Here, we present a comprehensive investigation of the 3D genome organization of pepper (Capsicum annuum) and its association with gene expression and genomic variants. We report the first chromosome-scale long-read genome assembly of pepper and generate Hi-C contact maps for four tissues. The contact maps indicate that 3D structure varies somewhat across tissues, but generally the genome was segregated into subcompartments that were correlated with transcriptional state. In addition, chromosomes were almost continuously spanned by TADs, with the most prominent found in large genomic regions that were rich in retrotransposons. A substantial fraction of TAD boundaries were demarcated by chromatin loops, suggesting loop extrusion is a major mechanism for TAD formation; many of these loops were bordered by genes, especially in highly repetitive regions, resulting in gene clustering in three dimensional space. Integrated analysis of Hi-C profiles and transcriptomes showed that change in 3D chromatin structures (e.g. subcompartments, TADs, and loops) was not the primary mechanism contributing to differential gene expression between tissues, but chromatin structure does play a role in transcription stability. TAD boundaries were significantly enriched for breaks of synteny and depletion of sequence variation, suggesting that TADs constrain patterns of genome structural evolution in plants. Together, our work provides insights into principles of 3D genome folding in large plant genomes and its association with function and evolution.

genomics↗

Chromosome-scale haplotype-phased genome assemblies of the male and female lines of wild asparagus (Asparagus kiusianus), a dioecious plant species

Asparagus kiusianus is a disease-resistant dioecious plant species and a wild relative of garden asparagus (A. officinalis). To enhance A. kiusianus genomic resources, advance plant science, and facilitate asparagus breeding, we determined the genome sequences of the male and female lines of A. kiusianus. Genome sequence reads obtained with a linked-read technology were assembled into four haplotype-phased contig sequences (~1.6 Gb each) for the male and female lines. The contig sequences were aligned onto the chromosome sequences of garden asparagus to construct pseudomolecule sequences. Approximately 55,000 potential protein-encoding genes were predicted in each genome assembly, and ~70% of the genome sequence was annotated as repetitive. Comparative analysis of the genomes of the two species revealed structural and sequence variants between the two species as well as between the male and female lines of each species. Genes with high sequence similarity with the male-specific sex determinant gene in A. officinalis, MSE1/AoMYB35/AspTDF1, were presented in the genomes of the male line but absent from the female genome assemblies. Overall, the genome sequence assemblies, gene sequences, and structural and sequence variants determined in this study will reveal the genetic mechanisms underlying sexual differentiation in plants, and will accelerate disease-resistance breeding in garden asparagus.

genomics↗

Profiling of the most reliable mutations from sequenced SARS-CoV-2 genomes scattered in Uzbekistan

Due to rapid mutations in the coronavirus genome over time and re-emergence of multiple novel variants of concerns (VOC), there is a continuous need for a periodic genome sequencing of SARS-CoV-2 genotypes of particular region. This is for on-time development of diagnostics, monitoring and therapeutic tools against virus in the global pandemics condition. Toward this goal, we have generated 18 high-quality whole-genome sequence data from 32 SARS-CoV-2 genotypes of PCR-positive COVID-19 patients, sampled from the Tashkent region of Uzbekistan. The nucleotide polymorphisms in the sequenced sample genomes were determined, including nonsynonymous (missense) and synonymous mutations in coding regions of coronavirus genome. Phylogenetic analysis grouped fourteen whole genome sample sequences (1, 2, 4, 5, 8, 10-15, 17, 32) into the G clade (or GR sub-clade) and four whole genome sample sequences (3, 6, 25, 27) into the S clade. A total of 128 mutations were identified, consisting of 45 shared and 83 unique mutations. Collectively, nucleotide changes represented one unique frameshift mutation, four upstream region mutations, six downstream region mutations, 50 synonymous mutations, and 67 missense mutations. The sequence data, presented herein, is the first coronavirus genomic sequence data from the Republic of Uzbekistan, which should contribute to enrich the global coronavirus sequence database, helping in future comparative studies. More importantly, the sequenced genomic data of coronavirus genotypes of this study should be useful for comparisons, diagnostics, monitoring, and therapeutics of COVID-19 disease in local and regional levels.

genomics↗

Complete sequence of a 641-kb insertion of mitochondrial DNA in the Arabidopsis thaliana nuclear genome

Intracellular transfers of mitochondrial DNA continue to shape nuclear genomes. Chromosome 2 of the model plant Arabidopsis thaliana contains one of the largest known nuclear insertions of mitochondrial DNA (numts). Estimated at over 600 kb in size, this numt is larger than the entire Arabidopsis mitochondrial genome. The primary Arabidopsis nuclear reference genome contains less than half of the numt because of its structural complexity and repetitiveness. Recent datasets generated with improved long-read sequencing technologies (PacBio HiFi) provide an opportunity to finally determine the accurate sequence and structure of this numt. We performed a de novo assembly using sequencing data from recent initiatives to span the Arabidopsis centromeres, producing a gap-free sequence of the Chromosome 2 numt, which is 641-kb in length and has 99.933% nucleotide sequence identity with the actual mitochondrial genome. The numt assembly is consistent with the repetitive structure previously predicted from fiber-based fluorescent in situ hybridization. Nanopore sequencing data indicate that the numt has high levels of cytosine methylation, helping to explain its biased spectrum of nucleotide sequence divergence and supporting previous inferences that it is transcriptionally inactive. The original numt insertion appears to have involved multiple mitochondrial DNA copies with alternative structures that subsequently underwent an additional duplication event within the nuclear genome. This work provides insights into numt evolution, addresses one of the last unresolved regions of the Arabidopsis reference genome, and represents a resource for distinguishing between highly similar numt and mitochondrial sequences in studies of transcription, epigenetic modifications, and de novo mutations. Significance statementNuclear genomes are riddled with insertions of mitochondrial DNA. The model plant Arabidopsis has one of largest of these insertions ever identified, which at over 600-kb in size represents one of the last unresolved regions in the Arabidopsis genome more than 20 years after the insertion was first identified. This study reports the complete sequence of this region, providing insights into the origins and subsequent evolution of the mitochondrial DNA insertion and a resource for distinguishing between the actual mitochondrial genome and this nuclear copy in functional studies.

genomics↗

Anticodon Table of the Chloroplast Genome and Identification of Putative Quadruplet Anticodons in Chloroplast tRNAs

The chloroplast genome of 5959 species was analyzed to construct the anticodon table of the chloroplast genome. Analysis of the chloroplast transfer ribonucleic acid (tRNA) revealed the presence of a putative quadruplet anticodon containing tRNAs in the chloroplast genome. The tRNAs with putative quadruplet anticodons were UAUG, UGGG, AUAA, GCUA, and GUUA, where the GUUA anticodon putatively encoded tRNAAsn. The study also revealed the complete absence of tRNA genes containing ACU, CUG, GCG, CUC, CCC, and CGG anticodons in the chloroplast genome from the species studied so far. The chloroplast genome was also found to encode tRNAs encoding N-formylmethionine (fMet), Ile2, selenocysteine, and pyrrolysine. The chloroplast genomes of mycoparasitic and heterotrophic plants have had heavy losses of tRNA genes. Furthermore, the chloroplast genome was also found to encode putative spacer tRNA, tRNA fragments (tRFs), tRNA-derived, stress-induced RNA (tiRNAs), and group I introns. An evolutionary analysis revealed that chloroplast tRNAs had evolved via multiple common ancestors and the GC% had more influence toward encoding the tRNA number in the chloroplast genome compared to the genome size.

genomics↗

Draft genome of six Cuban Anolis lizards and insights into genetic changes during the diversification

The detection of various type of genomic variants and their accumulation processes during species diversification and adaptive radiation is important for understanding the molecular and genetic basis of evolution. Anolis lizards in the West Indies are good models for studying the mechanism of the evolution because of the repeated evolution of their morphology and the ecology. In this study, we performed de novo genome assembly of six Cuban Anolis lizards with different ecomorphs and thermal habitats (Anolis isolepis, Anolis allisoni, Anolis porcatus, Anolis allogus, Anolis homolechis, and Anolis sagrei). As a result, we obtained six novel draft genomes with relatively long and high gene completeness, with scaffold N50 ranging from 5.56-39.79 Mb, and vertebrate Benchmarking Universal Single-Copy Orthologs completeness ranging from 77.5% to 86.9%. Subsequently, we performed comparative analysis of genomic contents including those of mainland Anolis lizards to estimate genetic variations that had emerged and accumulated during the diversification of Anolis lizards. Comparing the repeat element compositions and repeat landscapes revealed differences in the accumulation process between Cuban trunk-crown and trunk-ground species, LTR accumulation observed only in A. carolinensis, and separate expansions of several families of LINE in each of Cuban trunk-ground species. The analysis of duplicated genes suggested that the proportional difference of duplicated gene number among Cuban Anolis lizards may be associated to the difference of their habitat range. Furthermore, Pairwise Sequentially Markovian Coalescent analysis proposed that the effective population sizes of each species might have been affected by Cubas geohistory. Hence, these six novel draft genome assemblies and detected genetic variations can be a springboard for the further genetic elucidation of the Anolis lizards diversification. SignificanceAnolis lizard in the West Indies is excellent model for studying the mechanisms of speciation and adaptive evolution. Still, due to a lack of genome assemblies, genetic variations and accumulation process of them involved in the diversification remain largely unexplored. In this study, we reported the novel genome assemblies of six Cuban Anolis lizards and analyzed evolution of genome contents. From comparative genomic analysis and inferences of genetic variation accumulation process, we detected species- and lineage-specific transposon accumulation processes and gene copy number evolution, considered to be associated with the adaptation to their habitats. Additionally, we estimated past effective population sizes and the results suggested its relationship to Cubas geohistory.

genomics↗

Pan-cancer whole genome comparison of primary and metastatic solid tumors

Metastatic cancer remains almost inevitably a lethal disease. A better understanding of disease progression and response to therapies therefore remains of utmost importance. Here, we characterize the genomic differences between early-stage untreated primary tumors and late-stage treated metastatic tumors using a harmonized pan-cancer (re-)analysis of 7,152 whole-genome-sequenced tumors. In general, our analysis shows that metastatic tumors have a low intra-tumor heterogeneity, high genomic instability and increased frequency of structural variants with comparatively a modest increase in the number of small genetic variants. However, these differences are cancer type specific and are heavily impacted by the exposure to cancer therapies. Five cancer types, namely breast, prostate, thyroid, kidney clear carcinoma and pancreatic neuroendocrine, are a clear exception to the rule, displaying an extensive transformation of their genomic landscape in advanced stages. These changes were supported by increased genomic instability and involved substantial differences in tumor mutation burden, clock-based molecular signatures and the landscape of driver alterations as well as a pervasive increase in structural variant burden. The majority of cancer types had either moderate genomic differences (e.g., cervical and colorectal cancers) or highly consistent genomic portraits (e.g., ovarian cancer and skin melanoma) when comparing early- and late-stage disease. Exposure to treatment further scars the tumor genome and introduces an evolutionary bottleneck that selects for known therapy-resistant drivers in approximately half of treated patients. Our data showcases the potential of whole-genome analysis to understand tumor evolution and provides a valuable resource to further investigate the biological basis of cancer and resistance to cancer therapies.

genomics↗

Chromosome level reference genome for European flat oyster (Ostrea edulis L.)

The European flat oyster (Ostrea edulis L.) is a bivalve naturally distributed across Europe that was an integral part of human diets for centuries, until anthropogenic activities and disease outbreaks severely reduced wild populations. Despite a growing interest in genetic applications to support population management and aquaculture, a reference genome for this species is lacking to date. Here we report a chromosome-level assembly and annotation for the European Flat oyster genome, generated using Oxford Nanopore, Illumina, Dovetail OmniC proximity ligation and RNA sequencing. A contig assembly (N50: 2.38Mb) was scaffolded into the expected karyotype of 10 pseudo-chromosomes. The final assembly is 935.13 Mb, with a scaffold-N50 of 95.56 Mb, with a predicted repeat landscape dominated by unclassified elements specific to O. edulis. The assembly was verified for accuracy and completeness using multiple approaches, including a novel linkage map built with ddRAD-Seq technology, comprising 4,016 SNPs from four full-sib families (8 parents and 163 F1 offspring). Annotation of the genome integrating multi-tissue transcriptome data, comparative protein evidence and ab-initio gene prediction identified 35,699 protein-coding genes. Chromosome level synteny was demonstrated against multiple high-quality bivalve genome assemblies, including an O. edulis genome generated independently for a French O. edulis individual. Comparative genomics was used to characterize gene family expansions during Ostrea evolution that potentially facilitated adaptation. This new reference genome for European flat oyster will enable high-resolution genomics in support of conservation and aquaculture initiatives, and improves our understanding of bivalve genome evolution.

genomics↗

Comparative genomics of tarakihi (Nemadactylus macropterus) and five New Zealand fish species: assembly contiguity affects the identification of genic features but not transposable elements

Comparative analysis of whole-genome sequences can provide valuable insights into the evolutionary patterns of diversification and adaptation of species, including the genome contents and the regions under selection. However, such studies are lacking for fishes in New Zealand. To supplement the recently sequenced genome of tarakihi (Nemadactylus macropterus), the genomes of five additional percomorph species native to New Zealand (king tarakihi (Nemadactylus n.sp.), blue moki (Latridopsis ciliaris), butterfish (Odax pullus), barracouta (Thyrsites atun), and kahawai (Arripis trutta)) were determined and assembled using Illumina sequencing. While the proportion of repeat elements was highly correlated with the genome size (R2 = 0.97, P < 0.01), most of the metrics for the genic features (e.g. number of exons or intron length) were significantly correlated with assembly contiguity (| R2| = 0.79-0.97). A phylogenomic tree including eight additional high-quality fish genomes was reconstructed from sequences of shared gene families. The radiation of Percomorpha was estimated to have occurred c. 112 mya (mid-Cretaceous), while the Latridae have diverged from true Perciformes c. 83 mya (late Cretaceous). Evidence of positive selection was found in 65 genes in tarakihi and 209 genes in Latridae: the largest portion of these are involved in the ATP binding pathway and the integral structure of membranes. These results and the de novo genome sequences can be used to (1) inform future studies on both the strength and shortcomings of scaffold-level assemblies for comparative genomics and (2) provide insights into the evolutionary patterns and processes of genome evolution in bony fishes.

genomics↗

A high-quality chromosome-level genome assembly of rohu carp, Labeo rohita, and its utilization in SNP-based exploration of gene flow and sex determination

Labeo rohita (rohu) is a carp important to aquaculture in South Asia, with a production volume close to Atlantic salmon. While genetic improvements to rohu are ongoing, the genomic methods commonly used in other aquaculture improvement programs have historically been precluded in rohu, partially due to the lack of a high quality reference genome. Here we present a high-quality de novo genome produced using a combination of next-generation sequencing technologies, resulting in a 946 Mb genome consisting of 25 chromosomes and 2,844 unplaced scaffolds. Notably, while approximately half the size of the existing genome sequence, our genome represents 97.9% of the genome size newly estimated here using flow cytometry. Sequencing from 120 individuals was used in conjunction with this genome to predict the population structure, diversity, and divergence in three major rivers (Jamuna, Padma, and Halda), in addition to infer a likely sex determination mechanism in rohu. These results demonstrate the utility of the new rohu genome in modernizing some aspects of rohu genetic improvement programs.

genomics↗

Massively parallel characterization of insulator activity across the genome

Insulators are cis-regulatory sequences (CRSs) that can block enhancers from activating target promoters or act as barriers to block the spread of heterochromatin. Their name derives from their ability to insulate transgenes from genomic position effects, an important function in gene therapy and biotechnology applications that require high levels of sustained transgene expression. In theory, flanking transgenes with insulators protects them from position effects, but in practice, efforts to insulate transgenes meet with mixed success because the contextual requirements for insulator function in the genome are not well understood. A key question is whether insulators are modular elements that can function anywhere in the genome or whether they are adapted to function only in certain genomic locations. To distinguish between these two possibilities we developed MPIRE (Massively Parallel Integrated Regulatory Elements) and used it to measure the effects of three insulators (A2, cHS4, ALOXE3) and their mutants at thousands of locations across the genome. Our results show that each insulator functions in only a small number of genomic locations, and that insulator function depends on the sequence motifs that comprise each insulator. All three insulators can block enhancers in the genome, but specificity arises because each insulator blocks enhancers that are bound by different sets of transcription factors. All three insulators can block enhancers in the genome, but only ALOXE3 can act as a heterochromatin barrier. We conclude that insulator function is highly context dependent and that MPIRE is a robust and systematic method for revealing the context dependencies of insulators and other cis-regulatory elements across the genome.

genomics↗

Gap-free nuclear and mitochondrial genomes of Fusarium verticillioides strain HN2

Fusarium ear rot (FER) and Fusarium stalk rot (FSR) caused by the filamentous fungus Fusarium verticillioides have become increasingly serious around the world. Additionally, fumonisins produced by F. verticillioides threaten food and feed security. By adding the contribution of genomic resources to better understand the pathosystem including the mechanisms of F. verticillioides-maize interactions, and further improving the quality of the F. verticillioides genome, the gap-free nuclear genome and mitochondrial genome of F. verticillioides strain HN2 were sequenced and assembled. Using Oxford Nanopore long reads and next-generation sequencing short reads, the final 42.81-Mb genome was assembled into 12 contigs (N50 = 4.16-Mb). A total of 13,466 protein-coding genes were annotated, including 1,076 secreted proteins that contain 342 candidate effectors. In addition, we assembled the complete 53,764 bp mitochondrial genome. F. verticillioides strain 7600 genome assemblies are fragmented and high-quality reference genomes were needed. The genomes presented here will serve as an important resource for F. verticillioides research.

genomics↗