bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

OrganellarGenomeDRAW (OGDRAW) version 1.3.1: expanded toolkit for the graphical visualization of organellar genomes

Key points O_LIOGDRAW has become the standard tool for displaying maps of organellar genomes C_LIO_LIit converts GenBank entries into graphical maps C_LIO_LIa new version with improved functionality has been released C_LI AbstractOrganellar (plastid and mitochondrial) genomes play an important role in resolving phylogenetic relationships, and next-generation sequencing technologies have led to a burst in their availability. The ongoing massive sequencing efforts require software tools for routine assembly and annotation of organellar genomes as well as their display as physical maps. OrganellarGenomeDRAW (OGDRAW) has become the standard tool to draw graphical maps of plastid and mitochondrial genomes. Here were present a new version of OGDRAW equipped with a new front end. Besides several new features, OGDRAW has now access to a local copy of the organelle genome database of the NCBI RefSeq project. Together with batch processing of (multi-)GenBank files, this enables the user to easily visualize large sets of organellar genomes spanning entire taxonomic clades. The new OGDRAW server can be accessed at https://chlorobox.mpimp-golm.mpg.de/OGDraw.html.

bioinformatics

High resolution single-cell chromatin 3D modeling reveals coherent chromatin aggregation with varied structures in controlling genome function stability

The genome 3D architecture is thought to be related to regulating gene expression levels in cells and can be explained by genome-wide chromatin interactions which have been explored by chromosome conformation capture based techniques, especially Hi-C. Based on single-cell Hi-C data, we developed a new method in constructing experimental consistent 3D intact genome structures for individual cells with a resolution of 10kb or higher. The modeled structures showed marked variations of 3D genome organization across different cells. However, chromosome loci marked by different proteins, such as CTCF and post-translationally modified histones, are consistently non-specifically aggregated in space. Interestingly, similar aggregations between active enhancers and active promoters were observed, especially for those separated by genomic regions of the scale of megabase or larger. Such long-range associations between active enhancers and promoters are strongly correlated with spatial aggregation of chromosome loci marked by different proteins. Through analyzing the 3D structures of intact genome, we proposed that coherent gene activation profiles among individual cells can be achieved by the consistent aggregation of protein marked loci instead of maintaining identical folded conformations.

biophysics

Scalable multiple whole-genome alignment and locally collinear block construction with SibeliaZ

Multiple whole-genome alignment is a challenging problem in bioinformatics. Despite many successes, current methods are not able to keep up with the growing number, length, and complexity of assembled genomes, especially when computational resources are limited. Approaches based on compacted de Bruijn graphs to identify and extend anchors into locally collinear blocks have potential for scalability, but current methods do not scale to mammalian genomes. We present an algorithm, SibeliaZ-LCB, for identifying collinear blocks in closely related genomes based on analysis of the de Bruijn graph. We further incorporate this into a multiple whole-genome alignment pipeline called SibeliaZ. SibeliaZ shows run-time improvements over other methods while maintaining accuracy. On sixteen recently-assembled strains of mice, SibeliaZ runs in under 16 hours on a single machine, while other tools did not run to completion for eight mice within a week. SibeliaZ makes a significant step towards improving scalability of multiple whole-genome alignment and collinear block reconstruction algorithms on a single machine.

bioinformatics

Enterococcus faecium genome dynamics during long-term asymptomatic patient gut colonization

BackgroundE. faecium is a gut commensal of humans and animals. In addition, it has recently emerged as an important nosocomial pathogen through the acquisition of genetic elements that confer resistance to antibiotics and virulence. We performed a whole-genome sequencing based study on 96 multidrug-resistant E. faecium strains that asymptomatically colonized five patients with the aim to describe the genome dynamics of this species. ResultsThe patients were hospitalized on multiple occasions and isolates were collected over periods ranging from 15 months to 6.5 years. Ninety-five of the sequenced isolates belonged to E. faecium clade A1, which was previously determined to be responsible for the vast majority of clinical infections. The clade A1 strains clustered into six clonal groups of highly similar isolates, three of which entirely consisted of isolates from a single patient. We also found evidence of concurrent colonization of patients by multiple distinct lineages and transfer of strains between patients during hospitalisation. We estimated the evolutionary rate of two clonal groups that colonized a single patient at 12.6 and 25.2 single nucleotide polymorphisms (SNPs)/genome/year. A detailed analysis of the accessory genome of one of the clonal groups revealed considerable variation due to gene gain and loss events, including the chromosomal acquisition of a 37 kbp prophage and the loss of an element containing carbohydrate metabolism-related genes. We determined the presence and location of twelve different Insertion Sequence (IS) elements, with ISEfa5 showing a unique pattern of location in 24 of the 25 isolates, suggesting widespread ISEfa5 excision and insertion into the genome during gut colonization. ConclusionsOur findings show that the E. faecium genome is highly dynamic during asymptomatic colonization of the patient gut. We observe considerable genomic flexibility due to frequent horizontal gene transfer and recombination, which can contribute to the generation of genetic diversity within the species and, ultimately, can contribute to its success as a nosocomial pathogen.

microbiology

Dynamic Scan Procedure for Detecting Rare-Variant Association Regions in Whole Genome Sequencing Studies

Whole genome sequencing (WGS) studies are being widely conducted to identify rare variants associated with human diseases and disease-related traits. Classical single-marker association analyses for rare variants have limited power, and variant-set based analyses are commonly used to analyze rare variants. However, existing variant-set based approaches need to pre-specify genetic regions for analysis, and hence are not directly applicable to WGS data due to the large number of intergenic and intron regions that consist of a massive number of non-coding variants. The commonly used sliding window method requires pre-specifying fixed window sizes, which are often unknown as a priori, are difficult to specify in practice and are subject to limitations given genetic association region sizes are likely to vary across the genome and phenotypes. We propose a computationally-efficient and dynamic scan statistic method (Scan the Genome (SCANG)) for analyzing WGS data that flexibly detects the sizes and the locations of rare-variants association regions without the need of specifying a prior fixed window size. The proposed method controls the genome-wise type I error rate and accounts for the linkage disequilibrium among genetic variants. It allows the detected rare variants association region sizes to vary across the genome. Through extensive simulated studies that consider a wide variety of scenarios, we show that SCANG substantially outperforms several alternative rare-variant association detection methods while controlling for the genome-wise type I error rates. We illustrate SCANG by analyzing the WGS lipids data from the Atherosclerosis Risk in Communities (ARIC) study.

genetics

Reconstruction of clone- and haplotype-specific cancer genome karyotypes from bulk tumor samples

Many cancer genomes are extensively rearranged with highly aberrant chromosomal karyotypes. These genome rearrangements, or structural variants, can be detected in tumor DNA sequencing data by abnormal mapping of se-quence reads to the reference genome. However, nearly all cancer sequencing to date is of bulk tumor samples which consist of a heterogeneous mixture of normal cells and subpopulations of cancers cells, or clones, that harbor distinct somatic structural variants. We introduce a novel algorithm, Reconstructing Cancer Karyotypes (RCK), to reconstruct haplotype-specific karyotypes of one or more rearranged cancer genomes, or clones, that best explain the read alignments from a bulk tumor sample. RCK leverages specific evolutionary constraints on the somatic mutation process in cancer to reduce ambiguity in the deconvolution of admixed DNA sequence data into multiple haplotype-specific cancer karyotypes. In particular, RCK relies on generalizations of the infinite sites assumption that a genome re-arrangement is highly unlikely to occur at the same nucleotide position more than once during somatic evolution. RCKs comprehensive model allows us to incorporate information both from short and long-read sequencing technologies and is applicable to bulk tumor samples containing a mixture of an arbitrary number of derived genomes. We compared RCK to the state-of-the-art method ReMixT on a dataset of 17 primary and metastatic prostate cancer samples. We demonstrate that ReMixTs limited support for heterogeneity and lack of evolutionary constrains leads to reconstruction of implausible karyotypes. In contrast, RCKs infers cancer karyotypes that better explain read alignments from bulk tumor samples and are consistent with a reasonable evolutionary model. RCKs reconstructions of clone- and haplotype-specific karyotypes will aid further studies of the role of intra-tumor heterogeneity in cancer development and response to treatment. RCK is available at https://github.com/raphael-group/RCK.

cancer biology

GCNA preserves genome integrity and fertility across species

The propagation of species depends on the ability of germ cells to protect their genome in the face of numerous exogenous and endogenous threats. While these cells employ a number of known repair pathways, specialized mechanisms that ensure high-fidelity replication, chromosome segregation, and repair of germ cell genomes remain incompletely understood. Here, we identify Germ Cell Nuclear Acidic Peptidase (GCNA) as a highly conserved regulator of genome stability in flies, worms, zebrafish, and humans. GCNA contains a long acidic intrinsically disordered region (IDR) and a protease-like SprT domain. In addition to chromosomal instability and replication stress, GCNA mutants accumulate DNA-protein crosslinks (DPCs). GCNA acts in parallel with a second SprT domain protein Spartan. Structural analysis reveals that while the SprT domain is needed to limit meiotic and replicative damage, most of GCNAs function maps to its IDR. This work shows GCNA protects germ cells from various sources of damage, providing novel insights into conserved mechanisms that promote genome integrity across generations.\n\nHighlightsGCNA ensures genomic stability in germ cells and early embryos across species\n\nGCNA limits replication stress and DNA double stranded breaks\n\nGCNA restricts DNA-Protein Crosslinks within germ cells and early embryos\n\nThe IDR and SprT domains of GCNA govern distinct aspects of genome integrity\n\nGraphic Abstract\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=189 SRC=\"FIGDIR/small/570804_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (54K):\norg.highwire.dtl.DTLVardef@1928d82org.highwire.dtl.DTLVardef@885ff2org.highwire.dtl.DTLVardef@15322c9org.highwire.dtl.DTLVardef@110c615_HPS_FORMAT_FIGEXP M_FIG C_FIG

developmental biology

The exhaustive genomic scan approach, with an application to rare-variant association analysis

Region-based genome-wide scans are usually performed by use of a priori chosen analysis regions. Such an approach will likely miss the region comprising the strongest signal and, thus, may result in increased type II error rates and decreased power. Here, we propose a genomic exhaustive scan approach that analyzes all possible subsequences and does not rely on a prior definition of the analysis regions. As a prime instance, we present a computationally ultra-efficient implementation using the rare-variant collapsing test for phenotypic association, the genomic exhaustive collapsing scan (GECS). Our implementation allows for the identification of regions comprising the strongest signals in large, genome-wide rare-variant association studies while controlling the family-wise error rate via permutation. Application of GECS to two genomic data sets revealed several novel significantly associated regions for age-related macular degeneration and for schizophrenia. Our approach also offers a high potential for genome-wide scans for selection, methylation and other analyses.

genetics

A catalog of CasX genome editing sites in common model organisms

DpbCasX, also called Cas12e, is an RNA-guided DNA endonuclease isolated from Deltaproteobacteria. In this paper I characterized the CasX-compatible genome editing sites in the reference genomes of yeast (Saccharomyces cerevisiae), flatworms (Caenorhabditis elegans), flies (Drosophila melanogaster), zebrafish (Danio rerio), mouse (Mus musculus), rats (Rattus norvegicus), and humans (Homo sapiens). Across those genomes there were >27,000 CasX sites per megabase on average. More than 90% of genes in each genome had at least one unique site overlapping an exon, with median unique sites per gene of 6 - 45. I also annotated sites in the GRCm38 reference and 15 additional mouse strain genomes. The presence of specific guide sequences varied amongst the strains, with CAST/EiJ and PWK/PhJ showing the greatest divergence from the reference strain. The high density of CasX sites and number of exon overlapping sites suggests that CasX has the potential to be used as a common genome editor.

bioinformatics

OMGS: Optical Map-based Genome Scaffolding

Due to the current limitations of sequencing technologies, de novo genome assembly is typically carried out in two stages, namely contig (sequence) assembly and scaffolding. While scaffolding is computationally easier than sequence assembly, the scaffolding problem can be challenging due to the high repetitive content of eukaryotic genomes, possible mis-joins in assembled contigs and inaccuracies in the linkage information. Genome scaffolding tools either use paired-end/mate-pair/linked/Hi-C reads or genome-wide maps (optical, physical or genetic) as linkage information. Optical maps (in particular Bionano Genomics maps) have been extensively used in many recent large-scale genome assembly projects (e.g., goat, apple, barley, maize, quinoa, sea bass, among others). However, the most commonly used scaffolding tools have a serious limitation: they can only deal with one optical map at a time, forcing users to alternate or iterate over multiple maps. In this paper, we introduce a novel scaffolding algorithm called OMGS that for the first time can take advantages of multiple optical maps. OMGS solves several optimization problems to generate scaffolds with optimal contiguity and correctness. Extensive experimental results demonstrate that our tool outperforms existing methods when multiple optical maps are available, and produces comparable scaffolds using a single optical map. OMGS can be obtained from https://github.com/ucrbioinfo/OMGS

bioinformatics

Revealing virulence potential of clinical and environmental Aspergillus fumigatus isolates using Whole-genome sequencing

Aspergillus fumigatus is an opportunistic airborne pathogen and one of the most common causative agents of human fungal infections. A restricted number of virulence factors have been described but none of them lead to a differentiation of the virulence level among different strains. In this study, we analyzed the whole-genome sequence of a set of A. fumigatus isolates from clinical and environmental origin to compare their genomes and to determine their virulence profiles. For this purpose, a database containing 244 genes known to be associated with virulence was built. The genes were classified according to their biological function into factors involved in thermotolerance, resistance to immune responses, cell wall structure, toxins and secondary metabolites, allergens, nutrient uptake and signaling and regulation. No difference in virulence profiles was found between clinical isolates causing an infection and a colonizing clinical isolate, nor between isolates from clinical and environmental origin. We observed the presence of genetic repetitive elements located next to virulence related gene groups, which could potentially influence their regulation. In conclusion, our genomic analysis reveals that A. fumigatus, independently of their source of isolation, are potentially pathogenic at the genomic level, which may lead to fatal infections in vulnerable patients. However, other determinants such as genetic variations in virulence related genes and host-pathogen interactions most likely influence A. fumigatus pathogenicity and further studies should be performed.\n\nImportanceAspergillus spp. infections are among the most clinically relevant fungal infections also presenting treatment difficulties due to increasing antifungal resistance. The lack of key virulence factors and a broad genomic diversity complicates the development of targeted diagnosis and novel treatment strategies. A widely spread variability in virulence has been reported for experimental, clinical and environmental isolates. Here we provide supporting evidence that members of this species are fully capable of establishing an infection in immunosuppressed hosts according to their virulence content at the genomic level. Due to the possible clinical complications, studies are urgently required linking strains virulent phenotype with the genotype to better understand the virulence activation of this important fungal pathogen.

microbiology

Unusual metabolism and hypervariation in the genome of a Gracilibacteria (BD1-5) from an oil degrading community

The Candidate Phyla Radiation (CPR) comprises a large monophyletic group of bacterial lineages known almost exclusively based on genomes obtained using cultivation-independent methods. Within the CPR, Gracilibacteria (BD1-5) are particularly poorly understood due to undersampling and the inherent fragmented nature of available genomes. Here, we report the first closed, curated genome of a Gracilibacteria from an enrichment experiment inoculated from the Gulf of Mexico and designed to investigate hydrocarbon degradation. The gracilibacterium rose in abundance after the community switched to dominance by Colwellia. Notably, we predict that this gracilibacterium completely lacks glycolysis, the pentose phosphate and Entner-Doudoroff pathways. It appears to acquire pyruvate, acetyl-CoA and oxaloacetate via degradation of externally derived citrate, malate and amino acids and may use compound interconversion and oxidoreductases to generate and recycle reductive power. The initial genome assembly was fragmented in an unusual gene that is hypervariable within a repeat region. Such extreme local variation is rare, but characteristic of genes that confer traits under pressure to diversify within a population. Notably, the four major repeated 9-mer nucleotide sequences all generate a proline-threonine-aspartic acid (PTD) repeat. The genome of an abundant Colwellia psychrerythraea population has a large extracellular protein that also contains the repeated PTD motif. Although we do not know the host for the BD1-5 cell, the high relative abundance of the C. psychrerythraea population and the shared surface protein repeat may indicate an association between these bacteria.\n\nImportanceCPR bacteria are generally predicted to be symbionts due to their extensive biosynthetic deficits. Although monophyletic, they are not monolithic in terms of their lifestyles. The organism described here appears to have evolved an unusual metabolic platform not reliant on glucose or pentose sugars. Its biology appears to be centered around bacterial host-derived compounds and/or cell detritus. Amino acids likely provide building blocks for nucleic acids, peptidoglycan and protein synthesis. We resolved an unusual repeat region that would be invisible without genome curation. The nucleotide sequence is apparently under strong diversifying selection but the amino acid sequence is under stabilizing selection. The amino acid repeat also occurs in a surface protein of a coexisting bacterium, suggesting co-location and possibly interdependence.

microbiology

Whole genome phylogenies reflect long-tailed distributions of recombination rates in many bacterial species

Although homologous recombination is accepted to be common in bacteria, so far it has been challenging to accurately quantify its impact on genome evolution within bacterial species. We here introduce methods that use the statistics of single-nucleotide polymorphism (SNP) splits in the core genome alignment of a set of strains to show that, for many bacterial species, recombination dominates genome evolution. Each genomic locus has been overwritten so many times by recombination that it is impossible to reconstruct the clonal phylogeny and, instead of a consensus phylogeny, the phylogeny typically changes many thousands of times along the core genome alignment. We also show how SNP splits can be used to quantify the relative rates with which different subsets of strains have recombined in the past. We find that virtually every strain has a unique pattern of frequencies with which its lineages have recombined with those of other strains, and that the relative rates with which different subsets of strains share SNPs follow long-tailed distributions. Our findings show that bacterial populations are neither clonal nor freely recombining, but structured such that recombination rates between different lineages vary along a continuum spanning several orders of magnitude, with a unique pattern of rates for each lineage. Thus, rather than reflecting clonal ancestry, whole genome phylogenies reflect these long-tailed distributions of recombination rates.

evolutionary biology

Genome annotation of Poly(lactic acid) degrading Pseudomonas aeruginosa and Sphingobacterium sp.

Pseudomonas aeruginosa and Sphinogobacterium sp. are well known for their ability to decontaminate many environmental pollutants like PAHs, dyes, pesticides and plastics. The present study reports the annotation of genomes from P. aeruginosa and Sphinogobacterium sp. that were isolated from compost, based on their ability to degrade poly(lactic acid), PLA, at mesophillic temperatures (~30{degrees}C). Draft genomes of both the strains were assembled from Illumina reads, annotated and viewed with an aim of gaining insight into the genetic elements involved in degradation of PLA. The draft-assembled genome of strain Sphinogobacterium strain S2 was 5,604,691 bp in length with 435 contigs (maximum length of 434,971 bp) and an average G+C content of 43.5%. The assembled genome of P. aeruginosa strain S3 was 6,631,638 bp long with 303 contigs (maximum contig length of 659,181 bp) and an average G+C content 66.17 %. A total of 5,385 (60% with annotation) and 6,437 (80% with annotation) protein-coding genes were predicted for strains S2 and S3 respectively. Catabolic genes for biodegradation of xenobiotic and aromatic compounds were identified on both draft genomes. Both strains were found to have the genes attributable to the establishment and regulation of biofilm, with more extensive annotation for this in S3. The genome of P. aeruginosa S3 had the complete cascade of genes involved in the transport and utilization of lactate while Sphinogobacterium strain S2 lacked lactate permease, consistent with its inability to grow on lactate. As a whole, our results reveal and predict the genetic elements providing both strains with the ability to degrade PLA at mesophilic temperature.

microbiology

The scaling of genome size and cell size limits maximum rates of photosynthesis with implications for ecological strategies

A central challenge in plant ecology is to define the major axes of plant functional variation with direct consequences for fitness. Central to the three main components of plant fitness (growth, survival, and reproduction) is the rate of metabolic conversion of CO2 into carbon that can be allocated to various structures and functions. Here we (1) argue that a primary constraint on the maximum rate of photosynthesis per unit leaf area is the size and packing density of cells and (2) show that variation in genome size is a strong predictor of cell sizes, packing densities, and the maximum rate of photosynthesis across terrestrial vascular plants. Regardless of the genic content associated with variation in genome size, the simple biophysical constraints of encapsulating the genome define the lower limit of cell size and the upper limit of cell packing densities, as well as the range of possible cell sizes and densities. Genome size, therefore, acts as a first-order constraint on carbon gain and is predicted to define the upper limits of allocation to growth, reproduction, and defense. The strong effects of genome size on metabolism, therefore, have broad implications for plant biogeography and for other theories of plant ecology, and suggest that selection on metabolism may have a role in genome size evolution.

ecology

Developmentally Regulated Genome Editing in Terminally Differentiated N2-Fixing Heterocysts of Anabaena cylindrica ATCC 29414

Some vegetative cells of Anabaena cylindrica are programed to differentiate semi-regularly spaced, single heterocysts along filaments. Since heterocysts are terminally differentiated non-dividing cells, with the sole known function for solar-powered N2-fixation, is it necessary for a heterocyst to retain the entire genome ({approx}7.1 Mbp) from its progenitor vegetative cell? By sequencing the heterocyst genome, we discovered and confirmed that at least six DNA elements ({approx}0.12 Mbp) are deleted during heterocyst development. The six-element deletions led to the restoration of five genes (nifH1, nifD, hupL, primase P4 and a hypothetical protein gene) that were interrupted in vegetative cells. The deleted elements contained 172 genes present in the genome of vegetative cells. By sequence alignments of intact nif genes (nifH, nifD and hupL) from N2-fixing cyanobacteria (multicellular and unicellular) as well as other N2-fixing bacteria (non-cyanobacteria), we found that interrupted nif genes all contain the conserved core sequences that may be required for phage DNA insertion. Here, we discuss the nif genes interruption which uniquely occurs in heterocyst-forming cyanobacteria. To our best knowledge, this is first time to sequence the genome of heterocyst, a specially differentiated oxic N2-fixing cell. This research demonstrated that (1) different genomes may occur in distinct cell types in a multicellular bacterium; and (2) genome editing is coupled to cellular differentiation and/or cellular function in a heterocyst-forming cyanobacterium.

microbiology

CReasPy-cloning: a method for simultaneous cloning and engineering of megabase-sized genomes in yeast using the CRISPR-Cas9 system

Over the last decade a new strategy was developed to bypass the difficulties to genetically engineer some microbial species by transferring (or \"cloning\") their genome into another organism that is amenable to efficient genetic modifications and therefore acts as a living workbench. As such, the yeast Saccharomyces cerevisiae has been used to clone and engineer genomes from viruses, bacteria and algae. The cloning step requires the insertion of yeast genetic elements within the genome of interest, in order to drive its replication and maintenance as an artificial chromosome in the host cell. Current methods used to introduce these genetic elements are still unsatisfactory, due either to their random nature (transposon) or the requirement for unique restriction sites at specific positions (TAR cloning). Here we describe the CReasPy-Cloning, a new method that combines both the ability of Cas9 to cleave DNA at a user-specified locus and the yeasts highly efficient homologous recombination to simultaneously clone and engineer a bacterial chromosome in yeast. Using the 0.816 Mbp genome of Mycoplasma pneumoniae as a proof of concept, we demonstrate that our method can be used to introduce the yeast genetic element at any location in the bacterial chromosome while simultaneously deleting various genes or group of genes. We also show that CReasPy-cloning can be used to edit up to three independent genomic loci at the same time with an efficiency high enough to warrant the screening of a small (<50) number of clones, allowing for significantly shortened genome engineering cycle times.

bioengineering

Repetitive DNA content in the maize genome is uncoupled from population stratification at SNP loci

MotivationRepetitive DNA is a major component of plant genomes and is thought to be a driver of evolutionary novelty. Describing variation in repeat content among individuals and between populations is key to elucidating the evolutionary significance of repetitive DNA. However, the cost of producing references genomes has limited large-scale intraspecific comparisons to a handful of model organisms where multiple reference genomes are available.\n\nResultsWe examine repeat content variation in the genomes of 94 elite inbred maize lines using graph-based repeat clustering, a reference-free and rapid assay of repeat content. We examine population structure using genome-wide repeat profiles and demonstrate the stiff-stalk and non-stiff-stalk heterotic populations are homogenous with regard to global repeat content. In contrast and similar to previously reported results, the same individuals show clear differentiation, and aggregate into two populations, when examining population structure using genome-wide SNPs. Additionally, we develop a novel kmer based technique to examine the chromosomal distribution of repeat clusters in silico and show a cluster dependent statistically significant association with gene density.\n\nConclusionOur results indicate that repeat content variation in the heterotic populations of maize has not diverged and is uncoupled from population stratification at SNP loci. We also show that repeat families exhibit divergent patterns with regard to chromosomal distribution, some repeat clusters accumulate in regions of high gene density, whereas others aggregate in regions of low gene density.\n\nAuthors contributionsSRB and AB conceived the study, SRB performed the bioinformatic analysis, SRB wrote the paper with input from AB. email contacts: Simon Renny-Byfield: simon.renny-byfield@corteva.com, Andy Baumgarten: andy.baumgarten@corteva.com

plant biology