bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Pervasive population genomic consequences of genome duplication in Arabidopsis arenosa

Ploidy-variable species allow direct inference of the effects of chromosome copy number on fundamental evolutionary processes. While an abundance of theoretical work suggests polyploidy should leave distinct population genomic signatures, empirical data remains sparse. We sequenced [~]300 individuals from 39 populations of Arabidopsis arenosa, a naturally diploid-autotetraploid species. We find the impacts of polyploidy on population genomic processes are subtle yet pervasive, including reduced efficiency on linked and purifying selection as well as rampant gene flow from diploids. Initial masking of deleterious mutations, faster rates of nucleotide substitution, and interploidy introgression all conspire to shape the evolutionary potential of polyploids.

evolutionary biology

Cleaning clinical genomic data: Simple identification and removal of recurrently miscalled variants in single genomes

Identification of sequence variation from short-read sequence data is subject to common-yet-intermittent miscalling that occurs in a sequence intrinsic manner. We identify that recurrent false positive single nucleotide variants are strongly present in databases of human sequence variation and demonstrate how each individual sample generates a unique set of recurrent false positive variants. These recurrent miscalls result from known difficulties aligning short-read sequence data between redundant genomic regions. We could replicate, catalogue and remove three quarters of these recurrent miscalls for any given exome with as little as ten rounds of read resampling, realignment and recalling. The removal of such misleading variants reduces the search space for identification of disease causing variants.\n\nList of Abbreviations

genomics

Measuring error rates in genomic perturbation screens: gold standards for human functional genomics

Technological advancement has opened the door to systematic genetics in mammalian cells. Genome-scale loss-of-function screens can assay fitness defects induced by partial gene knockdown, using RNA interference, or complete gene knockout, using new CRISPR techniques. These screens can reveal the basic blueprint required for cellular proliferation. Moreover, comparing healthy to cancerous tissue can uncover genes that are essential only in the tumor; these genes are targets for the development of specific anticancer therapies. Unfortunately, progress in this field has been hampered by offtarget effects of perturbation reagents and poorly quantified error rates in large-scale screens. To improve the quality of information derived from these screens, and to provide a framework for understanding the capabilities and limitations of CRISPR technology, we derive gold-standard reference sets of essential and nonessential genes, and provide a Bayesian classifier of gene essentiality that outperforms current methods on both RNAi and CRISPR screens. Our results indicate that CRISPR technology is more sensitive than RNAi, and that both techniques have nontrivial false discovery rates that can be mitigated by rigorous analytical methods.

Systems Biology

Rational design and whole-genome predictions of single guide RNAs for efficient CRISPR/Cas9-mediated genome editing in Ciona

The CRISPR/Cas9 system has emerged as an important tool for various genome engineering applications. A current obstacle to high throughput applications of CRISPR/Cas9 is the imprecise prediction of highly active single guide. RNAs (sgRNAs). We previously implemented the CRISPR/Cas9 system to induce tissue-specific mutations in the tunicate Ciona. In the present study, we designed and tested 83 single guide RNA (sgRNA) vectors targeting 23 genes expressed in the cardiopharyngeal progenitors and surrounding tissues of Ciona embryo. Using high-throughput sequencing of mutagenized alleles, we identified guide sequences that correlate with sgRNA mutagenesis activity and used this information for the rational design of all possible sgRNAs targeting the Ciona transcriptome. We also describe a one-step cloning-free protocol for the assembly of sgRNA expression cassettes. These cassettes can be directly electroporated as unpurified PCR products into Ciona embryos for sgRNA expression in vivo, resulting in high frequency of CRISPR/Cas9-mediated mutagenesis in somatic cells of electroporated embryos.\n\nWe found a strong correlation between the frequency of an Ebf loss-of-function phenotype and the mutagenesis efficacies of individual Ebf-targeting sgRNAs tested using this method. We anticipate that our approach can be scaled up to systematically design and deliver highly efficient sgRNAs for the tissue-specific investigation of gene functions in Ciona.

Developmental Biology

The Staphylococcus aureus Two-Component System AgrAC Displays Four Distinct Genomic Arrangements That Delineate Genomic Virulence Factor Signatures

Two-component systems (TCSs) consist of a histidine kinase and a response regulator. Here, we evaluated the conservation of the AgrAC TCS among 149 completely sequenced S. aureus strains. It is composed of four genes: agrBDCA. We found that: i) AgrAC system (agr) was found in all but one of the 149 strains; ii) The agr positive strains were further classified into four agr types based on AgrD protein sequences, iii) the four agr types not only specified the chromosomal arrangement of the agr genes but also the sequence divergence of AgrC histidine kinase protein, which confers signal specificity, iv) the sequence divergence was reflected in distinct structural properties especially in the transmembrane region and second extracellular binding domain, and v) there was a strong correlation between the agr type and the virulence genomic profile of the organism. Taken together, these results demonstrate that bioinformatic analysis of the agr locus leads to a classification system that correlates with the presence of virulence factors and protein structural properties.

microbiology

Revealing missing isoforms encoded in the human genome by integrating genomic, transcriptomic and proteomic data

Biological and biomedical research relies on comprehensive understanding of protein-coding transcripts. However, the total number of human proteins is still unknown due to the prevalence of alternative splicing and is much larger than the number of human genes.\n\nIn this paper, we detected 31,566 novel transcripts with coding potential by filtering our ab initio predictions with 50 RNA-seq datasets from diverse tissues/cell lines. PCR followed by MiSeq sequencing showed that at least 84.1% of these predicted novel splice sites could be validated. In contrast to known transcripts, the expression of these novel transcripts were highly tissue-specific. Based on these novel transcripts, at least 36 novel proteins were detected from shotgun proteomics data of 41 breast samples. We also showed L1 retrotransposons have a more significant impact on the origin of new transcripts/genes than previously thought. Furthermore, we found that alternative splicing is extraordinarily widespread for genes involved in specific biological functions like protein binding, nucleoside binding, neuron projection, membrane organization and cell adhesion. In the end, the total number of human transcripts with protein-coding potential was estimated to be at least 204,950.\n\nAuthor summaryThe identification of all human proteins is an important and open problem. In this report we first develop an ab initio predictor to collect candidate gene models as many as possible. Next, comprehensive sets of RNA-seq data from diverse tissues and cell lines are used to select confident transcripts. Experimental validation of a subset of predictions confirms a high accuracy for the predicted coding transcript set and has added about 30,000 new protein-coding transcripts to the existing corpus of knowledge in this area. This is significant progress given that the existing protein-coding transcript number in public databases is about 60,000. Our newly found transcripts are more tissue specific. Based on our results, we show that L1's high impact on gene origin and genes with high number of transcripts are enriched in specific functions. At last, we estimate that the total number of human protein-coding transcripts is in excess of 200,000.

Genomics

Genome-wide identification and characterisation of HOT regions in the human genome

HOT (high-occupancy target) regions, which are bound by a surprisingly large number of transcription factors, are considered to be among the most intriguing findings of recent years. An improved understanding of the roles that HOT regions play in biology would be afforded by knowing the constellation of factors that constitute these domains and by identifying HOT regions across the spectrum of human cell types. We characterised and validated HOT regions in embryonic stem cells (ESCs) and produced a catalogue of HOT regions in a broad range of human cell types. We found that HOT regions are associated with genes that control and define the developmental processes of the respective cell and tissue types. We also showed evidence of the developmental persistence of HOT regions at primitive enhancers and demonstrate unique signatures of HOT regions that distinguish them from typical enhancers and super-enhancers. Finally, we performed an epigenetic analysis to reveal the dynamic epigenetic regulation of HOT regions upon H1 differentiation. Taken together, our results provide a resource for the functional exploration of HOT regions and extend our understanding of the key roles of HOT regions in development and differentiation.

Genomics

First landscape of binding to chromosomes for a domesticated mariner transposase in the human genome: diversity of genomic targets of SETMAR isoforms in two colorectal cell lines

Setmar is a 3-exons gene coding a SET domain fused to a Hsmar1 transposase. Its different transcripts theoretically encode 8 isoforms with SET moieties differently spliced. In vitro, the largest isoform binds specifically to Hsmar1 DNA ends and with no specificity to DNA when it is associated with hPso4. In colon cell lines, we found they bind specifically to two chromosomal targets depending probably on the isoform, Hsmar1 ends and sites with no conserved motifs. We also discovered that the isoforms profile was different between cell lines and patient tissues, suggesting the isoforms encoded by this gene in healthy cells and their functions are currently not investigated.

genomics

A genome-wide resource for high-throughput genomic tagging of yeast ORFs

Here we describe a C-SWAT library for high-throughput tagging of Saccharomyces cerevisiae ORFs. It consists of 5661 strains with an acceptor module inserted after each ORF, which can be efficiently replaced with tags or regulatory elements. We validate the library with targeted sequencing and demonstrate its use by tagging the yeast proteome with bright fluorescent proteins, determining how sequences downstream of ORFs influence protein expression and localizing previously undetected proteins.

genomics

Animals transcribe at least half of the genome

One central goal of genome biology is to understand how the usage of the genome differs between organisms. Our knowledge of genome composition, needed for downstream inferences, is critically dependent on gene annotations, yet problems associated with gene annotation and assembly errors are usually ignored in comparative genomics. Here we analyze the genomes of 68 species across all animal groups and some single-cell eukaryotes for general trends in genome usage and composition, taking into account problems of gene annotation. We show that, regardless of genome size, essentially all animals have comparable amounts of introns and intergenic sequence, with nearly all deviations dominated by increased intergenic sequence. Genomes of model organisms have ratios much closer to 1:1, suggesting that the majority of published genomes of non-model organisms are underannotated and consequently omit substantial numbers of genes, with likely negative impact on evolutionary interpretations. Finally, our results also indicate that most animals transcribe half or more of their genomes arguing against differences in genome usage between animal groups, and also suggesting that the transcribed portion is more dependent on genome size than previously thought.\n\nAuthors SummaryWithin our anthropocentric genomic framework, many analyses tends to try to define humans, mammals, or vertebrates relative to the so-called \"lower\" animals. This implicitly posits that vertebrates are complex organisms with large genomes and invertebrates are simple organisms with small genomes. This has the problem that genome size is therefore presumed to correlate with complexity and ignores any unknown complexity of vast numbers of invertebrate groups, many with large genomes. Animals vary widely in genome size, by almost three orders of magnitude, but when sequencing new animal genomes preference is given to those with smaller genomes for reasons of cost. In trying to understand how genomes are used in general, there is an added layer of complication from quality of the assembly and annotation. We have examined genome usage across a wide range of animals and have described ways to account for errors of low-quality annotations. We also show that the genomes of invertebrates and vertebrates are not so different, and that when large-genome invertebrates are added, genome size alone appears to be defining the fraction of the genome that is genes.

Genomics

Parallel Altitudinal Clines Reveal Adaptive Evolution Of Genome Size In Zea mays

While the vast majority of genome size variation in plants is due to differences in repetitive sequence, we know little about how selection acts on repeat content in natural populations. Here we investigate parallel changes in intraspecific genome size and repeat content of domesticated maize (Zea mays) landraces and their wild relative teosinte across altitudinal gradients in Mesoamerica and South America. We combine genotyping, low coverage whole-genome sequence data, and flow cytometry to test for evidence of selection on genome size and individual repeat abundance. We find that population structure alone cannot explain the observed variation, implying that clinal patterns of genome size are maintained by natural selection. Our modeling additionally provides evidence of selection on individual heterochromatic knob repeats, likely due to their large individual contribution to genome size. To better understand the phenotypes driving selection on genome size, we conducted a growth chamber experiment using a population of highland teosinte exhibiting extensive variation in genome size. We find weak support for a positive correlation between genome size and cell size, but stronger support for a negative correlation between genome size and the rate of cell production. Reanalyzing published data of cell counts in maize shoot apical meristems, we then identify a negative correlation between cell production rate and flowering time. Together, our data suggest a model in which variation in genome size is driven by natural selection on flowering time across altitudinal clines, connecting intraspecific variation in repetitive sequence to important differences in adaptive phenotypes.\n\nAuthor summaryGenome size in plants can vary by orders of magnitude, but this variation has long been considered to be of little to no functional consequence. Studying three independent adaptations to high altitude in Zea mays, we find that genome size experiences parallel pressures from natural selection, causing a linear reduction in genome size with increasing altitude. Though reductions in repetitive content are responsible for the genome size change, we find that only those individual loci contributing most to the variation in genome size are individually targeted by selection. To identify the phenotype influenced by genome size, we study how variation in genome size within a single teosinte population impacts leaf growth and cell division. We find that genome size variation correlates negatively with the rate of cell division, suggesting that individuals with larger genomes require longer to complete a mitotic cycle. Finally, we reanalyze data from maize inbreds to show that faster cell division is correlated with earlier flowering, connecting observed variation in genome size to an important adaptive phenotype.

evolutionary biology

Resolving the complex Bordetella pertussis genome using barcoded nanopore sequencing

The genome of Bordetella pertussis is complex, with high GC content and many repeats, each longer than 1,000 bp. Short-read DNA sequencing is unable to resolve the structure of the genome; however, long-read sequencing offers the opportunity to produce single-contig B. pertussis assemblies using sequencing reads which are longer than the repetitive sections. We used an R9.4 MinION flow cell and barcoding to sequence five B. pertussis strains in a single sequencing run. We then trialled combinations of the many nanopore-user-community-built long-read analysis tools to establish the current optimal assembly pipeline for B. pertussis genome sequences. Our best long-read-only assemblies were produced by Canu read correction followed by assembly with Flye and polishing with Nanopolish, whilst the best hybrids (using nanopore and Illumina reads together) were produced by Canu correction followed by Unicycler. This pipeline produced closed genome sequences for four strains, revealing inter-strain genomic rearrangement. However, read mapping to the Tohama I reference genome suggests that the remaining strain contains an ultra-long duplicated region (over 100 kbp), which was not resolved by our pipeline. We have therefore demonstrated the ability to resolve the structure of several B. pertussis strains per single barcoded nanopore flow cell, but the genomes with highest complexity (e.g. very large duplicated regions) remain only partially resolved using the standard library preparation and will require an alternative library preparation method. For full strain characterisation, we recommend hybrid assembly of long and short reads together; for comparison of genome arrangement, assembly using long reads alone is sufficient.\n\nDATA SUMMARYO_LIFinal sequence read files (fastq) for all 5 strains have been deposited in the SRA, BioProject PRJNA478201, accession numbers SAMN09500966, SAMN09500967, SAMN09500968, SAMN09500969, SAMN09500970\nC_LIO_LIA full list of accession numbers for Illumina sequence reads is available in Table S1\nC_LIO_LIAssembly tests, basecalled read sets and reference materials are available from figshare: https://figshare.com/projects/Resolving_the_complex_Bordetella_pertussis_genome_using_barcoded_nanopore_sequencing/31313\nC_LIO_LIGenome sequences for B. pertussis strains UK36, UK38, UK39, UK48 and UK76 have been deposited in GenBank; accession numbers: CP031289, CP031112, CP031113, QRAX00000000, CP031114\nC_LIO_LISource code and full commands used are available from Github: https://github.com/nataliering/Resolving-the-complex-Bordetella-pertussis-genome-using-barcoded-nanopore-sequencing\nC_LI\n\nIMPACT STATEMENTOver the past two decades, whole genome sequencing has allowed us to understand microbial pathogenicity and evolution on an unprecedented level. However, repetitive regions, like those found throughout the B. pertussis genome, have confounded our ability to resolve complex genomes using short-read sequencing technologies alone. To produce closed B. pertussis genome sequences it is necessary to use a sequencing technology which can generate reads longer than these problematic genomic regions. Using barcoded nanopore sequencing, we show that multiple B. pertussis genomes can be resolved per flow cell. Use of our assembly pipeline to resolve further B. pertussis genomes will advance understanding of how genome-level differences affect the phenotypes of strains which appear monomorphic at nucleotide-level.\n\nThis work expands the recently emergent theme that even the most complex genomes can be resolved with sufficiently long sequencing reads. Additionally, we utilise a more widely accessible alternative sequencing platform to the Pacific Biosciences platform already used by large research centres such as the CDC. Our optimisation process, moreover, shows that the analysis tools favoured by the sequencing community do not necessarily produce the most accurate assemblies for all organisms; pipeline optimisation may therefore be beneficial in studies of unusually complex genomes.

microbiology

Bacterial genome reduction as a result of short read sequence assembly

High-throughput comparative genomics has changed our view of bacterial evolution and relatedness. Many genomic comparisons, especially those regarding the accessory genome that is variably conserved across strains in a species, are performed using assembled genomes. For completed genomes, an assumption is made that the entire genome was incorporated into the genome assembly, while for draft assemblies, often constructed from short sequence reads, an assumption is made that genome assembly is an approximation of the entire genome. To understand the potential effects of short read assemblies on the estimation of the complete genome, we downloaded all completed bacterial genomes from GenBank, simulated short reads, assembled the simulated short reads and compared the resulting assembly to the completed assembly. Although most simulated assemblies demonstrated little reduction, others were reduced by as much as 25%, which was correlated with the repeat structure of the genome. A comparative analysis of lost coding region sequences demonstrated that up to 48 CDSs or up to ~112,000 bases of coding region sequence, were missing from some draft assemblies compared to their finished counterparts. Although this effect was observed to some extent in 32% of genomes, only minimal effects were observed on pan-genome statistics when using simulated draft genome assemblies. The benefits and limitations of using draft genome assemblies should be fully realized before interpreting data from assembly-based comparative analyses.

genomics

A near complete haplotype-phased genome of the dikaryotic wheat stripe rust fungus Puccinia striiformis f. sp. tritici reveals high inter-haplome diversity

A long-standing biological question is how evolution has shaped the genomic architecture of dikaryotic fungi. To answer this, high quality genomic resources that enable haplotype comparisons are essential. Short-read genome assemblies for dikaryotic fungi are highly fragmented and lack haplotype-specific information due to the high heterozygosity and repeat content of these genomes. Here we present a diploidaware assembly of the wheat stripe rust fungus Puccinia striiformis f. sp. tritici based on long-reads using the FALCON-Unzip assembler. RNA-seq datasets were used to infer high quality gene models and identify virulence genes involved in plant infection referred to as effectors. This represents the most complete Puccinia striiformis f. sp. tritici genome assembly to date (83 Mb, 156 contigs, N50 1.5 Mb) and provides phased haplotype information for over 92% of the genome. Comparisons of the phase blocks revealed high inter-haplotype diversity of over 6%. More than 25% of all genes lack a clear allelic counterpart. When investigating genome features that potentially promote the rapid evolution of virulence, we found that candidate effector genes are spatially associated with conserved genes commonly found in basidiomycetes. Yet candidate effectors that lack an allelic counterpart are more distant from conserved genes than allelic candidate effectors, and are less likely to be evolutionarily conserved within the P. striiformis species complex and Pucciniales. In summary, this haplotype-phased assembly enabled us to discover novel genome features of a dikaryotic plant pathogenic fungus previously hidden in collapsed and fragmented genome assemblies.\n\nImportanceCurrent representations of eukaryotic microbial genomes are haploid, hiding the genomic diversity intrinsic to diploid and polyploid life forms. This hidden diversity contributes to the organisms evolutionary potential and ability to adapt to stress conditions. Yet it is challenging to provide haplotype-specific information at a whole-genome level. Here, we take advantage of long-read DNA sequencing technology and a tailored-assembly algorithm to disentangle the two haploid genomes of a dikaryotic pathogenic wheat rust fungus. The two genomes display high levels of nucleotide and structural variations, which leads to allelic variation and the presence of genes lacking allelic counterparts. Non-allelic candidate effector genes, which likely encode important pathogenicity factors, display distinct genome localization patterns and are less likely to be evolutionary conserved than those which are present as allelic pairs. This genomic diversity may promote rapid host adaptation and/or be related to the age of the sequenced isolate since last meiosis.

microbiology

One species, two genomes: A critical assessment of inter isolate variation and identification of assembly incongruence in Haemonchus contortus

BackgroundNumerous quality issues may compromise genomic datas representation of its underlying organism. In this study, we compared two genomes published by different research groups for the parasitic nematode Haemonchus contortus, corresponding to divergent isolates. We analyzed differences between the genomes, attempting to ascertain which were attributable to legitimate biological differences, and which to technical error in one or both genomes.\n\nResultsWe found discrepancies between the H. contortus genomes in both assembly and annotation. The genomes differed in representation of genes that are highly conserved across eukaryotes, with clear evidence of misassembly underlying conserved genes missing from one genome or the other. Only 45% of genes in one genome were orthologous to genes in the other genome, with one genome exhibiting almost as much orthology to C. elegans as its counterpart H. contortus strain. The two genomes differed substantially in probable causes underlying this unexpectedly low orthology. One genome included many more inparalogues than the other, and more frequently assembled inparalogues together on the same portions of contiguous sequence. It also exhibited cases of better-conserved gene position relative to C. elegans.\n\nConclusionThe discrepancies between the two genomes far exceeded those expected as a consequence of biological differences between the two H. contortus isolates. This implies substantial quality issues in one or both genomes, suggesting that researchers must exercise caution when using genomic data for newly sequenced species.

genomics

Reconstruction of ancestral genomes in presence of gene gain and loss.

Since most dramatic genomic changes are caused by genome rearrangements as well as gene duplications and gain/loss events, it becomes crucial to understand their mechanisms and reconstruct ancestral genomes of the given genomes. This problem was shown to be NP-complete even in the \"simplest\" case of three genomes, thus calling for heuristic rather than exact algorithmic solutions. At the same time, a larger number of input genomes may actually simplify the problem in practice as it was earlier illustrated with MGRA, a state-of-the-art software tool for reconstruction of ancestral genomes of multiple genomes.\n\nOne of the key obstacles for MGRA and other similar tools is presence of breakpoint reuses when the same breakpoint region is broken by several different genome rearrangements in the course of evolution. Furthermore, such tools are often limited to genomes composed of the same genes with each gene present in a single copy in every genome. This limitation makes these tools inapplicable for many biological datasets and degrades the resolution of ancestral reconstructions in diverse datasets.\n\nWe address these deficiencies by extending the MGRA algorithm to genomes with unequal gene contents. The developed next-generation tool MGRA2 can handle gene gain/loss events and shares the ability of MGRA to reconstruct ancestral genomes uniquely in the case of limited breakpoint reuse. Furthermore, MGRA2 employs a number of novel heuristics to cope with higher breakpoint reuse and process datasets inaccessible for MGRA. In practical experiments, MGRA2 shows superior performance for simulated and real genomes as compared to other ancestral genomes reconstruction tools. The MGRA2 tool is distributed as an open-source software and can be downloaded from GitHub repository http://github.com/ablab/mgra/. It is also available in the form of a web-server at http://mgra.cblab.org, which makes it readily accessible for inexperienced users.

Bioinformatics

Inferring synteny between genome assemblies: a systematic evaluation

Identification of synteny between genomes of closely related species is an important aspect of comparative genomics. However, it is unknown to what extent draft assemblies lead to errors in such analysis. To investigate this, we fragmented genome assemblies of model nematodes to various extents and conducted synteny identification and downstream analysis. We first show that synteny between species can be underestimated up to 40% and find disagreements between popular tools that infer synteny blocks. This inconsistency and further demonstration of erroneous gene ontology enrichment tests throws into question the robustness of previous synteny analysis when gold standard genome sequences remain limited. In addition, determining the true evolutionary relationship is compromised by assembly improvement using a reference guided approach with a closely related species. Annotation quality, however, has minimal effect on synteny if the assembled genome is highly contiguous. Our results highlight the need for gold standard genome assemblies for synteny identification and accurate downstream analysis.\n\nAuthor summaryGenome assemblies across all domains of life are currently produced routinely. Initial analysis of any new genome usually includes annotation and comparative genomics. Synteny provides a framework in which conservation of homologous genes and gene order is identified between genomes of different species. The availability of human and mouse genomes paved the way for algorithm development in large-scale synteny mapping, which eventually became an integral part of comparative genomics. Synteny analysis is regularly performed on assembled sequences that are fragmented, neglecting the fact that most methods were developed using complete genomes. Here, we systematically evaluate this interplay by inferring synteny in genome assemblies with different degrees of contiguation. As expected, our investigation reveals that assembly quality can drastically affect synteny analysis, from the initial synteny identification to downstream analysis. Importantly, we found that improving a fragmented assembly using synteny with the genome of a related species can be dangerous, as this a priori assumes a potentially false evolutionary relationship between the species. The results presented here re-emphasize the importance of gold standard genomes to the science community, and should be achieved given the current progress in sequencing technology.

bioinformatics

Ancient genomic variation underlies recent and repeated ecological adaptation

Adaptation in the wild often involves standing genetic variation (SGV), which allows rapid responses to selection on ecological timescales. However, we still know little about how the evolutionary histories and genomic distributions of SGV influence local adaptation in natural populations. Here, we address this knowledge gap using the threespine stickleback fish (Gasterosteus aculeatus) as a model. We extend the popular restriction site-associated DNA sequencing (RAD-seq) method to produce phased haplotypes approaching 700 base pairs (bp) in length at each of over 50,000 loci across the stickleback genome. Parallel adaptation in two geographically isolated freshwater pond populations consistently involved fixation of haplotypes that are identical-by-descent. In these same genomic regions, sequence divergence between marine and freshwater stickleback, as measured by dXY, reaches ten-fold higher than background levels and structures genomic variation into distinct marine and freshwater haplogroups. By combining this dataset with a de novo genome assembly of a related species, the ninespine stickleback (Pungitius pungitius), we find that this habitat-associated divergent variation averages six million years old, nearly twice the genome-wide average. The genomic variation that is involved in recent and rapid local adaptation in stickleback has actually been evolving throughout the 15-million-year history since the two species lineages split. This long history of genomic divergence has maintained large genomic regions of ancient ancestry that include multiple chromosomal inversions and extensive linked variation. These discoveries of ancient genetic variation spread broadly across the genome in stickleback demonstrate how selection on ecological timescales is a result of genome evolution over geological timescales, and vice versa.\n\nIMPACT STATEMENTAdaptation to changing environments requires a source of genetic variation. When environments change quickly, species often rely on variation that is already present - so-called standing genetic variation - because new adaptive mutations are just too rare. The threespine stickleback, a small fish species living throughout the Northern Hemisphere, is well-known for its ability to rapidly adapt to new environments. Populations living in coastal oceans are heavily armored with bony plates and spines that protect them from predators. These marine populations have repeatedly invaded and adapted to freshwater environments, losing much of their armor and changing in shape, size, color, and behavior.\n\nAdaptation to freshwater environments can occur in mere decades and probably involves lots of standing genetic variation. Indeed, one of the clearest examples we have of adaptation from standing genetic variation comes from a gene, eda, that controls the shifts in armor plating. This discovery involved two surprises that continue to shape our understanding of the genetics of adaptation. First, freshwater stickleback from across the Northern Hemisphere share the same version, or allele, of this gene. Second, the marine and freshwater alleles arose millions of years ago, even though the freshwater populations studied arose much more recently. While it has been hypothesized that other genes in the stickleback genome may share these patterns, large-scale surveys of genomic variation have been unable to test this prediction directly.\n\nHere, we use new sequencing technologies to survey DNA sequence variation across the stickleback genome for patterns like those at the eda gene. We find that every region of the genome associated with marine-freshwater genetic differences shares this pattern to some degree. Moreover, many of these regions are as old or older than eda, stretching back over 10 million years in the past and perhaps even predating the species we now call the threespine stickleback. We conclude that natural selection has maintained this variation over geological timescales and that the same alleles we observe in freshwater stickleback today are the same as those under selection in ancient, now-extinct freshwater habitats. Our findings highlight the need to understand evolution on macroevolutionary timescales to understand and predict adaptation happening in the present day.

evolutionary biology