bioRxiv ScienceSearch

Biology subjects

Kono, T. J. Y.

Publications and source records attributed to Kono, T. J. Y..

5 recordsLinked to original sources

RNAlater and flash freezing storage methods nonrandomly influence observed gene expression in RNAseq experiments

RNA-sequencing is a popular next-generation sequencing technique for assaying genome-wide gene expression profiles. Nonetheless, it is susceptible to biases that are introduced by sample handling prior gene expression measurements. Two of the most common methods for preserving samples in both field-based and laboratory conditions are submersion in RNAlater and flash freezing in liquid nitrogen. Flash freezing in liquid nitrogen can be impractical, particularly for field collections. RNAlater is a solution for stabilizing tissue for longer-term storage as it rapidly permeates tissue to protect cellular RNA. In this study, we assessed genome-wide expression patterns in 30 day old fry collected from the same brood at the same time point that were flash-frozen in liquid nitrogen and stored at -80{degrees}C or submerged and stored in RNAlater at room temperature, simulating conditions of fieldwork. We show that sample storage is a significant factor influencing observed differential gene expression. In particular, genes with elevated GC content exhibit higher observed expression levels in liquid nitrogen flash-freezing relative to RNAlater-storage. Further, genes with higher expression in RNAlater relative to liquid nitrogen experience disproportionate enrichment for functional categories, many of which are involved in RNA processing. This suggests that RNAlater may elicit a physiological response that has the potential to bias biological interpretations of expression studies. The biases introduced to observed gene expression arising from mimicking many field-based studies are substantial and should not be ignored.

genomics

The role of gene flow in rapid and repeated evolution of cave related traits in Mexican tetra, Astyanax mexicanus

Understanding the molecular basis of repeated evolved phenotypes can yield key insights into the evolutionary process. Quantifying the amount of gene flow between populations is especially important in interpreting mechanisms of repeated phenotypic evolution, and genomic analyses have revealed that admixture is more common between diverging lineages than previously thought. In this study, we resequenced and analyzed nearly 50 whole genomes of the Mexican tetra from three blind cave populations, two surface populations, and outgroup samples. We confirmed that cave populations are polyphyletic and two Astyanax mexicanus lineages are present in our dataset. The two lineages likely diverged [~]257k generations ago, which, assuming 1 generation per year, is substantially younger than previous mitochondrial estimates of 5-7mya. Divergence of cave populations from their phylogenetically closest surface population likely occurred between [~]161k - 191k generations ago. The favored demographic model for most population pairs accounts for divergence with secondary contact and heterogeneous gene flow across the genome, and we rigorously identified abundant gene flow between cave and surface fish, between caves, and between separate lineages of cave and surface fish. Therefore, the evolution of cave-related traits occurred more rapidly than previously thought, and trogolomorphic traits are maintained despite substantial gene flow with surface populations. After incorporating these new demographic estimates, our models support that selection may drive the evolution of cave-derived traits, as opposed to the classic hypothesis of disuse and drift. Finally, we show that a key QTL is enriched for genomic regions with very low divergence between caves, suggesting that regions important for cave phenotypes may be transferred between caves via gene flow. In sum, our study shows that shared evolutionary history via gene flow must be considered in studies of independent, repeated trait evolution.

evolutionary biology

Tandem Duplicate Genes in Maize are Abundant and Date to Two Distinct Periods of Time

Tandem duplicate genes are proximally duplicated and as such occur in the same genomic neighborhood. Using the maize B73 and PH207 de novo genome assemblies, we identified thousands of tandem gene duplicates that account for ~10% of the genes. These tandem duplicates have a bimodal distribution of estimated ages corresponding to known periods of genomic instability. Tandem duplicates had a number of associated features that suggest origins in nonhomologous recombination based on smaller size distribution and higher rate of containing LTRs than non-tandem duplicates. Within relatively recent tandem duplicate genes, ~26% appear to be undergoing degeneration or divergence in function from the ancestral copy. Our results show that tandem duplicates are abundant in maize, arose in bursts throughout maize evolutionary history under multiple potential mechanisms, and may provide a substrate for novel phenotypic variation.

genomics

Limited role of differential fractionation in genome content variation and function in maize (Zea mays L.) inbred lines

Maize is a diverse paleotetraploid species with widespread presence/absence variation and copy number variation. One mechanism through which presence/absence variation can arise is differential fractionation. Fractionation refers to the loss of duplicate gene pairs from one of the maize subgenomes during diploidization and differential fractionation refers to non-shared gene loss events between individuals. We investigated the prevalence of presence/absence variation resulting from differential fractionation in the syntenic portion of the genome using two whole genome de novo assemblies of the inbred lines B73 and PH207. Between these two genomes, syntenic genes were highly conserved with less than 1% of syntenic genes being subject to differential fractionation. The few variable syntenic genes that were identified are unlikely to contribute to functional phenotypic variation, as there is a significant depletion of these genes in annotated gene sets. In further comparisons of 60 diverse inbred lines, non-syntenic genes were six times more likely to be variable compared to syntenic genes, suggesting that comparisons among additional genome assemblies are not likely to result in the discovery of large-scale presence/absence variation among syntenic genes.\n\nSIGNIFICANCE STATEMENTThere is a large amount of presence/absence variation for gene content in maize. One mechanism that has been hypothesized to contribute to this variation is differential fractionation between individuals following the maize whole genome duplication event. Using comparative genomics, with sorghum and rice representing the ancestral state, we observed little evidence of differential fractionation among elite inbred lines and the few differentially fractionated genes identified did not appear to confer functional significance.

plant biology

Comparative genomics approaches accurately predict deleterious variants in plants

Recent advances in genome resequencing have led to increased interest in prediction of the functional consequences of genetic variants. Variants at phylogenetically conserved sites are of particular interest, because they are more likely than variants at phylogenetically variable sites to have deleterious effects on fitness and contribute to phenotypic variation. Numerous comparative genomic approaches have been developed to predict deleterious variants, but the approaches are nearly always assessed based on their ability to identify known disease-causing mutations in humans. Determining the accuracy of deleterious variant predictions in nonhuman species is important to understanding evolution, domestication, and potentially to improving crop quality and yield. To examine our ability to predict deleterious variants in plants we generated a curated database of 2,910 Arabidopsis thaliana mutants with known phenotypes. We evaluated seven approaches and found that while all performed well, their relative ranking differed from prior benchmarks in humans. We conclude that deleterious mutations can be reliably predicted in A. thaliana and likely other plant species, but that the relative performance of various approaches does not necessarily translate from one species to another.

bioinformatics