bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Microbial contamination in the genome of the domesticated olive

Draft genomes of both the wild and domesticated olive were recently published. While working with these genomes, we identified contamination in the domesticated olive genome representing .06% of basepairs and 1.3% of scaffolds. We used targeted and untargeted approaches to identify the contaminating sequences, the majority of which were Aureobasidium pullulans. We applied the same method to the wild olive genome and did not find evidence of contamination. Although contamination in the domesticated olive genome was not prolific, it could lead to biased results either from functional content encoded in the contaminant sequences, or from using it to separate microbial reads from olive reads in microbiome studies.

genomics

Accurate tracking of the mutational landscape of diploid hybrid genomes reveals genetic background effects

BackgroundGenome evolution promotes diversity within a population via mutations, recombination, and whole-genome duplication. However, quantifying precisely these factors in diploid hybrid genomes is challenging. Here we present an integrated experimental and computational workflow to accurately track the mutational landscape of yeast diploid hybrids (MuLoYDH) in terms of single-nucleotide variants, small insertions/deletions, copy-number variants and loss-of-heterozygosity. ResultsHaploid Saccharomyces parents were combined into diploid hybrids with fully phased genome and controlled levels of heterozygosity. The resulting hybrids represented the ancestral state and were evolved under different laboratory protocols. Variant simulations enabled to efficiently integrate competitive and standard mapping, depending on local levels of heterozygosity and read length. Experimental validations proved high accuracy and resolution of our computational approach. Finally, applying MuLoYDH to four different diploids revealed striking genetic background effects. Homozygous S. cerevisiae showed ~4-fold higher mutation rate compared to S. paradoxus. In contrast, interspecies hybrids exhibited mutation rates similar to intraspecies hybrids despite 10-fold higher heterozygosity. MuLoYDH unveiled that a substantial fraction of the genome (~200 bp per generation) was shaped by loss-of-heterozygosity and this process was strongly inhibited by high levels of heterozygosity. ConclusionsWe report a comprehensive framework for characterizing the mutational spectrum of yeast diploid hybrids with unprecedented resolution, which can be generalised to other genetic systems. Applying MuLoYDH to laboratory-evolved hybrids provided novel quantitative insights into the evolutionary processes that mould yeast genomes.

genomics

Vitamin D: marker, cause or consequence of depression? An exploration using genomics

BackgroundTrials testing the effect of vitamin D or omega-3 fatty acids (n3-PUFA) supplementation on major depressive disorder (MDD) reported conflicting findings. These trials were boosted by epidemiological evidence suggesting an inverse association of circulating 25-hydroxyvitamin D (25-OH-D) and n3-PUFA levels with MDD. Observational associations may emerge from unresolved confounding, shared genetic risk, or direct causal relationships. We explored the nature of these associations exploiting data and statistical tools from genomics. MethodsResults from GWAS on 25-OH-D (N = 79366), n3-PUFA (N = 24925) and MDD (135458 cases, 344901 controls) were applied to individual-level data (>2,000 subjects with measures of genotype, DSM-IV lifetime MDD diagnoses and circulating 25-OH-D and n3-PUFA) and summary-level data analyses. Shared genetic risk between traits was tested by polygenic risk scores (PRS). Two-sample Mendelian Randomization (2SMR) analyses tested the potential bidirectional causality between traits. OutcomeIn individual-level data, PRS were associated with the phenotype of the same trait (PRS 25-OH-D p = 1.4e-20, PRS N3-PUFA p = 9.3e-6, PRS MDD p = 1.4e-4), but not with the other phenotypes, suggesting a lack of shared genetic effects. In summary-level data, 2SMR analyses provided no evidence of a causal role on MDD of 25-OH-D (p = 0.50) or n3-PUFA (p = 0.16), or for a causal role of MDD on 25-OH-D (p = 0.25) or n3-PUFA (p = 0.66). ConclusionsApplying genomics tools indicated that that shared genetic risk or direct causality between 25-OH-D, n3-PUFA and MDD is unlikely: unresolved confounding may explain the associations reported in observational studies. These findings represent a cautionary tale for testing supplementation of these compounds in preventing or treating MDD. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSMeta-analyses of trials testing the effect of vitamin D or omega-3 fatty acids (n3-PUFA) supplementation on major depressive disorder (MDD) reported conflicting findings, including small clinical effect or no effect. These trials were boosted by epidemiological evidence suggesting an inverse association of circulating 25-hydroxyvitamin D (25-OH-D) and n3-PUFA levels with MDD. However, observational associations may emerge from different scenarios, including unresolved confounding, shared genetic risk, or direct causal relationships. Added value of this studyGenomics provides unique opportunities to investigate shared risk and causality between traits applying new statistical tools and results from genome-wide association studies (GWAS). In the present study we examined the nature of the association of 25-OH-D and n3-PUFA with MDD using the latest data and tools from genomics. We found no significant evidence of shared genetic risk or direct causality between vitamin D or n-3 PUFA and MDD; at this stage, unresolved confounding should be considered the most likely explanation for the association reported by observational studies. Implications of all the available evidenceFindings from the present study, in conjunction with previous conflicting evidence from clinical studies, represent a cautionary tale for further research testing the potential therapeutic effect of vitamin D and n3-PUFA supplementation on depression, as the expectations of a direct causal effect of these compounds on mood should be substantially reconsidered. Genomic tools could be efficiently employed to examine the nature of observational associations emerging in epidemiology, providing some indications on the most promising associations to be prioritized in subsequent intervention studies.

genomics

Quality of Whole Genome Sequencing from Blood versus Saliva Derived DNA in Cardiac Patients

Whole-genome sequencing (WGS) is becoming an increasingly important tool for detecting genomic variation. Blood derived DNA is the current standard for WGS for research or clinical purposes. We compared the level of microbial contamination, sequencing coverage, as well as yield and concordance of single-nucleotide polymorphism (SNP) and copy number variant (CNV) calls in WGS from paired blood and saliva samples from 5 pediatric heart disease patients. We found that although saliva samples contained a higher proportion of sequence reads that map to the human oral microbiome, these reads were readily excluded by mapping the reads to the human reference genome. Sequencing coverage was low only in 1 of 5 saliva samples. Over 95% SNPs (including rare SNPs) but <80% CNVs called in blood genomes were detected in paired saliva genomes. These findings suggest that most good quality saliva samples can serve as an alternative to blood samples for detection of sequence variants from WGS in cardiovascular disease patients.

genomics

Traces of past transposable element presence in Brassicaceae genome dark matter

AO_SCPLOWBSTRACTC_SCPLOWTransposable elements (TEs) are mobile, repetitive DNA sequences that make the largest contribution to genome bulk. They thus contribute to the so-called "dark matter of the genome", the part of the genome in which nothing is immediately recognizable as biologically functional. We developed a new method, based on k-mers, to identify degenerate TE sequences. With this new algorithm, we detect up to 10% of the A. thaliana genome as derived from as yet unidentified TEs, bringing the proportion of the genome known to be derived from TEs up to 50%. A significant proportion of these sequences overlapped conserved non-coding sequences identified in crucifers and rosids, and transcription factor binding sites. They are overrepresented in some gene regulation networks, such as the flowering gene network, suggesting a functional role for these sequences that have been conserved for more than 100 million years, since the spread of flowering plants in the Cretaceous.

genomics

A high-resolution, chromosome-assigned Komodo dragon genome reveals adaptations in the cardiovascular, muscular, and chemosensory systems of monitor lizards

Monitor lizards are unique among ectothermic reptiles in that they have a high aerobic capacity and distinctive cardiovascular physiology which resembles that of endothermic mammals. We have sequenced the genome of the Komodo dragon (Varanus komodoensis), the largest extant monitor lizard, and present a high resolution de novo chromosome-assigned genome assembly for V. komodoensis, generated with a hybrid approach of long-range sequencing and single molecule physical mapping. Comparing the genome of V. komodoensis with those of related species showed evidence of positive selection in pathways related to muscle energy metabolism, cardiovascular homeostasis, and thrombosis. We also found species-specific expansions of a chemoreceptor gene family related to pheromone and kairomone sensing in V. komodoensis and several other lizard lineages. Together, these evolutionary signatures of adaptation reveal genetic underpinnings of the unique Komodo sensory, cardiovascular, and muscular systems, and suggest that selective pressure altered thrombosis genes to help Komodo dragons evade the anticoagulant effects of their own saliva. As the only sequenced monitor lizard genome, the Komodo dragon genome is an important resource for understanding the biology of this lineage and of reptiles worldwide.

genomics

Claudin-low-like mouse mammary tumors show distinct transcriptomic patterns uncoupled from genomic drivers

Claudin-low breast cancer is a molecular subtype associated with poor prognosis and without targeted treatment options. The claudin-low subtype is defined by certain biological characteristics, some of which may be clinically actionable, such as high immunogenicity. In mice, the medroxyprogesterone acetate (MPA) and 7,12-dimethylbenzanthracene (DMBA) induced mammary tumor model yields a heterogeneous set of tumors, a subset of which display claudin-low features. Neither the genomic characteristics of MPA/DMBA-induced claudin-low tumors, nor those of human claudin-low breast tumors, have been thoroughly explored. The transcriptomic characteristics and subtypes of MPA/DMBA-induced mouse mammary tumors were determined using gene expression microarrays. Somatic mutations and copy number aberrations in MPA/DMBA-induced tumors were identified from whole exome sequencing data. A publicly available dataset was queried to explore the genomic characteristics of human claudin-low breast cancer and to validate findings in the murine tumors. Half of MPA/DMBA-induced tumors showed a claudin-low-like subtype. All tumors carried mutations in known driver genes. While the specific genes carrying mutations varied between tumors, there was a consistent mutational signature with an overweight of T>A transversions in TG dinucleotides. Most tumors carried copy number aberrations with a potential oncogenic driver effect. Overall, several genomic events were observed recurrently, however none accurately delineated claudin-low-like tumors. Human claudin-low breast cancers carried a distinct set of genomic characteristics, in particular a relatively low burden of mutations and copy number aberrations. The gene expression characteristics of claudin-low-like MPA/DMBA-induced tumors accurately reflected those of human claudin-low tumors, including epithelial-mesenchymal transition phenotype, high level of immune activation and low degree of differentiation. There was an elevated expression of the immunosuppressive genes PTGS2 (encoding COX-2) and CD274 (encoding PD-L1) in human and murine claudin-low tumors. Our findings show that the claudin-low breast cancer subtype is not demarcated by specific genomic aberrations, but carries potentially targetable characteristics warranting further research. Author SummaryBreast cancer is comprised of several distinct disease subtypes with different etiologies, prognoses and therapeutic targets. The claudin-low breast cancer subtype is relatively poorly understood, and no specific treatment exists targeting its unique characteristics. Animal models accurately representing human disease counterparts are vital for developing novel therapeutics, but for the claudin-low breast cancer subtype, no such uniform model exists. Here, we show that exposing mice to the carcinogen DMBA and the hormone MPA causes a diverse range of mammary tumors to grow, and half of these have a gene expression pattern similar to that seen in human claudin-low breast cancer. These tumors have numerous changes in their DNA, with clear differences between each tumor, however no specific DNA aberrations clearly demarcate the claudin-low subtype. We also analyzed human breast cancers and show that human claudin-low tumors have several clear patterns in their DNA aberrations, but no specific features accurately distinguish claudin-low from non-claudin-low breast cancer. Finally, we show that both human and murine claudin-low tumors express high levels of genes associated with suppression of immune response. In sum, we highlight claudin-low breast cancer as a clinically relevant subtype with a complex etiology, and with potential unexploited therapeutic targets.

genomics

An Integrated Framework for Genome Analysis Reveals Numerous Previously Unrecognizable Structural Variants in Leukemia Patients' Samples

While genomic analysis of tumors has stimulated major advances in cancer diagnosis, prognosis and treatment, current methods fail to identify a large fraction of somatic structural variants in tumors. We have applied a combination of whole genome sequencing and optical genome mapping to a number of adult and pediatric leukemia samples, which revealed in each of these samples a large number of structural variants not recognizable by current tools of genomic analyses. We developed computational methods to determine which of those variants likely arose as somatic mutations. The method identified 97% of the structural variants previously reported by karyotype analysis of these samples and revealed an additional fivefold more such somatic rearrangements. The method identified on average tens of previously unrecognizable inversions and duplications and hundreds of previously unrecognizable insertions and deletions. These structural variants recurrently affected a number of leukemia associated genes as well as cancer driver genes not previously associated with leukemia and genes not previously associated with cancer. A number of variants only affected intergenic regions but caused cis-acting alterations in expression of neighboring genes. Analysis of TCGA data indicates that the status of several of the recurrently mutated genes identified in this study significantly affect survival of AML patients. Our results suggest that current genomic analysis methods fail to identify a majority of structural variants in leukemia samples and this lacunae may hamper diagnostic and prognostic efforts.

genomics

Genome-wide association analysis for resistance to infectious pancreatic necrosis virus identifies candidate genes involved in viral replication and immune response in rainbow trout (Oncorhynchus mykiss)

Infectious pancreatic necrosis (IPN) is a viral disease with considerable negative impact on the rainbow trout (Oncorhynchus mykiss) aquaculture industry. The aim of the present work was to detect genomic regions that explain resistance to infectious pancreatic necrosis virus (IPNV) in rainbow trout. A total of 2,278 fish from 58 full-sib families were challenged with IPNV. Of the challenged fish, 768 individuals were genotyped (488 resistant and 280 susceptible), using a 57K single nucleotide polymorphisms (SNPs) panel Axiom(R), Affymetrix(R). A genome-wide association study (GWAS) was performed using the phenotypes time to death (TD) and binary survival (BS), along with the genotypes of the challenged fish using a Bayesian model (Bayes C). Heritabilities for resistance to IPNV estimated using pedigree information, were 0.39 and 0.32 for TD and BS, respectively. Heritabilities for resistance to IPNV estimated using genomic information, were 0.50 and 0.54 for TD and BS, respectively. The Bayesian GWAS detected a SNP located on chromosome 5 explaining 18% of the genetic variance for TD. A SNP located on chromosome 23 was detected explaining 9% of the genetic variance for BS. The proximity of Sentrin-specific protease 5 (SENP5) to a significant SNP makes it a candidate gene for resistance against IPNV. However, the moderate-low proportion of variance explained by the detected marker leads to the conclusion that the incorporation of all genomic information, through genomic selection, would be the most appropriate approach to accelerate genetic progress for the improvement of resistance against IPNV in rainbow trout.

genomics

Phased genome sequence of an interspecific hybrid flowering cherry, Somei-Yoshino (Cerasus x yedoensis)

We report the phased genome sequence of an interspecific hybrid, the flowering cherry Somei-Yoshino (Cerasus x yedoensis). The sequence was determined by single-molecule real-time sequencing technology and assembled using a trio-binning strategy in which allelic variation was resolved to obtain phased sequences. The resultant assembly consisting of two haplotype genomes spanned 690.1 Mb with 4,552 contigs and an N50 length of 1.0 Mb. We predicted 95,076 high-confidence genes, including 94.9% of the core eukaryotic genes. Based on a high-density genetic map, we established a pair of eight pseudomolecule sequences, with highly conserved structures between two genome sequences with 2.4 million sequence variants. A whole genome resequencing analysis of flowering cherry varieties suggested that Somei-Yoshino is derived from a cross between C. spachiana and either C. speciose or its derivative. Transcriptome data for flowering date revealed comprehensive changes in gene expression in floral bud development toward flowering. These genome and transcriptome data are expected to provide insights into the evolution and cultivation of flowering cherry and the molecular mechanism underlying flowering.

genomics

A genomic analysis and transcriptomic atlas of gene expression in Psoroptes ovis reveals feeding- and stage-specific patterns of allergen expression

Psoroptic mange, caused by infestation with the ectoparasitic mite, Psoroptes ovis, is highly contagious, resulting in intense pruritus and represents a major welfare and economic concern for the livestock industry Worldwide. Control relies on injectable endectocides and organophosphate dips, but concerns over residues, environmental contamination, and the development of resistance threaten the sustainability of this approach, highlighting interest in alternative control methods. However, development of vaccines and identification of chemotherapeutic targets is hampered by the lack of P. ovis transcriptomic and genomic resources. Building on the recent publication of the P. ovis draft genome, here we present a genomic analysis and transcriptomic atlas of gene expression in P. ovis revealing feeding- and stage-specific patterns of gene expression, including novel multigene families and allergens. Network-based clustering revealed 14 gene clusters demonstrating either single- or multi-stage specific gene expression patterns, with 3,075 female-specific, 890 male-specific and 112, 217 and 526 transcripts showing larval, protonymph and tritonymph specific-expression, respectively. Detailed analysis of P. ovis allergens revealed stage-specific patterns of allergen gene expression, many of which were also enriched in \"fed\" mites and tritonymphs, highlighting an important feeding-related allergenicity in this developmental stage. Pair-wise analysis of differential expression between life-cycle stages identified patterns of sex-biased gene expression and also identified novel P. ovis multigene families including known allergens and novel genes with high levels of stage-specific expression. The genomic and transcriptomic atlas described here represents a unique resource for the acarid-research community, whilst the OrcAE platform makes this freely available, facilitating further community-led curation of the draft P. ovis genome.

genomics

A draft genome sequence of the miniature parasitoid wasp, Megaphragma amalphitanum

Body size reduction, also known as miniaturization, is an important evolutionary process that affects a number of physiological and phenotypic traits and helps animals to conquer new ecological niches. However, this process is poorly understood at the molecular level. Here, we report genomic and transcriptomic features of arguably the smallest known insect - the parasitoid wasp, Megaphragma amalphitanum (Hymenoptera: Trichogrammatidae). In contrast to expectations, we find that the genome and transcriptome sizes of this parasitoid wasp are comparable to other members of the Chalcidoidea superfamily. Moreover, the gene content of M. amalphitanum compared to other chalcid wasps is remarkably conserved. Among the very rare cases of apparent gene loss is centrosomin, which encodes an important centrosome component; the absence of this protein might be related to the large number of anucleate neurons in M. amalphitanum. Intriguingly, we also observed significant changes in M. amalphitanum transposable element dynamics over time, whereby an initial burst was followed by suppression of activity, possibly due to a recent reinforcement of the genome defense machinery. Thus, while the M. amalphitanum genomic data reveal certain features that may be linked to the unusual biological properties of this organism, miniaturization is not associated with a large decrease in genome complexity.

genomics

The genome of the blind soil-dwelling and ancestrally wingless dipluran Campodea augens, a key reference hexapod for studying the emergence of insect innovations

The dipluran two-pronged bristletail Campodea augens is a blind ancestrally wingless hexapod with the remarkable capacity to regenerate lost body appendages such as its long antennae. As sister group to Insecta (sensu stricto), Diplura are key to understanding the early evolution of hexapods and the origin and evolution of insects. Here we report the 1.2-Gbp draft genome of C. augens and results from comparative genomic analyses with other arthropods. In C. augens we uncovered the largest chemosensory gene repertoire of ionotropic receptors in the animal kingdom, a massive expansion which might compensate for the loss of vision. We found a paucity of photoreceptor genes mirroring at the genomic level the secondary loss of an ancestral external photoreceptor organ. Expansions of detoxification and carbohydrate metabolism gene families might reflect adaptations for foraging behaviour, and duplicated apoptotic genes might underlie its high regenerative potential.\n\nThe C. augens genome represents one of the key references for studying the emergence of genomic innovations in insects, the most diverse animal group, and opens up novel opportunities to study the under-explored biology of diplurans.

genomics

The genomic diversification of clonally propagated grapevines

Vegetatively propagated clones accumulate somatic mutations. The purpose of this study was to better understand the consequences of clonal propagation and involved defining the nature of somatic mutations throughout the genome. Fifteen Zinfandel winegrape clone genomes were sequenced and compared to one another using a highly contiguous genome reference produced from one of the clones, Zinfandel 03.\n\nThough most heterozygous variants were shared, somatic mutations accumulated in individual and subsets of clones. Overall, heterozygous mutations were most frequent in intergenic space and more frequent in introns than exons. A significantly larger percentage of CpG, CHG, and CHH sites in repetitive intergenic space experienced transition mutations than genic and non-repetitive intergenic spaces, likely because of higher levels of methylation in the region and the increased likelihood of methylated cytosines to spontaneously deaminate. Of the minority of mutations that occurred in exons, larger proportions of these were putatively deleterious when they occurred in relatively few clones.\n\nThese data support three major conclusions. First, repetitive intergenic space is a major driver of clone genome diversification. Second, clonal propagation is associated with the accumulation of putatively deleterious mutations. Third, the data suggest selection against deleterious variants in coding regions such that mutations are less frequent in coding than noncoding regions of the genome.

genomics

Magnus Representation of Genome Sequences

We introduce an alignment-free method, the Magnus Representation, to analyze genome sequences. The Magnus Representation captures higher-order information in genome sequences. We combine our approach with the idea of k-mers to define an effectively computable Mean Magnus Vector. We perform phylogenetic analysis on three datasets: mosquito-borne viruses, filoviruses, and bacterial genomes. Our results on ebolaviruses are consistent with previous phylogenetic analyses, and confirm the modern viewpoint that the 2014 West African Ebola outbreak likely originated from Central Africa. Our analysis also confirms the close relationship between Bundibugyo ebolavirus and Tai Forest ebolavirus. For bacterial genomes, our method is able to classify relatively well at the family and genus level, as well as at higher levels such as phylum level. The bacterial genomes are also separated well into Gram-positive and Gram-negative subgroups.

genomics

Genomic and phylogenetic analysis of Salmonella Typhimurium and its monophasic variants responsible for invasive endemic infections in Colombia

Salmonellosis is an endemic human infection, associated with both sporadic cases and outbreaks throughout Colombia. Typhimurium is the most common Colombian serovar of Salmonella enterica, responsible for 32.5% of the Salmonella infections. Whole genome sequencing (WGS) is being used increasingly in Europe and the USA to study the epidemiology of Salmonella, but there has not yet been a WGS-based analysis of Salmonella associated with bloodstream infection in Colombia. Here, we analysed 209 genome sequences of Colombian S. Typhimurium and monophasic S. 4,[5],12:i:-isolates from Colombia from 1999 to 2017. We used a core genome-based maximum likelihood tree to define seven distinct clusters which were predominantly Sequence Type (ST) 19 isolates. We also identified the first ST313 and monophasic ST34 isolates to be reported in Colombia. The history of each cluster was reconstructed with a Bayesian tree to reveal a timeline of evolution. Cluster 7 was closely related to European multidrug-resistant (MDR) DT104. Cluster 4 became the dominant variant of Salmonella in 2016, and resistance to nalidixic acid was associated with a plasmid-encoded qnrB19 gene. Our findings suggest multiple transfers of S. Typhimurium between Europe and Colombia.\n\nAuthor summaryThe large-scale genome sequencing of Salmonella Typhimurium and monophasic Salmonella 4,[5],12:i:-involved bloodstream isolates from Colombia. The two serovars were responsible for about 1/3 of Salmonella infections in Colombia in the past 20 years. To identify the population structure we used Whole Genome Sequencing, performed in silico sequence typing, obtained phylogenetic trees, inferred the evolutionary history, detected the plasmids and prophages, and associated the antibiotic resistance (AMR) genotype with phenotype. Different clusters showed temporal replacement. The Colombian sequence type 313 was distinct from African lineages due to the absence of a key virulence-related gene, bstA. One of the Colombian clusters is likely to belong to the global epidemic of DT104, according to the evolutionary history and the AMR profile. The most common cluster in recent years was resistant to nalidixic acid and carried a plasmid-mediated antibiotic resistant gene qnrB19. Our findings will inform the ongoing efforts to combat Salmonellosis by Colombian public health departments.

genomics

The Genomics of Selfing in Maize (Zea mays ssp. mays): Catching Purging in the Act

In plants, self-fertilization is both an important reproductive strategy and a valuable genetic tool. In theory, selfing increases homozygosity at a rate of 0.50 per generation. Increased homozygosity can uncover recessive deleterious variants and lead to inbreeding depression, unless it is countered by the loss of these variants by genetic purging. Here we investigated the dynamics of purging on genomic scale by testing three predictions. The first was that heterozygous, putatively deleterious SNPs were preferentially lost from the genome during continued selfing. The second was that the loss of deleterious SNPs varied as a function of recombination rate, because recombination increases the efficacy of selection by uncoupling linked variants. Finally, we predicted that genome size (GS) decreases during selfing, due to the purging of deleterious transposable element (TE) insertions. We tested these three predictions by following GS and SNP variants in a series of selfed maize (Zea mays ssp. mays) lines over six generations. In these lines, putatively deleterious alleles were purged, and purging was more pronounced in highly recombining regions. Homozygosity increased more slowly than expected; instead of increasing by 50% each generation, it increased by 35% to 40%. Finally, three lines showed dramatic decreases in GS, losing an average of 398 Mb from their genomes over the short timeframe of our experiment. TEs were the principal component of loss, and GS loss was more likely for lineages that began with more TE and more chromosomal knob repeats. Overall, this study documented remarkable GS loss - as much DNA as three Arabidopsis thaliana genomes, on average - in only a few generations of selfing.

genomics

Identifying TCDD-resistance genes via murine and rat comparative genomics and transcriptomics

The aryl hydrocarbon receptor (AHR) mediates many of the toxic effects of 2,3,7,8-tetrachlorodibenzo-p-dioxin (TCDD). However, the AHR alone is insufficient to explain the widely different outcomes among organisms. Attempts to identify unknown factor(s) have been confounded by genetic variability of model organisms. Here, we evaluated three transgenic mouse lines, each expressing a different rat AHR isoform (rWT, DEL, and INS), as well as C57BL/6 and DBA/2 mice. We supplement these with whole-genome sequencing and transcriptomic analyses of the corresponding rat models: Long-Evans (L-E) and Han/Wistar (H/W) rats. These integrated multi-species genomic and transcriptomic data were used to identify genes associated with TCDD-response phenotypes.\n\nWe identified several genes that show consistent transcriptional changes in both transgenic mice and rats. Hepatic Pxdc1 was significantly repressed by TCDD in C57BL/6, rWT mice, and in L-E rat. Three genes demonstrated different AHRE-1 (full) motif occurrences within their promoter regions: Cxxc5 had fewer occurrences in H/W, as compared with L-E; Sugp1 and Hgfac (in either L-E or H/W respectively). These genes also showed different patterns of mRNA abundance across strains.\n\nThe AHR isoform explains much of the transcriptional variability: up to 50% of genes with altered mRNA abundance following TCDD exposure are associated with a single AHR isoform (30% and 10% unique to DEL and rWT respectively following 500 g/kg TCDD). Genomic and transcriptomic evidence allowed identification of genes potentially involved in phenotypic outcomes: Pxdc1 had differential mRNA abundance by phenotype; Cxxc5 had altered AHR binding sites and differential mRNA abundance.\n\nAuthor SummaryEnvironmental contaminants such as dioxins cause many toxic responses, anything from chloracne (common in humans) to death. These toxic responses are mostly regulated by the Ahr, a ligand-activated transcription factor with roles in drug metabolism and immune responses, however other contributing factors remain unclear. Studies are complicated by the underlying genetic heterogeneity of model organisms. Our team evaluated a number of mouse and rat models, including two strains of mouse, two strains of rat and three transgenic mouse lines which differ only at the Ahr locus, that present widely different sensitivities to the most potent dioxin: 2,3,7,8 tetrachlorodibenzo-p-dioxin (TCDD). We identified a number of changes to gene expression that were associated with different toxic responses. We then contrasted these findings with results from whole-genome sequencing of the H/W and L-E rats and found some key genes, such as Cxxc5 and Mafb, which might contribute to TCDD toxicity. These transcriptomic and genomic datasets will provide a valuable resource for future studies into the mechanisms of dioxin toxicities.

genomics