bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Fantastic beasts and how to sequence them: genomic approaches for obscure model organisms.

Application of genomic approaches to \"obscure model organisms\" (OMOs), meaning species with little or no genomic resources, enables increasingly sophisticated studies of genomic basis of evolution, acclimatization and adaptation in real ecological contexts. Here, I highlight sequencing solutions and data handling techniques most suited for genomic analysis of OMOs.\n\nGlossary- Allele Frequency Spectrum, AFS (same as Site Frequency Spectrum, SFS): histogram of the number of segregating variants depending on their frequency in one or more populations.\n- Restriction site-Associated DNA (RAD) sequencing: family of diverse genotyping methods that sequence short fragments of the genome adjacent to recognition site(s) for specific restriction endonuclease(s).\n- Linkage Disequilibrium (LD): in this review, correlation of genotypes at a pair of markers across individuals.\n- LD block: typical distance between markers in the genome across which their genotypes remain correlated.\n- Genome scan: profiling of genotypes along the genome looking for unusual patterns. Often used to look for signatures of natural selection or introgression.\n- \"Denser-than-LD\" genotyping: genotyping of several polymorphic markers per LD block.\n- Highly contiguous reference: genome or transcriptome reference sequence containing the least amount of fragmentation.\n- Phased data: data showing which SNP alleles belong to the same homologous chromosome copy.\n- Cross-tissue gene expression analysis: looking for individual-specific shifts in gene expression detectable across multiple tissues. Such shifts are predominantly genetic in nature.

genomics

Using DNase Hi-C techniques to map global and local three-dimensional genome architecture at high resolution

The folding and three-dimensional (3D) organization of chromatin in the nucleus critically impacts genome function. The past decade has witnessed rapid advances in genomic tools for delineating 3D genome architecture. Among them, chromosome conformation capture (3C)-based methods such as Hi-C are the most widely used techniques for mapping chromatin interactions. However, traditional Hi-C protocols rely on restriction enzymes (REs) to fragment chromatin and are therefore limited in resolution. We recently developed DNase Hi-C for mapping 3D genome organization, which uses DNase I for chromatin fragmentation. DNase Hi-C overcomes RE-related limitations associated with traditional Hi-C methods, leading to improved methodological resolution. Furthermore, combining this method with DNA capture technology provides a high-throughput approach (targeted DNase Hi-C) that allows for mapping fine-scale chromatin architecture at exceptionally high resolution. Hence, targeted DNase Hi-C will be valuable for delineating the physical landscapes of cis-regulatory networks that control gene expression and for characterizing phenotype-associated chromatin 3D signatures. Here, we provide a detailed description of method design and step-by-step working protocols for these two methods.\n\nHighlightsO_LIDNase Hi-C, a method for comprehensive mapping of chromatin contacts on a whole-genome scale, is based on random chromatin fragmentation by DNase I digestion instead of sequence-specific restriction enzyme (RE) digestion.\nC_LIO_LITargeted DNase Hi-C, which combines DNase Hi-C with DNA capture technology, is a high-throughput method for mapping fine-scale chromatin architecture of genomic loci of interest at a resolution comparable to that of genomic annotations of functional elements.\nC_LIO_LIDNase Hi-C and targeted DNase Hi-C provide the first high-throughput way to overcome the RE-digestion-associated resolution limit of 3C-based methods.\nC_LIO_LIStep-by-step whole-genome and targeted DNase Hi-C protocols for mapping global and local 3D genome architecture, respectively, are described.\nC_LI

genomics

UniProt Genomic Mapping for Deciphering Functional Effects of Missense Variants

Understanding the association of genetic variation with its functional consequences in proteins is essential for the interpretation of genomic data and identifying causal variants in diseases. Integration of protein function knowledge with genome annotation can assist in rapidly comprehending genetic variation within complex biological processes. Here, we describe mapping UniProtKB human sequences and positional annotations such as active sites, binding sites, and variants to the human genome (GRCh38) and the release of a public genome track hub for genome browsers. To demonstrate the power of combining protein annotations with genome annotations for functional interpretation of variants, we present specific biological examples in disease-related genes and proteins. Computational comparisons of UniProtKB annotations and protein variants with ClinVar clinically annotated SNP data show that 32% of UniProtKB variants co-locate with 8% of ClinVar SNPs. The majority of co-located UniProtKB disease-associated variants (86%) map to pathogenic ClinVar SNPs. UniProt and ClinVar are collaborating to provide a unified clinical variant annotation for genomic, protein and clinical researchers. The genome track hubs, and related UniProtKB files, are downloadable from the UniProt FTP site and discoverable as public track hubs at the UCSC and Ensembl genome browsers.

genomics

A critical comparison of technologies for a plant genome sequencing project

A high quality genome sequence of your model organism is an essential starting point for many studies. Old clone based methods are slow and expensive, whereas faster, cheaper short read only assemblies can be incomplete and highly fragmented, which minimises their usefulness. The last few years have seen the introduction of many new technologies for genome assembly. These new technologies and new algorithms are typically benchmarked on microbial genomes or, if they scale appropriately, human. However, plant genomes can be much more repetitive and larger than human, and plant biology makes obtaining high quality DNA free from contaminants difficult. Reflecting their challenging nature we observe that plant genome assembly statistics are typically poorer than for vertebrates. Here we compare Illumina short read, PacBio long read, 10x Genomics linked reads, Dovetail Hi-C and BioNano Genomics optical maps, singly and combined, in producing high quality long range genome assemblies of the potato species S. verrucosum. We benchmark the assemblies for completeness and accuracy, as well as DNA, compute requirements and sequencing costs. We expect our results will be helpful to other genome projects, and that these datasets will be used in benchmarking by assembly algorithm developers.

genomics

Resolving the Full Spectrum of Human Genome Variation using Linked-Reads

Large-scale population based analyses coupled with advances in technology have demonstrated that the human genome is more diverse than originally thought. To date, this diversity has largely been uncovered using short read whole genome sequencing. However, standard short-read approaches, used primarily due to accuracy, throughput and costs, fail to give a complete picture of a genome. They struggle to identify large, balanced structural events, cannot access repetitive regions of the genome and fail to resolve the human genome into its two haplotypes. Here we describe an approach that retains long range information while harnessing the advantages of short reads. Starting from only [~]1ng of DNA, we produce barcoded short read libraries. The use of novel informatic approaches allows for the barcoded short reads to be associated with the long molecules of origin producing a novel datatype known as Linked-Reads. This approach allows for simultaneous detection of small and large variants from a single Linked-Read library. We have previously demonstrated the utility of whole genome Linked-Reads (lrWGS) for performing diploid, de novo assembly of individual genomes (Weisenfeld et al. 2017). In this manuscript, we show the advantages of Linked-Reads over standard short read approaches for reference based analysis. We demonstrate the ability of Linked-Reads to reconstruct megabase scale haplotypes and to recover parts of the genome that are typically inaccessible to short reads, including phenotypically important genes such as STRC, SMN1 and SMN2. We demonstrate the ability of both lrWGS and Linked-Read Whole Exome Sequencing (lrWES) to identify complex structural variations, including balanced events, single exon deletions, and single exon duplications. The data presented here show that Linked-Reads provide a scalable approach for comprehensive genome analysis that is not possible using short reads alone.

genomics

Signatures of host specialization and a recent transposable element burst in the dynamic one-speed genome of the fungal barley powdery mildew pathogen

Powdery mildews are biotrophic pathogenic fungi infecting a number of economically important plants. The grass powdery mildew, Blumeria graminis, has become a model organism to study host specialization of obligate biotrophic fungal pathogens. We resolved the large-scale genomic architecture of B. graminis forma specialis hordei (Bgh) to explore the potential influence of its genome organization on the co-evolutionary process with its host plant, barley (Hordeum vulgare). The near-chromosome level assemblies of the Bgh reference isolate DH14 and one of the most diversified isolates, RACE1, enabled a comparative analysis of these haploid genomes, which are highly enriched with transposable elements (TEs). We found largely retained genome synteny and gene repertoires, yet detected copy number variation (CNV) of secretion signal peptide-containing protein-coding genes (SPs) and locally disrupted synteny blocks. Genes coding for sequence-related SPs are often locally clustered, but neither the SP clusters nor TEs are enriched in specific genomic regions. Extended comparative analysis with different host-specific B. graminis formae speciales revealed the existence of a core suite of SPs, but also isolate-specific SP sets as well as congruence of SP CNV and phylogenetic relationship. We further detected evidence for a recent, lineage-specific expansion of TEs in the Bgh genome. The characteristics of the Bgh genome (largely retained synteny, CNV of SP genes, recently proliferated TEs and a lack of compartmentalization) are consistent with a \"one-speed\" genome that differs in its architecture and (co-)evolutionary pattern from the \"two-speed\" genomes reported for several other filamentous phytopathogens.

genomics

Nuclear and mitochondrial genomes of the hybrid fungal plant pathogen Verticillium longisporum display a mosaic structure

Allopolyploidization, genome duplication through interspecific hybridization, is an important evolutionary mechanism that can enable organisms to adapt to environmental changes or stresses. This increased adaptive potential of allopolyploids can be particularly relevant for plant pathogens in their quest for host immune response evasion. Allodiploidization likely caused the shift in host range of the fungal pathogen plant Verticillium longisporum, as V. longisporum mainly infects Brassicaceae plants in contrast to haploid Verticillium spp. In this study, we investigated the allodiploid genome structure of V. longisporum and its evolution in the hybridization aftermath. The nuclear genome of V. longisporum displays a mosaic structure, as numerous contigs consists of sections of both parental origins. V. longisporum encountered extensive genome rearrangements, whereas the contribution of gene conversion is negligible. Thus, the mosaic genome structure mainly resulted from genomic rearrangements between parental chromosome sets. Furthermore, a mosaic structure was also found in the mitochondrial genome, demonstrating its bi-parental inheritance. In conclusion, the nuclear and mitochondrial genomes of V. longisporum parents interacted dynamically in the hybridization aftermath. Conceivably, novel combinations of DNA sequence of different parental origin facilitated genome stability after hybridization and consecutive niche adaptation of V. longisporum.

genomics

Whole genome sequence of Mapuche-Huilliche Native Americans

BackgroundWhole human genome sequencing initiatives provide a compendium of genetic variants that help us understand population history and the basis of genetic diseases. Current data mostly focuses on Old World populations and information on the genomic structure of Native Americans, especially those from the Southern Cone is scant.\n\nResultsHere we present a high-quality complete genome sequence of 11 Mapuche-Huilliche individuals (HUI) from Southern Chile (85% genomic and 98% exonic coverage at > 30X), with 96-97% high confidence calls. We found approximately 3.1x106 single nucleotide variants (SNVs) per individual and identified 403,383 (6.9%) of novel SNVs that are not included in current sequencing databases. Analyses of large-scale genomic events detected 680 copy number variants (CNVs) and 4,514 structural variants (SVs), including 398 and 1,910 novel events, respectively. Global ancestry composition of HUI genomes revealed that the cohort represents a marginally admixed population from the Southern Cone, whose genetic component is derived from early Native American ancestors. In addition, we found that HUI genomes display highly divergent and novel variants with potential functional impact that converge in ontological categories essential in cell metabolic processes.\n\nConclusionsMapuche-Huilliche genomes contain a unique set of small- and large-scale genomic variants in functionally linked genes, which may contribute to susceptibility for the development of common complex diseases or traits in admixed Latinos and Native American populations. Our data represents an ancestral reference panel for population-based studies in Native and admixed Latin American populations.

genomics

Systematic Discovery of Conservation States for Single-Nucleotide Annotation of the Human Genome

Comparative genomics sequence data is an important source of information for interpreting genomes. Genome-wide annotations based on this data have largely focused on univariate scores or binary calls of evolutionary constraint. Here we present a complementary whole genome annotation approach, ConsHMM, which applies a multivariate hidden Markov model to learn de novo different conservation states based on the combinatorial and spatial patterns of which species align to and match a reference genome in a multiple species DNA sequence alignment. We applied ConsHMM to a 100-way vertebrate sequence alignment to annotate the human genome at single nucleotide resolution into 100 different conservation states. These states have distinct enrichments for other genomic information including gene annotations, chromatin states, and repeat families, which were used to characterize their biological significance. Conservation states have greater or complementary predictive information than standard constraint based measures for a variety of genome annotations. Bases in constrained elements have distinct heritability enrichments depending on the conservation state assignment, demonstrating their relevance to analyzing phenotypic associated variation. The conservation states also highlight differences in the conservation patterns of bases prioritized by a number of scores used for variant prioritization. The ConsHMM method and conservation state annotations provide a valuable resource for interpreting genomes and genetic variation.

genomics

Bacillus safensis FO-36b and Bacillus pumilus SAFR-032: A Whole Genome Comparison of Two Spacecraft Assembly Facility Isolates

BackgroundBacillus strains producing highly resistant spores have been isolated from cleanrooms and space craft assembly facilities. Organisms that can survive such conditions merit planetary protection concern and if that resistance can be transferred to other organisms, a health concern too. To further efforts to understand these resistances, the complete genome of Bacillus safensis strain FO-36b, which produces spore resistant to peroxide and radiation was determined. The genome was compared to the complete genome of B. pumilus SAFR-032, as well as draft genomes of B. safensis JPL-MERTA-8-2 and the type strain B. pumilus ATCC7061T. In addition, comparisons were made to 61 draft genomes that have been mostly identified as strains of B. pumilus or B. safensis.\n\nResultsThe FO-36b gene order is essentially the same as that in SAFR-032 and other B. pumilus strains [1]. The annotated genome has 3850 open reading frames and 40 noncoding RNAs and riboswitches. Of these, 307 are not shared by SAFR-032, and 65 are also not shared by either MERTA or ATCC7061T. The FO-36b genome was found to have ten unique reading frames and two phage-like regions, which have homology with the Bacillus bacteriophage SPP1 (NC_004166) and Brevibacillus phage Jimmer1 (NC_029104). Differing remnants of the Jimmer1 phage are found in essentially all safensis/pumilus strains. Seven unique genes are part of these phage elements. Comparison of gyrA sequences from FO-36b, SAFR-032, ATCC7061T, and 61 other draft genomes separate the various strains into three distinct clusters. Two of these are subgroups of B. pumilus while the other houses all the B. safensis strains.\n\nConclusionsIt is not immediately obvious that the presence or absence of any specific gene or combination of genes is responsible for the variations in resistance seen. It is quite possible that distinctions in gene regulation can change the level of expression of key proteins thereby changing the organisms resistance properties without gain or loss of a particular gene. What is clear is that phage elements contribute significantly to genome variability. The larger comparison of multiple strains indicates that many strains named as B. pumilus actually belong to the B. safensis group.

genomics

MinION re-sequencing of Giardia genomes and de novo assembly of a new Giardia isolate

BackgroundGenomes of the parasite Giardia duodenalis are relatively small for eukaryotic genomes, yet there are only six publicly available. Difficulties in assembling the tetraploid G. duodenalis genome from short read sequencing data likely contribute to this lack of genomic information. We sequenced three isolates of G. duodenalis (AWB, BGS, and beaver) on the Oxford Nanopore Technologies MinION whose long reads have the potential to address genomic areas that are problematic for short reads.\n\nResultsUsing a hybrid approach that combines MinION long reads and Illumina short reads to take advantage of the continuity of the long reads and the accuracy of the short reads we generated reference quality genomes for each isolate. The genomes for two of the isolates were evaluated against the available reference genomes for comparison. The third genome for which there is no previous data was then assembled. The long reads were used to find structural variants in each isolate to examine heterozygosity. Consistent with previous findings based on SNPs, Giardia BGS was found to be considerably more heterozygous than the other isolates that are from Assemblage A. We also find an enrichment of variant-specific surface proteins in some of the structural variant regions.\n\nConclusionsOur results show that the MinION can be used to generate reference quality genomes in Giardia and further be used to identify structural variant regions that are an important source of genetic variation not previously examined in these parasites.

genomics

A high-quality, long-read de novo genome assembly to aid conservation of Hawaii’s last remaining crow species

Genome-level data can provide researchers with unprecedented precision to examine the causes and genetic consequences of population declines, and to apply these results to conservation management. Here we present a high-quality, long-read, de novo genome assembly for one of the worlds most endangered bird species, the Alala. As the only remaining native crow species in Hawaii, the Alala survived solely in a captive breeding program from 2002 until 2016, at which point a long-term reintroduction program was initiated. The high-quality genome assembly was generated to lay the foundation for both comparative genomics studies, and the development of population-level genomic tools that will aid conservation and recovery efforts. We illustrate how the quality of this assembly places it amongst the very best avian genomes assembled to date, comparable to intensively studied model systems. We describe the genome architecture in terms of repetitive elements and runs of homozygosity, and we show that compared with more outbred species, the Alala genome is substantially more homozygous. We also provide annotations for a subset of immunity genes that are likely to be important for conservation applications, and we discuss how this genome is currently being used as a roadmap for downstream conservation applications.

genomics

A high-quality grapevine downy mildew genome assembly reveals rapidly evolving and lineage-specific putative host adaptation genes

Downy mildews are obligate biotrophic oomycete pathogens that cause devastating plant diseases on economically important crops. Plasmopara viticola is the causal agent of grapevine downy mildew, a major disease in vineyards worldwide. We sequenced the genome of Pl. viticola with PacBio long reads and obtained a new 92.94 Mb assembly with high continuity (359 scaffolds for a N50 of 706.5 kb) due to a better resolution of repeat regions. This assembly presented a high level of gene completeness, recovering 1,592 genes encoding secreted proteins involved in plant-pathogen interactions. Pl. viticola had a two-speed genome architecture, with secreted protein-encoding genes preferentially located in gene-sparse, repeat-rich regions and evolving rapidly, as indicated by pairwise dN/dS values. We also used short reads to assemble the genome of Plasmopara muralis, a closely related species infecting grape ivy (Parthenocissus tricuspidata). The lineage-specific proteins identified by comparative genomics analysis included a large proportion of RxLR cytoplasmic effectors and, more generally, genes with high dN/dS values. We identified 270 candidate genes under positive selection, including several genes encoding transporters and components of the RNA machinery potentially involved in host specialization. Finally, the Pl. viticola genome assembly generated here will allow the development of robust population genomics approaches for investigating the mechanisms involved in adaptation to biotic and abiotic selective pressures in this species.\n\nDATA AVAILABILITYRaw reads and genome assemblies have been deposited in GenBank (BioProjects PRJNA329579 for Pl. viticola and PRJNA448661 for Pl. muralis). Genome assemblies, gene annotations and analysis files (e.g. orthology relationships, full tables for GO enrichment analyses, pairwise dN/dS values and branch-site tests) have been deposited in Dataverse (Pl. viticola assembly and annotation: doi.org/10.15454/4NYHD6, Pl. muralis assembly and annotation: doi.org/10.15454/Q1QJYK, analysis files: doi.org/10.15454/8NZ8X9). Links to the data and information about the grapevine downy mildew genome project can be found at http://grapevine-downy-mildew-genome.com/.

genomics

SMRT long-read sequencing and Direct Label and Stain optical maps allow the generation of a high-quality genome assembly for the European barn swallow (Hirundo rustica rustica)

BackgroundThe barn swallow (Hirundo rustica) is a migratory bird that has been the focus of a large number of ecological, behavioural and genetic studies. To facilitate further population genetics and genomic studies, here we present a reference genome assembly for the European subspecies (H. r. rustica).\n\nFindingsAs part of the Genome10K (G10K) effort on generating high quality vertebrate genomes, we have assembled a highly contiguous genome assembly using Single Molecule Real-Time (SMRT) DNA sequencing and several Bionano optical map technologies. We compared and integrated optical maps derived both from the Nick, Label, Repair and Stain and from the Direct Label and Stain (DLS) technologies. As proposed by Bionano, the DLS more than doubled the scaffold N50 with respect to the nickase. The dual enzyme hybrid scaffold led to a further marginal increase in scaffold N50 and an overall increase of confidence in the scaffolds. After removal of haplotigs, the final assembly is approximately 1.21 Gbp in size, with a scaffold N50 value of over 25.95 Mbp.\n\nConclusionsThis high-quality genome assembly represents a valuable resource for further studies of population genetics and genomics in the barn swallow, and for studies concerning the evolution of avian genomes. It also represents one of the very first genomes assembled by combining SMRT long-read sequencing with the new Bionano DLS technology for scaffolding. The quality of this assembly demonstrates the potential of this methodology to substantially increase the contiguity of genome assemblies.

genomics

Coverage-versus-Length plots, a simple quality control step for de novo yeast genome sequence assemblies

Illumina sequencing has revolutionized yeast genomics, with prices for commercial draft genome sequencing now below $200. The popular SPAdes assembler makes it simple to generate a de novo genome assembly for any yeast species. However, whereas making genome assemblies has become routine, understanding what they contain is still challenging. Here, we show how graphing the information that SPAdes provides about the length and coverage of each scaffold can be used to investigate the nature of an assembly, and to diagnose possible problems. Scaffolds derived from mitochondrial DNA, ribosomal DNA, and yeast plasmids can be identified by their high coverage. Contaminating data, such as cross-contamination from other samples in a multiplex sequencing run, can be identified by its low coverage. Scaffolds derived from the bacteriophage PhiX174 and Lambda DNAs that are frequently used as molecular standards in Illumina protocols can also be detected. Assemblies of yeast genomes with high heterozygosity, such as interspecies hybrids, often contain two types of scaffold: regions of the genome where the two alleles assembled into two separate scaffolds and each has a coverage level C, and regions where the two alleles co-assembled (collapsed) into a single scaffold that has a coverage level 2C. Visualizing the data with Coverage-versus-Length (CVL) plots, which can be done using Microsoft Excel or Google Sheets, provides a simple method to understand the structure of a genome assembly and detect aberrant scaffolds or contigs. We provide a Python script that allows assemblies to be filtered to remove contaminants identified in CVL plots.\n\n100-word article summaryWe describe a simple new method, Coverage-versus-Length plots, for examining de novo genome sequence assemblies. These plots enable researchers to detect scaffolds that have unusually high or unusually low coverage, which allows contaminants, and scaffolds that come from atypical parts of the organisms DNA complement, to be detected. We show that contaminants are common in yeast genomes sequenced in multiplex Illumina runs. We provide instructions for making plots using Microsoft Excel or Google Sheets, and software for filtering assemblies to remove contaminants. Contaminants can be detected and removed, even without knowing their source.

genomics

When genomes collide: multiple modes of germline misregulation in a dysgenic syndrome of Drosophila virilis

In sexually reproducing species the union of gametes that are not closely related can result in genomic incompatibility. Hybrid dysgenic syndromes represent a form of genomic incompatibility that can arise when transposable element (TE) abundance differs between two parents. When TEs lacking in the female parent are transmitted paternally, a lack of corresponding silencing small RNAs (piRNAs) transmitted through the female germline can lead to TE mobilization in progeny. The epigenetic nature of this phenomenon is demonstrated by the fact that genetically identical females of the reciprocal cross are normal. Here we show that in the hybrid dysgenic syndrome of Drosophila virilis, an excess of paternally inherited TE families leads not only to increased expression of these TEs, but also coincides with derepression of TEs in equal abundance within parents. Moreover, TE derepression is stable as flies age and associated with piRNA biogenesis defects for only some TEs. At the same time, TE activation is associated with a genome wide shift in the distribution of endogenous gene expression and an increase in abundance of off-target genic piRNAs. To identify regions of the maternal genome that most protect against dysgenesis, we performed an F3 backcross analysis. We find that pericentric regions play a dominant role in maternal protection. This F3 backcross approach additionally allowed us to clarify the properties of genic paramutation in D. virilis. Overall, results support a model in which early germline events in dysgenesis establish a chronic, stable state of mis-expression that is maintained through adulthood.\n\nSuch early events in the germline that are mediated by parent-of-origin effects may be important in determining patterns of gene expression in natural populations.\n\nAuthor SummaryTransposable elements (TE) are selfish elements that code for the function of copying themselves. More than half the human genome is comprised of such elements. Studies in the fruit flies Drosophila melanogaster and D. virilis have been important in demonstrating a role for RNA silencing by piwi-interacting RNAs (piRNAs) in protecting the genome against these harmful elements. These small RNAs are capable of recognizing TE mRNAs and mediating their destruction by Argonaute proteins. They are also transmitted by the female germline to offspring in order to maintain a stable genome across generations. When males carrying a particular TE family are crossed with females lacking the element, the mother is unable to provide genome defense via complementary piRNAs that target the element. This leads to excess TE activation in the germline and sterility. This phenomenon is known as hybrid dysgenesis. In this article we characterize the genomic landscape of TE destabilization that occurs in hybrid dysgenesis in D. virilis. Previous studies had demonstrated that multiple TEs mobilized during hybrid dysgenesis. We demonstrate that this mobilization of multiple TEs is associated with activation of additional TEs in the germline. In addition, we find that TE activation leads to the production of off-target genic piRNAs that cause reduced expression of highly expressed genes. Finally, we show that genic off-target effects of piRNA silencing can contribute to parent-of-origin effects on gene expression. Similar phenomena may influence patterns of gene expression in the germline of natural populations.

Genetics

Efficient isolation of specific genomic regions retaining molecular interactions by the iChIP system using recombinant exogenous DNA-binding proteins

BackgroundComprehensive understanding of mechanisms of genome functions requires identification of molecules interacting with genomic regions of interest in vivo. We have developed the insertional chromatin immunoprecipitatin (iChIP) technology to isolate specific genomic regions retaining molecular interactions and identify their associated molecules. iChIP consists of locus-tagging and affinity purification. The recognition sequences of an exogenous DNA-binding protein such as LexA are inserted into a genomic region of interest in the cell to be analyzed. The exogenous DNA-binding protein fused with a tag(s) is expressed in the cell and the target genomic region is purified with antibody against the tag(s). In this study, we developed the iChIP system using recombinant DNA-binding proteins to make iChIP more straightforward.\n\nResultsIn this system, recombinant 3xFNLDD-D (r3xFNLDD-D) consisting of the 3xFLAG-tag, a nuclear localization signal, the DNA-binding domain plus the dimerization domain of the LexA protein, and the Dock-tag is used for isolation of specific genomic regions. 3xFNLDD-D was expressed using a silkworm-baculovirus expression system and purified by affinity purification. iChIP using r3xFNLDD-D could efficiently isolate the single-copy chicken Pax5 (cPax5) locus, in which LexA binding elements were inserted, with negligible contamination of other genomic regions. In addition, we could detect RNA associated with the cPax5 locus using this form of the iChIP system combined with RT-PCR.\n\nConclusionsThe iChIP system using r3xFNLDD-D can isolate specific genomic regions retaining molecular interactions without expression of the exogenous DNA-binding protein in the cell to be analyzed. iChIP using r3xFNLDD-D would be more straightforward and useful for analysis of specific genomic regions to elucidate their functions.

Biochemistry

4C-ker: A method to reproducibly identify genome-wide interactions captured by 4C-Seq experiments

4C-Seq has proven to be a powerful technique to identify genome-wide interactions with a single locus of interest (or \"bait\") that can be important for gene regulation. However, analysis of 4C-Seq data is complicated by the many biases inherent to the technique. An important consideration when dealing with 4C-Seq data is the differences in resolution of signal across the genome that result from differences in 3D distance separation from the bait. This leads to the highest signal in the region immediately surrounding the bait and increasingly lower signals in far-cis and trans. Another important aspect of 4C-Seq experiments is the resolution, which is greatly influenced by the choice of restriction enzyme and the frequency at which it can cut the genome. Thus, it is important that a 4C-Seq analysis method is flexible enough to analyze data generated using different enzymes and to identify interactions across the entire genome. Current methods for 4C-Seq analysis only identify interactions in regions near the bait or in regions located in far-cis and trans, but no method comprehensively analyzes 4C signals of different length scales. In addition, some methods also fail in experiments where chromatin fragments are generated using frequent cutter restriction enzymes. Here, we describe 4C-ker, a Hidden-Markov Model based pipeline that identifies regions throughout the genome that interact with the 4C bait locus. In addition we incorporate methods for the identification of differential interactions in multiple 4C-seq datasets collected from different genotypes or experimental conditions. Adaptive window sizes are used to correct for differences in signal coverage in near-bait regions, far-cis and trans chromosomes. Using several datasets, we demonstrate that 4C-ker outperforms all existing 4C-Seq pipelines in its ability to reproducibly identify interaction domains at all genomic ranges with different resolution enzymes.\n\nAUTHORS SUMMARYCircularized chromosome conformation capture, or 4C-Seq is a technique developed to identify regions of the genome that are in close spatial proximity to a single locus of interest ( bait). This technique is used to detect regulatory interactions between promoters and enhancers and to characterize the nuclear environment of different regions within and across different cell types. So far, existing methods for 4C-Seq data analysis do not comprehensively identify interactions across the entire genome due to biases in the technique that are related to the decrease in 4C signal that results from increased 3D distance from the bait. To compensate for these weaknesses in existing methods we developed 4C-ker, a method that explicitly models these biases to improve the analysis of 4C-Seq to better understand the genome wide interaction profile of an individual locus.

Bioinformatics