bioRxiv ScienceSearch

Biology subjects

Mik Black

Publications and source records attributed to Mik Black.

3 recordsLinked to original sources

Copy number variants in the sheep genome detected using multiple approaches

Background.Background. Copy number variants (CNVs) are a type of polymorphism found to underlie phenotypic variation, both in humans and livestock. Most surveys of CNV in livestock have been conducted in the cattle genome, and often utilise only a single approach for the detection of copy number differences. Here we performed a study of CNV in sheep, using multiple methods to identify and characterise copy number changes. Comprehensive information from small pedigrees (trios) was collected using multiple platforms (array CGH, SNP chip and whole genome sequence data), with these data then analysed via multiple approaches to identify and verify CNVs.\n\nResults.In total, 3,488 autosomal CNV regions (CNVRs) were identified from 30 sheep. The average length of the identified CNVRs was 19kb (range of 1kb to 3.6Mb), with shorter CNVRs being more frequent than longer CNVRs. The total length of all CNVRs was 67.6Mbps, which equates to 2.7% of the sheep autosomes. For individuals this value ranged from 0.24 to 0.55%, and the majority of CNVRs were identified in single animals. Rather than being uniformly distributed throughout the genome, CNVRs tended to be clustered. Application of three independent approaches for CNVR detection facilitated a comparison of validation rates. CNVs identified on the Roche-NimbleGen 2.1M CGH array generally had low validation rates, while whole genome sequence data had the highest validation rate.\n\nConclusions.This study represents the first comprehensive survey of the distribution, prevalence and characteristics of CNVR in sheep. Multiple approaches were used to detect CNV regions and it appears that the best method for verifying CNVR on a large scale involves using a combination of detection methodologies. The characteristics of the 3,488 autosomal CNV regions identified in this study are comparable to other CNV regions reported in the literature and provide a valuable addition to the small subset of published sheep CNVs.

Genomics

The distribution and impact of common copy-number variation in the genome of the domesticated apple, Malus x domestica Borkh.

BackgroundCopy number variation (CNV) is a common feature of eukaryotic genomes, and a growing body of evidence suggests that genes affected by CNV are enriched in processes that are associated with environmental responses. Here we use next generation sequence (NGS) data to detect copy-number variable regions (CNVRs) within the Malus x domestica genome, as well as to examine their distribution and impact.\n\nMethodsCNVRs were detected using NGS data derived from 30 accessions of M. x domestica analyzed using the read-depth method, as implemented in the CNVrd2 software. To improve the reliability of our results, we developed a quality control and analysis procedure that involved checking for organelle DNA, not repeat masking, and the determination of CNVR identity using a permutation testing procedure.\n\nResultsOverall, we identified 876 CNVRs, which spanned 3.5% of the apple genome. To verify that detected CNVRs were not artifacts, we analyzed the B-allele-frequencies (BAF) within a single nucleotide polymorphism (SNP) array dataset derived from a screening of 185 individual apple accessions and found the CNVRs were enriched for SNPs having aberrant BAFs (P < 1e-13, Fishers Exact test). Putative CNVRs overlapped 845 gene models and were enriched for resistance (R) gene models (P < 1e-22, Fishers exact test). Of note was a cluster of resistance gene models on chromosome 2 near a region containing multiple major gene loci conferring resistance to apple scab.\n\nConclusionWe present the first analysis and catalogue of CNVRs in the M. x domestica genome. The enrichment of the CNVRs with R gene models and their overlap with gene loci of agricultural significance draw attention to a form of unexplored genetic variation in apple. This research will underpin further investigation of the role that CNV plays within the apple genome.

Genomics

Exploring DNA structures in real-time polymerase kinetics using Pacific Biosciences sequencer data

Pausing of DNA polymerase can indicate the presence of a DNA structure that differs from the canonical double-helix. Here we detail a method to investigate how polymerase pausing in the Pacific Biosciences sequencer reads can be related to DNA structure. The Pacific Biosciences sequencer uses optics to view a polymerase and its interaction with a single DNA molecule in real-time, offering a unique way to detect potential alternative DNA structures. We have developed a new way to examine polymerase kinetics and relate it to the DNA sequence by using a wavelet transform of read information from the sequencer. We use this method to examine how polymerase kinetics are related to nucleotide base composition. We then examine tandem repeat sequences known for their ability to form different DNA structures: (CGG)n and (CG)n repeats which can, respectively, form G-quadruplex DNA and Z-DNA. We find pausing around the (CGG)n repeat that may indicate the presence of G-quadruplexes in some of the sequencer reads. The (CG)n repeat does not appear to cause polymerase pausing, but its kinetics signature nevertheless suggests the possibility that alternative nucleotide conformations may sometimes be present. We discuss the implications of using our method to discover DNA sequences capable of forming alternative structures. The analyses presented here can be reproduced on any Pacific Biosciences kinetics data for any DNA pattern of interest using an R package that we have made publicly available.\n\nAuthor SummaryDNA can be found in various forms that differ from the double-helix first discovered by Watson and Crick in 1953. These alternative DNA structures depend on the DNA sequence, and researchers continue to explore which sequences have the potential to form alternative structures. Here we advance the use of Pacific Biosciences sequencer data to explore potential alternative DNA structures. The Pacific Bio-sciences sequencer provides an unprecedented way to examine the interaction between DNA polymerase and DNA by following a single polymerase in real time as it copies a DNA molecule. The pausing of DNA polymerase is a common method for exploring the DNA sequences that have the potential to form alternative DNA structures, and Pacific Biosciences data has previously been used to measure polymerase pausing at a slipped strand structure. DNA polymerase is known to pause at some of these alternative structures, such as the structure known as the G-quadruplex, a DNA structure that has potentially importing regulatory significance. We examine polymerase kinetics around a G-quadruplex, and find evidence of polymerase pausing in the Pacific Biosciences kinetics. We provide a method, with publicly available code, so that others can examine these polymerase kinetics for any sequence of interest.

Bioinformatics