bioRxiv ScienceSearch

Biology subjects

Shane McCarthy

Publications and source records attributed to Shane McCarthy.

6 recordsLinked to original sources

Whole genome view of the consequences of a population bottleneck using 2926 genome sequences from Finland and United Kingdom

Isolated populations with enrichment of variants due to recent population bottlenecks provide a powerful resource for identifying disease-associated genetic variants and genes. As a model of an isolate population, we sequenced the genomes of 1463 Finnish individuals as part of the Sequencing Initiative Suomi (SISu) Project. We compared the genomic profiles of the 1463 Finns to a sample of 1463 British individuals that were sequenced in parallel as part of the UK10K Project. Whereas there were no major differences in the allele frequency of common variants, a significant depletion of variants in the rare frequency spectrum was observed in Finns when comparing the two populations. On the other hand, we observed >2.1 million variants that were twice as frequent among Finns compared to Britons and 800,000 variants that were more than 10 times more frequent in Finns. Furthermore, in Finns we observed a relative proportional enrichment of variants in the minor allele frequency range between 2 - 5% (p < 2.2x10-16). When stratified by their functional annotations, loss-of-function (LoF) variants showed the highest proportional enrichment in Finns (p = 0.0291). In the noncoding part of the genome, variants in conserved regions (p = 0.002) and promoters (p = 0.01) were also significantly enriched in the Finnish samples. These functional categories represent the highest a priori power for downstream association studies of rare variants using population isolates.

Genetics

Using reference-free compressed data structures to analyse sequencing reads from thousands of human genomes

We are rapidly approaching the point where we have sequenced millions of human genomes. There is a pressing need for new data structures to store raw sequencing data and efficient algorithms for population scale analysis. Current reference based data formats do not fully exploit the redundancy in population sequencing nor take advantage of shared genetic variation. In recent years, the Burrows-Wheeler transform (BWT) and FM-index have been widely employed as a full text searchable index for read alignment and de novo assembly. We introduce the concept of a population BWT and use it to store and index the sequencing reads of 2,705 samples from the 1000 Genomes Project. A key feature is that as more genomes are added, identical read sequences are increasingly observed and compression becomes more efficient. We assess the support in the 1000 Genomes read data for every base position of two human reference assembly versions, identifying that 3.2 Mbp with population support was lost in the transition from GRCh37 with 13.7 Mbp added to GRCh38. We show that the vast majority of variant alleles can be uniquely described by overlapping 31-mers and show how rapid and accurate SNP and indel genotyping can be carried out across the genomes in the population BWT. We use the population BWT to carry out non-reference queries to search for the presence of all known viral genomes, and discover human T-lymphotropic virus 1 integrations in six samples in a recognised epidemiological distribution.

Genomics

Exploring the genetic architecture of inflammatory bowel disease by whole genome sequencing identifies association at ADCY7

In order to further resolve the genetic architecture of the inflammatory bowel diseases, ulcerative colitis and Crohns disease, we sequenced the whole genomes of 4,280 patients at low coverage, and compared them to 3,652 previously sequenced population controls across 73.5 million variants. To increase power we imputed from these sequences into new and existing GWAS cohorts, and tested for association at ~12 million variants in a total of 16,432 cases and 18,843 controls. We discovered a 0.6% frequency missense variant in ADCY7 that doubles risk of ulcerative colitis, and offers insight into a new aspect of disease biology. Despite good statistical power, we did not identify any other new low-frequency risk variants, and found that such variants as a class explained little heritability. We did detect a burden of very rare, damaging missense variants in known Crohns disease risk genes, suggesting that more comprehensive sequencing studies will continue to improve our understanding of the biology of complex diseases.

Genetics

Reference-based phasing using the Haplotype Reference Consortium panel

Haplotype phasing is a fundamental problem in medical and population genetics. Phasing is generally performed via statistical phasing within a genotyped cohort, an approach that can attain high accuracy in very large cohorts but attains lower accuracy in smaller cohorts. Here, we instead explore the paradigm of reference-based phasing. We introduce a new phasing algorithm, Eagle2, that attains high accuracy across a broad range of cohort sizes by efficiently leveraging information from large external reference panels (such as the Haplotype Reference Consortium, HRC) using a new data structure based on the positional BurrowsWheeler transform. We demonstrate that Eagle2 attains a {approx}20x speedup and {approx}10% increase in accuracy compared to reference-based phasing using SHAPEIT2. On European-ancestry samples, Eagle2 with the HRC panel achieves >2x the accuracy of 1000 Genomes-based phasing. Eagle2 is open source and freely available for HRC-based phasing via the Sanger Imputation Service and the Michigan Imputation Server.

Genetics

A reference panel of 64,976 haplotypes for genotype imputation

We describe a reference panel of 64,976 human haplotypes at 39,235,157 SNPs constructed using whole genome sequence data from 20 studies of predominantly European ancestry. Using this resource leads to accurate genotype imputation at minor allele frequencies as low as 0.1%, a large increase in the number of SNPs tested in association studies and can help to discover and refine causal loci. We describe remote server resources that allow researchers to carry out imputation and phasing consistently and efficiently.

Genetics

Purging of deleterious variants due to drift and founder effect in Italian populations with extended autozygosity

Purging through inbreeding occurs when consanguineous marriages increases the rate at which deleterious alleles are present in a homozygous state. In this study we carried out low-read depth (4-10x) whole-genome sequencing in 568 individuals from three Italian founder populations, and compared it to data from other Italian and European populations from the 1000 Genomes Project. We show extended consanguinity and depletion of homozygous genotypes at potentially detrimental sites in the founder populations compared to outbred populations. However these patterns are not compatible with the hypothesis of consanguinity driving the purging of highly deleterious mutations according to simulations. Therefore we conclude that genetic drift and the founder effect should be responsible for the observed purging of deleterious variants.

Evolutionary Biology