bioRxiv ScienceSearch

Biology subjects

Heng Li

Publications and source records attributed to Heng Li.

6 recordsLinked to original sources

Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly

The human reference genome assembly plays a central role in nearly all aspects of todays basic and clinical research. GRCh38 is the first coordinate-changing assembly update since 2009 and reflects the resolution of roughly 1000 issues and encompasses modifications ranging from thousands of single base changes to megabase-scale path reorganizations, gap closures and localization of previously orphaned sequences. We developed a new approach to sequence generation for targeted base updates and used data from new genome mapping technologies and single haplotype resources to identify and resolve larger assembly issues. For the first time, the reference assembly contains sequence-based representations for the centromeres. We also expanded the number of alternate loci to create a reference that provides a more robust representation of human population variation. We demonstrate that the updates render the reference an improved annotation substrate, alter read alignments in unchanged regions and impact variant interpretation at clinically relevant loci. We additionally evaluated a collection of new de novo long-read haploid assemblies and conclude that while the new assemblies compare favorably to the reference with respect to continuity, error rate, and gene completeness, the reference still provides the best representation for complex genomic regions and coding sequences. We assert that the collected updates in GRCh38 make the newer assembly a more robust substrate for comprehensive analyses that will promote our understanding of human biology and advance our efforts to improve health.

Genomics

Exploratory analysis and error modeling of a sequencing technology

Next generation DNA sequencing methods have created an unprecedented leap in sequence data generation, thus novel computational tools and statistical models are required to optimize and assess the resulting data. In this report, we explore underlying causes of error for the Illumina Genome Analyzer (IGA) sequencing technology and attempt to quantify their effects using a human bacterial artificial chromosome sequenced to 60,000 fold coverage. Seven potential error predictors are considered: Phred score, read entropy, tile coordinates, local tile density, base position within read, nucleotide call, and lane. With these parameters, logistic regression and log-linear models are constructed and used to show that each of the potential predictors contributes to error (P<1x10-4). With this additional information, we apply the logistic model and achieve a 3% improvement in both the sensitivity and specificity to detect IGA errors. Further, we demonstrate that these modeling approaches can be used as a feedback loop to inform laboratory methods and identify specific machine or run bias.

Genomics

The contribution of rare variation to prostate cancer heritability

Although genome-wide association studies (GWAS) have found more than a hundred common susceptibility alleles for prostate cancer, the GWAS reported variants jointly explain only 33% of risk to siblings, leaving the majority of the familial risk unexplained. We use targeted sequencing of 63 known GWAS risk regions in 9,237 men from four ancestries (African, Latino, Japanese, and European) to explore the role of low-frequency variation in risk for prostate cancer. We find that the sequenced variants explain significantly more of the variance in the trait than the known GWAS variants, thus showing that part of the missing familial risk lies in poorly tagged causal variants at known risk regions. We report evidence for genetic heterogeneity in SNP effect sizes across different ancestries. We also partition heritability by minor allele frequency (MAF) spectrum using variance components methods, and find that a large fraction of heritability (0.12, s.e. 0.05; 95% CI [0.03, 0.21]) is explained by rare variants (MAF<0.01) in men of African ancestry. We use the heritability attributable to rare variants to estimate the coupling between selection and allelic effects at 0.48 (95% CI of [0.19, 0.78]) under the Eyre-Walker model. These results imply that natural selection has driven down the frequency of many prostate cancer risk alleles over evolutionary history. Overall our results show that a substantial fraction of the risk for prostate cancer in men of African ancestry lies in rare variants at known risk loci and suggests that rare variants make a significant contribution to heritability of common traits.

Genetics

SAM/BAM format v1.5 extensions for de novo assemblies

SummaryThe plain text Sequence Alignment/Map (SAM) file format and its companion binary form (BAM) are a generic alignment format for storing read alignments against reference sequences (and unmapped reads) together with structured meta-data (Li et al., 2009). Driven by the needs of the 1000 Genomes Project which sequenced many individual human genomes, early SAM/BAM usage focused on pairwise alignments of reads to a reference. However, through the CIGAR P operator multiple sequence alignments can also be preserved. Herein we describe clarifications and additions in version 1.5 of the specification to facilitate storing de novo sequence alignments: Padded reference sequences (with gap characters), annotation of reads or regions of the reference, and the option of embedding the reference sequence within the file.\n\nAvailabilityThe latest public release of the specification is at http://samtools.sourceforge.net/SAM1.pdf, with in development drafts at https://github.com/samtools/hts-specs/ under version control.\n\nContactpeter.cock@hutton.ac.uk

Bioinformatics

No evidence that natural selection has been less effective at removing deleterious mutations in Europeans than in West Africans

Non-African populations have experienced major bottlenecks in the time since their split from West Africans, which has led to the hypothesis that natural selection to remove weakly deleterious mutations may have been less effective in non-Africans. To directly test this hypothesis, we measure the per-genome accumulation of deleterious mutations across diverse humans. We fail to detect any significant differences, but find that archaic Denisovans accumulated non-synonymous mutations at a higher rate than modern humans, consistent with the longer separation time of modern and archaic humans. We also revisit the empirical patterns that have been interpreted as evidence for less effective removal of deleterious mutations in non-Africans than in West Africans, and show they are not driven by differences in selection after population separation, but by neutral evolution.

Evolutionary Biology

Ancient human genomes suggest three ancestral populations for present-day Europeans

We sequenced genomes from a [~]7,000 year old early farmer from Stuttgart in Germany, an [~]8,000 year old hunter-gatherer from Luxembourg, and seven [~]8,000 year old hunter-gatherers from southern Sweden. We analyzed these data together with other ancient genomes and 2,345 contemporary humans to show that the great majority of present-day Europeans derive from at least three highly differentiated populations: West European Hunter-Gatherers (WHG), who contributed ancestry to all Europeans but not to Near Easterners; Ancient North Eurasians (ANE), who were most closely related to Upper Paleolithic Siberians and contributed to both Europeans and Near Easterners; and Early European Farmers (EEF), who were mainly of Near Eastern origin but also harbored WHG-related ancestry. We model these populations deep relationships and show that EEF had [~]44% ancestry from a \"Basal Eurasian\" lineage that split prior to the diversification of all other non-African lineages.

Genetics