bioRxiv ScienceSearch

Biology subjects

Diogo Meyer

Publications and source records attributed to Diogo Meyer.

2 recordsLinked to original sources

Mapping bias overestimates reference allele frequencies at the HLA genes in the 1000 Genomes Project phase I data

Next Generation Sequencing (NGS) technologies have become the standard for data generation in studies of population genomics, as the 1000 Genomes Project (1000G). However, these techniques are known to be problematic when applied to highly polymorphic genomic regions, such as the Human Leukocyte Antigen (HLA) genes. Because accurate genotype calls and allele frequency estimations are crucial to population ge-nomics analises, it is important to assess the reliability of NGS data. Here, we evaluate the reliability of genotype calls and allele frequency estimates of the SNPs reported by 1000G (phase I) at five HLA genes (HLA-A, -B, -C, -DRB1, -DQB1). We take advantage of the availability of HLA Sanger sequencing of 930 of the 1,092 1000G samples, and use this as a gold standard to benchmark the 1000G data. We document that 18.6% of SNP genotype calls in HLA genes are incorrect, and that allele frequencies are estimated with an error higher than {+/-}0.1 at approximately 25% of the SNPs in HLA genes. We found a bias towards overestimation of reference allele frequency for the 1000G data, indicating mapping bias is an important cause of error in frequency estimation in this dataset. We provide a list of sites that have poor allele frequency estimates, and discuss the outcomes of including those sites in different kinds of analyses. Since the HLA region is the most polymorphic in the human genome, our results provide insights into the challenges of using of NGS data at other genomic regions of high diversity.\n\nData available in public repositories\n\nhttps://github.com/deboraycb/reliability_hla_1000g

Genomics

Heterogeneity of dN/dS ratios at the classical HLA class I genes over divergence time and across the allelic phylogeny

The classical class I HLA loci of humans show an excess of nonsynonymous with respect to synonymous substitutions at codons of the antigen recognition site (ARS), a hallmark of adaptive evolution. Additionally, high polymporphism, linkage disequilibrium and disease associations suggest that one or more balancing selection regimes have acted upon these genes. However, several questions about these selective regimes remain open. First, it is unclear if stronger evidence for selection on deep timescales is due to changes in the intensity of selection over time or to a lack of power of most methods to detect selection on recent timescales. Another question concerns the functional entities which define the selected phenotype. While most analysis focus on selection acting on individual alleles, it is also plausible that phylogenetically defined groups of alleles (\"lineages\") are targets of selection. To address these questions we analyzed how dN/dS ({omega}) varies with respect to divergence times between alleles and phylogenetic placement (position of branches). We find that{omega} for ARS codons of class I HLA genes increases with divergence time and is higher for inter-lineage branches. Throughout our analyses, we used non-selected codons to control for possible effects of inflation of{omega} associated to intra-specific analysis, and showed that our results are not artifactual. Our findings indicate the importance of considering the timescale effect when analysing{omega} over a wide spectrum of divergences. Finally, our results support the divergent allele advantage model, whereby heterozygotes with more divergent alleles have higher fitness than those carrying similar alleles.

Evolutionary Biology