bioRxiv ScienceSearch

Biology subjects

Mikhail Lipatov

Publications and source records attributed to Mikhail Lipatov.

3 recordsLinked to original sources

False Negatives Are a Significant Feature of Next Generation Sequencing Callsets

Short-read, next-generation sequencing (NGS) is now broadly used to identify rare or de novo mutations in population samples and disease cohorts. However, NGS data is known to be error-prone and post-processing pipelines have primarily focused on the removal of spurious mutations or \"false positives\" in downstream genome datasets. Less attention has been paid to characterizing the fraction of missing mutations or \"false negatives\" (FN). We design a phylogeny-aware tool to determine false negatives [PhyloFaN] and describe how read coverage and reference bias affect the FN rate. Using thousand-fold coverage NGS data from both Illumina HiSeq and Complete Genomics platforms derived from the 1000 Genomes Project, we first characterize the false negative rate in human mtDNA genomes. The false negative rate for the publically available callsets is 17-20%, even for extremely high coverage haploid data. We demonstrate that high FN rates are not limited to mtDNA by comparing autosomal data from 28 publically available full genomes to intergenic Sanger sequenced regions for each individual. We examine both low-coverage Illumina and high-coverage Complete Genomics genomes. We show that the FN rate varies between [~]6%-18% and that false-positive rates are considerably lower (<3%). The FN rate is strongly dependent on calling pipeline parameters, as well as read coverage. Our results demonstrate that missing mutations are a significant feature of genomic datasets and imply additional fine-tuning of bioinformatics pipelines is needed. We provide a tool which can be used to quantify the FN rate for haploid genomic experiments, without additional generation of validation data.\n\nData depositionData and software are freely available on the Henn Lab website: https://ecoevo.stonybrook.edu/hennlab/data-software/\n\nSoftwareGITHUB via https://ecoevo.stonybrook.edu/hennlab/data-software/

Bioinformatics

Maximum Likelihood Estimation of Biological Relatedness from Low Coverage Sequencing Data

1The inference of biological relatedness from DNA sequence data has a wide array of applications, such as in the study of human disease, anthropology and ecology. One of the most common analytical frameworks for performing this inference is to genotype individuals for large numbers of independent genomewide markers and use population allele frequencies to infer the probability of identity-by-descent (IBD) given observed genotypes. Current implementations of this class of methods assume genotypes are known without error. However, with the advent of 2nd generation sequencing data there are now an increasing number of situations where the confidence attached to a particular genotype may be poor because of low coverage. Such scenarios may lead to biased estimates of the kinship coefficient,{varepsilon} We describe an approach that utilizes genotype likelihoods rather than a single observed best genotype to estimate{phi} and demonstrate that we can accurately infer relatedness in both simulated and real 2nd generation sequencing data from a wide variety of human populations down to at least the third degree when coverage is as low as 2x for both individuals, while other commonly used methods such as PLINK exhibit large biases in such situations. In addition the method appears to be robust when the assumed population allele frequencies are diverged from the true frequencies for realistic levels of genetic drift. This approach has been implemented in the C++ software lcMLkin.

Genetics

Distance from Sub-Saharan Africa Predicts Mutational Load in Diverse Human Genomes

The Out-of-Africa (OOA) dispersal ~50,000 years ago is characterized by a series of founder events as modern humans expanded into multiple continents. Population genetics theory predicts an increase of mutational load in populations undergoing serial founder effects during range expansions. To test this hypothesis, we have sequenced full genomes and high-coverage exomes from 7 geographically divergent human populations from Namibia, Congo, Algeria, Pakistan, Cambodia, Siberia and Mexico. We find that individual genomes vary modestly in the overall number of predicted deleterious alleles. We show via spatially explicit simulations that the observed distribution of deleterious allele frequencies is consistent with the OOA dispersal, particularly under a model where deleterious mutations are recessive. We conclude that there is a strong signal of purifying selection at conserved genomic positions within Africa, but that many predicted deleterious mutations have evolved as if they were neutral during the expansion out of Africa. Under a model where selection is inversely related to dominance, we show that OOA populations are likely to have a higher mutation load due to increased allele frequencies of nearly neutral variants that are recessive or partially recessive.

Genomics