bioRxiv Science⌕ Search

Biology subjects

Temple, S. D.

Publications and source records attributed to Temple, S. D..

5 recordsLinked to original sources

Multiple-testing corrections in selection scans using identity-by-descent segments

Failing to correct for multiple testing in selection scans can lead to false discoveries of recent genetic adaptations. The scanning statistics in selection studies are often too complicated to theoretically derive a genome-wide significance level or empirically validate control of the family-wise error rate (FWER). By modeling the autocorrelation of identity-by-descent (IBD) rates, we propose a computationally efficient method to determine genome-wide significance levels in an IBD-based scan for recent positive selection. In whole genome simulations, we show that our method has approximate control of the FWER and can adapt to the spacing of tests along the genome. We also show that these scans can have more than fifty percent power to reject the null model in hard sweeps with a selection coefficient s >= 0.01 and a sweeping allele frequency between twenty-five and seventy-five percent. A few human genes and gene complexes have statistically significant excesses of IBD segments in thousands of samples of African, European, and South Asian ancestry groups from the Trans-Omics for Precision Medicine project and the United Kingdom Biobank. Among the significant loci, many signals of recent selection are shared across ancestry groups. One shared selection signal at a skeletal cell development gene is extremely strong in African ancestry samples. HighlightsO_LIWe propose a method to address multiple testing when scanning along the genome for excess identity-by-descent rates. C_LIO_LIIn whole genome simulations, we calculate that the family-wise error rates of our method are close to the desired family-wise significance level. C_LIO_LIWe perform six selection scans in two consortium datasets covering different ancestry groups and reference genome builds. C_LIO_LIFor a genomic region on chromosome 16, we report extremely high identity-by-descent rates in African ancestry groups and replication in European and South Asian ancestry groups. C_LI

genetics↗

The mutation rate of SARS-CoV-2 is highly variable between sites and is influenced by sequence context, genomic region, and RNA structure

RNA viruses like SARS-CoV-2 have a high mutation rate, which contributes to their rapid evolution. The rate of mutations depends on the mutation type (e.g., A[->]C, A[->]G, etc.) and can vary between sites in the viral genome. Understanding this variation can shed light on the mutational processes at play, and is crucial for quantitative modeling of viral evolution. Using the millions of available SARS-CoV-2 full-genome sequences, we estimate rates of synonymous mutations for all 12 possible nucleotide mutation types and examine how much these rates vary between sites. We find a surprisingly high level of variability and several striking patterns: the rates of four mutation types suddenly increase at one of two gene boundaries; the rates of most mutation types strongly depend on a sites local sequence context, with up to 56-fold differences between contexts; consistent with a previous study, the rates of some mutation types are lower at sites engaged in RNA secondary structure. A simple log-linear model of these features explains [~]15-60% of the fold-variation of mutation rates between sites, depending on mutation type; more complex models only modestly improve predictive power out of sample. We estimate the fitness effect of each mutation based on the number of times it actually occurs versus the number of times it is expected to occur based on the model. We identify several small regions of the genome where synonymous or noncoding mutations occur much less often than expected, indicative of strong purifying selection on the RNA sequence that is independent of protein sequence. Overall, this work expands our basic understanding of SARS-CoV-2s evolution by characterizing the viruss mutation process at the level of individual sites and uncovering several striking mutational patterns that arise from unknown mechanisms.

bioinformatics↗

Fast simulation of identity-by-descent segments

The worst-case runtime complexity to simulate haplotype segments identical by descent (IBD) is quadratic in sample size. We propose two main techniques to reduce the compute time, both of which are motivated by coalescent and recombination processes. We provide mathematical results that explain why our algorithm should outperform a naive implementation with high probability. In our experiments, we observe average compute times to simulate detectable IBD segments around a locus that scale approximately linearly in sample size and take a couple of seconds for sample sizes that are less than ten thousand diploid individuals. In contrast, we find that existing methods to simulate IBD segments take minutes to hours for sample sizes exceeding a few thousand diploid individuals. When using IBD segments to study recent positive selection around a locus, our efficient simulation algorithm makes feasible statistical inferences, e.g., parametric bootstrapping in analyses of large biobanks, that would be otherwise intractable.

genetics↗

Identity-by-descent in large samples

If two haplotypes share the same alleles for an extended gene tract, these haplotypes are likely to be derived identical-by-descent from a recent common ancestor. Identity-by-descent segment lengths are correlated via unobserved ancestral tree and recombination processes, which commonly presents challenges to the derivation of theoretical results in population genetics. We show that the proportion of detectable identity-by-descent segments around a locus is normally distributed when the sample size and the scaled population size are large. We generalize this central limit theorem to cover flexible demographic scenarios, multi-way identity-by-descent segments, and multivariate identity-by-descent rates. We use efficient simulations to study the distributional behavior of the detectable identity-by-descent rate. One consequence of non-normality in finite samples is that a genome-wide scan looking for excess identity-by-descent rates may be subject to anti-conservative control of family-wise error rates. HighlightsO_LIWe show the asymptotic normality of the detectable identity-by-descent rate, a mean of correlated binary random variables that arises in population genetics studies. C_LIO_LIWe generalize our main central limit theorem to cover scenarios of nonconstant population sizes, multi-way identity-by-descent segments, and identity-by-descent rates of multiple samples from the same population. C_LIO_LIIn enormous simulation studies, we use an efficient algorithm to characterize distributional properties of the detectable identity-by-descent rate. C_LI

genetics↗

Modeling recent positive selection in Americans of European ancestry

Recent positive selection can result in an excess of long identity-by-descent (IBD) haplotype segments. The statistical methods that we propose here address three major objectives in studying selective sweeps: scanning for regions of interest, identifying possible sweeping alleles, and estimating a selection coefficient s. First, we implement a selection scan to locate regions of excess IBD rate. Second, we develop a statistic to rank alleles that are in strong linkage disequilibrium with a putative sweeping allele. We aggregate these scores to estimate the allele frequency of the sweeping allele, even if it is not genotyped. Third, we propose an estimator for the selection coefficient and quantify uncertainty using the parametric bootstrap. Comparing against state-of-the-art methods in extensive simulations, we show that our methods are better at identifying sweeping alleles that are at low frequency and at estimating S when S [≥] 0.015. We apply these methods to study positive selection in European ancestry samples from the TOPMed project. We analyze eight loci where the IBD rate is more than four standard deviations above the population median. The IBD rate at LCT is thirty-five standard deviations above the population median, and our estimates of its selection coefficient imply strong selection within the past two hundred generations. Overall, we present robust and accurate approaches to study very recent adaptive evolution without knowing the identity of the causal allele or using time series data.

genetics↗