bioRxiv Science⌕ Search

Biology subjects

Pivirotto, A.

Publications and source records attributed to Pivirotto, A..

3 recordsLinked to original sources

A new comparative framework for estimating selection on synonymous substitutions.

Selection on synonymous codon usage is a well known and widespread phenomenon, yet existing models often do not account for it or its effect on synonymous substitution rates. In this article, we develop and expand the capabilities of Multiclass Synonymous Substitution (MSS) models, which account for such selection by partitioning synonymous substitutions into two or more classes and estimating a relative substitution rate for each class, while accounting for important confounders like mutation bias. We identify extensive heterogeneity among relative synonymous substitution rates in an empirical dataset of [~]12,000 gene alignments from twelve Drosophila species. We validate model performance using data simulated under a forward population genetic simulation, demonstrating that MSS models are robust to model misspecification. MSS rates are significantly correlated with other covariates of selection on codon usage (population-level polymorphism data and tRNA abundance data), suggesting that models can detect weak signatures of selection on codon usage. With the MSS model, we can now study selection on synonymous substitutions in diverse taxa, independent of any a priori assumptions about the forces driving that selection.

bioinformatics↗

Allele age estimators designed for whole genome datasets show only a modest decrease in accuracy when applied to whole exome datasets.

Personalized genomics in the healthcare system is becoming increasingly accessible as the costs of sequencing decreases. With the increase in the number of genomes, larger numbers of rare variants are being discovered, leading to important initiatives in identifying the functional impacts in relation to disease phenotypes. One way to characterize these variants is to estimate the time the mutation entered the population. However, allele age estimators such as those implemented in the programs Relate, Genealogical Estimator of Variant Age (GEVA), and Runtc, were developed based on the assumption that datasets include the entire genome. We examined the performance of each of these estimators on simulated exome data under a neutral constant population size model, as well as under population expansion and background selection models. We found that each provides usable estimates of allele age from whole-exome datasets. Relate performs the best amongst all three estimators with Pearson coefficients of 0.83 and 0.73 (with respect to true simulated values, for neutral constant and expansion population model, respectively) with a 12 percent and 20 percent decrease in correlation between whole genome and whole exome estimations. Of the three estimators, Relate is best able to parallelize to yield quick results with little resources, however, Relate is currently only able to scale to thousands of samples making it unable to match the hundreds of thousands of samples being currently released. While more work is needed to expand the capabilities of current methods of estimating allele age, these methods show a modest decrease in performance in the estimation of the age of mutations. Article SummaryIncreasing availability of whole exome sequencing yields large numbers of rare variants that have direct impact on disease phenotypes. Many methods of identifying the functional impact of mutations exist including the estimation of the time a mutation entered a population. Popular methods of estimating this time assume whole genome data in the estimate of the allele age based on haplotypes. We simulated genome and exome data under a constant and expansion population demography model and found that there is a decrease in performance in all three methods on exome data of 15-30% depending on the method. Testing the robustness of the best performing method, Relate, further simulations introducing background selection and varying the sample size were also undertaken with similar results.

genomics↗

Analyses of allele age and impact reveal human beneficial alleles are older than neutral controls

A classic population genetic prediction is that alleles experiencing directional selection should swiftly traverse allele frequency space, leaving detectable reductions in genetic variation in linked regions. However, despite this expectation, identifying clear footprints of beneficial allele passage has proven to be surprisingly challenging. We addressed the basic premise underlying this expectation by estimating the ages of large numbers of beneficial and deleterious alleles in a human population genomic data set. Deleterious alleles were found to be young, on average, given their allele frequency. However, beneficial alleles were older on average than non-coding, non-regulatory alleles of the same frequency. This finding is not consistent with directional selection and instead indicates some type of balancing selection. Among derived beneficial alleles, those fixed in the population show higher local recombination rates than those still segregating, consistent with a model in which new beneficial alleles experience an initial period of balancing selection due to linkage disequilibrium with deleterious recessive alleles. Alleles that ultimately fix following a period of balancing selection will leave a modest soft sweep impact on the local variation, consistent with the overall paucity of species-wide hard sweeps in human genomes. Impact StatementAnalyses of allele age and evolutionary impact reveal that beneficial alleles in a human population are often older than neutral controls, suggesting a large role for balancing selection in adaptation.

evolutionary biology↗