bioRxiv ScienceSearch

Biology subjects

Giulio Genovese

Publications and source records attributed to Giulio Genovese.

5 recordsLinked to original sources

Mosaic Mutations in Blood DNA Sequence Are Associated with Solid Tumor Cancers

Recent findings in understanding the causal role of blood-detectable somatic protein-truncating DNA variants in leukemia prompt questions about generalizability of such observations for other cancer types. We used exome sequencing to compare 22 different cancer phenotypes from TCGA data (~8,000 samples) with more than 6,000 controls using a case-control study design and demonstrate that mosaic protein truncating variants in these genes are also associated with solid-tumor cancers. We analyzed tumor DNA samples from TCGA and observed that the cancer-associated mosaic variants are absent from the tumors.\n\nThrough analysis of different cancer phenotypes we observe gene-specificity for mosaic mutations. PPM1D in previous reports has been linked to breast and ovarian cancer, which our analysis confirms as a specifically associated to ovarian cancer. Additionally, glioblastoma, melanoma and lung cancers show gene specific burden of the mosaic protein truncating mutations. Taken together, these results extend existing observations broadly and link solid-tumor cancers to somatic blood DNA changes.

Genetics

Ultra-rare disruptive and damaging mutations influence educational attainment in the general population

Ultra-rare inherited and de novo disruptive variants in highly constrained (HC) genes are enriched in neurodevelopmental disorders 1-5. However, their impact on cognition in the general population has not been explored. We hypothesize that disruptive and damaging ultra-rare variants (URVs) in HC genes not only confer risk to neurodevelopmental disorders, but also influence general cognitive abilities measured indirectly by years of education (YOE). We tested this hypothesis in 14,133 individuals with whole exome or genome sequencing data. The presence of one or more URVs was associated with a decrease in YOE (3.1 months less for each additional mutation; P-value=3.3x10-8) and the effect was stronger in HC genes enriched for brain expression (6.5 months less, P-value=3.4x10-5). The effect of these variants was more pronounced than the estimated effects of runs of homozygosity and pathogenic copy number variation 6-9. Our findings suggest that effects of URVs in HC genes are not confined to severe neurodevelopmental disorder, but influence the cognitive spectrum in the general population

Genetics

Leveraging distant relatedness to quantify human mutation and gene conversion rates

The rate at which human genomes mutate is a central biological parameter that has many implications for our ability to understand demographic and evolutionary phenomena. We present a method for inferring mutation and gene conversion rates using the number of sequence differences observed in identical-by-descent (IBD) segments together with a reconstructed model of recent population size history. This approach is robust to, and can quantify, the presence of substantial genotyping error, as validated in coalescent simulations. We applied the method to 498 trio-phased Dutch individuals from the Genome of the Netherlands (GoNL) project, sequenced at an average depth of 13x. We infer a point mutation rate of 1.66 {+/-} 0.04 x 10-8 per base per generation, and a rate of 1.26 {+/-} 0.06 x 10-9 for < 20 bp indels. Our estimated average genome-wide mutation rate is higher than most pedigree-based estimates reported thus far, but lower than estimates obtained using substitution rates across primates. By quantifying how estimates vary as a function of allele frequency, we infer the probability that a site is involved in non-crossover gene conversion as 5.99 {+/-} 0.69 x 10-6, consistent with recent reports. We find that recombination does not have observable mutagenic effects after gene conversion is accounted for, and that local gene conversion rates reflect recombination rates. We detect a strong enrichment for recent deleterious variation among mismatching variants found within IBD regions, and observe summary statistics of local IBD sharing to closely match previously proposed metrics of background selection, but find no significant effects of selection on our estimates of mutation rate. We detect no evidence for strong variation of mutation rates in a number of genomic annotations obtained from several recent studies.

Genetics

The weighting is the hardest part: on the behavior of the likelihood ratio test and score test under weight misspecification in rare variant association studies

Rare variant association studies are at a critical inflexion point with the increasing availability of exome-sequencing data. A popular test of association is the sequence kernel association test (SKAT). Weights are embedded within SKAT to reflect the hypothesized contribution of the variants to the trait variance. Correct weighting is expected to boost power, and yet the correct weights are generally unknown. It is therefore important to assess the effect of weight misspecification in SKAT.\n\nWe evaluated the behavior of the score and likelihood ratio tests (LRT) under weight misspecification. Simulation and empirical results revealed that LRT is generally more robust and more powerful than score test in such a circumstance. For instance, when the simulated betas were larger for rarer than for more common variants, (incorrectly) assigning equal weights reduced the power of the LRT by [~] 5%, while the score tests power dropped by [~] 30%.\n\nTo optimize weighting we proposed a data-driven weighting scheme. With this scheme and LRT we detected significant enrichment of rare case mutations (MAF<5%; P-value=7e-04) of a set of constrained genes in the Swedish schizophre-nia case-control cohort with exome-sequencing data.\n\nThe score test is currently preferred for its computational efficiency and power. Indeed, assuming correct specification, in some circumstances the score test is the most powerful test. However, LRT has the compelling qualities of being generally more powerful and more robust to misspecification. This is an important result given that, arguably, misspecified models are likely to be the rule rather than the exception in weighting-based approaches.

Genetics

Modeling Linkage Disequilibrium Increases Accuracy of Polygenic Risk Scores

Polygenic risk scores have shown great promise in predicting complex disease risk, and will become more accurate as training sample sizes increase. The standard approach for calculating risk scores involves LD-pruning markers and applying a P-value threshold to association statistics, but this discards information and may reduce predictive accuracy. We introduce a new method, LDpred, which infers the posterior mean causal effect size of each marker using a prior on effect sizes and LD information from an external reference panel. Theory and simulations show that LDpred outperforms the pruning/thresholding approach, particularly at large sample sizes. Accordingly, prediction R2 increased from 20.1% to 25.3% in a large schizophrenia data set and from 9.8% to 12.0% in a large multiple sclerosis data set. A similar relative improvement in accuracy was observed for three additional large disease data sets and when predicting in non-European schizophrenia samples. The advantage of LDpred over existing methods will grow as sample sizes increase.

Bioinformatics