bioRxiv ScienceSearch

Biology subjects

Daly, M.

Publications and source records attributed to Daly, M..

13 recordsLinked to original sources

Reduced representation sequencing for symbiotic anthozoans: are reference genomes necessary to eliminate endosymbiont contamination and make robust phylogeographic inference?

Anthozoan cnidarians form the backbone of coral reefs. Their success relies on endosymbiosis with photosynthetic dinoflagellates in the family Symbiodiniaceae. Photosymbionts represent a hurdle for researchers using population genomic techniques to study these highly imperiled and ecologically critical species because sequencing datasets harbor unknown mixtures of anthozoan and photosymbiont loci. Here we use range-wide sampling and a double-digest restriction-site associated DNA sequencing (ddRADseq) of the sea anemone Bartholomea annulata to explore how symbiont loci impact the interpretation of phylogeographic patterns and population genetic parameters. We use the genome of the closely related Exaiptasia diaphana (previously Aiptasia pallida) to create an anthozoan-only dataset from a genomic dataset containing both B. annulata and its symbiodiniacean symbionts and then compare this to the raw, holobiont dataset. For each, we investigate spatial patterns of genetic diversity and use coalescent model-based approaches to estimate demographic history and population parameters. The Florida Straits are the only phylogeographic break we recover for B. annulata, with divergence estimated during the last glacial maximum. Because B. annulata hosts multiple members of Symbiodiniaceae, we hypothesize that, under moderate missing data thresholds, de novo clustering algorithms that identify orthologs across datasets will have difficulty identifying shared non-coding loci from the photosymbionts. We infer that, for anthozoans hosting diverse members of Symbiodinaceae, clustering algorithms act as de facto filters of symbiont loci. Thus, while at least some photosymbiont loci remain, these are swamped by orders of magnitude greater numbers of anthozoan loci and thus represent genetic \"noise,\" rather than contributing genetic signal.

genomics

Genomic signatures of sympatric speciation with historical and contemporary gene flow in a tropical anthozoan

Sympatric diversification is increasingly thought to have played an important role in the evolution of biodiversity around the globe. However, an in situ sympatric origin for co-distributed taxa is difficult to demonstrate empirically because different evolutionary processes can lead to similar biogeographic outcomes-especially in ecosystems with few hard barriers to dispersal that can facilitate allopatric speciation followed by secondary contact (e.g. marine habitats). Here we use a genomic (ddRADseq), model-based approach to delimit a cryptic species complex of tropical sea anemones that are co-distributed on coral reefs throughout the Tropical Western Atlantic. We use coalescent simulations in fastsimcoal2 to test competing diversification scenarios that span the allopatric-sympatric continuum. We recover support that the corkscrew sea anemone Bartholomea annulata (Le Sueur, 1817) is a cryptic species complex, co-distributed throughout its range. Simulation and model selection analyses suggest these lineages arose in the face of historical and contemporary gene flow, supporting a sympatric origin, but an alternative secondary contact model also receives appreciable model support. Leveraging the genome of Exaiptasia pallida we identify five loci under divergent selection between cryptic B. annulata lineages that fall within mRNA transcripts or CDS regions. Our study provides a rare empirical, genomic example of sympatric speciation in a tropical anthozoan-a group that includes reef-building corals. Finally, these data represent the first range-wide molecular study of any tropical sea anemone, underscoring that anemone diversity is under described in the tropics, and highlighting the need for additional systematic studies into these ecologically and economically important species.

evolutionary biology

Signals of polygenic adaptation on height have been overestimated due to uncorrected population structure in genome-wide association studies

Genetic predictions of height differ among human populations and these differences are too large to be explained by genetic drift. This observation has been interpreted as evidence of polygenic adaptation. Differences across populations were detected using SNPs genome-wide significantly associated with height, and many studies also found that the signals grew stronger when large numbers of subsignificant SNPs were analyzed. This has led to excitement about the prospect of analyzing large fractions of the genome to detect subtle signals of selection and claims of polygenic adaptation for multiple traits. Polygenic adaptation studies of height have been based on SNP effect size measurements in the GIANT Consortium meta-analysis. Here we repeat the height analyses in the UK Biobank, a much more homogeneously designed study. Our results show that polygenic adaptation signals based on large numbers of SNPs below genome-wide significance are extremely sensitive to biases due to uncorrected population structure.

evolutionary biology

Enrichment of rare protein truncating variants in amyotrophic lateral sclerosis patients

To discover novel genetic risk factors underlying amyotrophic lateral sclerosis (ALS), we aggregated exomes from 3,864 cases and 7,839 ancestry matched controls. We observed a significant excess of ultra-rare and rare protein-truncating variants (PTV) among ALS cases, which was primarily concentrated in constrained genes; however, a significant enrichment in PTVs does persist in the remaining exome. Through gene level analyses, known ALS genes, SOD1, NEK1, and FUS, were the most strongly associated with disease status. We also observed suggestive statistical evidence for multiple novel genes including DNAJC7, which is a highly constrained gene and a member of the heat shock protein family (HSP40). HSP40 proteins, along with HSP70 proteins, facilitate protein homeostasis, such as folding of newly synthesized polypeptides, and clearance of degraded proteins. When these processes are not regulated, misfolding and accumulation of degraded proteins can occur leading to aberrant protein aggregation, one of the pathological hallmarks of neurodegeneration.

genomics

Genome-wide association study implicates CHRNA2 in cannabis use disorder

Introductory paragraphCannabis is the most frequently used illicit psychoactive substance worldwide1. Life time use has been reported among 35-40% of adults in Denmark2 and the United States3. Cannabis use is increasing in the population4-6 and among users around 9% become dependent7. The genetic risk component is high with heritability estimates of 518-70%9. Here we report the first genome-wide significant risk locus for cannabis use disorder (CUD, P=9.31x10-12) that replicates in an independent population (Preplication=3.27x10-3, Pmetaanalysis=9.09x10-12). The finding is based on a genome-wide association study (GWAS) of 2,387 cases and 48,985 controls followed by replication in 5,501 cases and 301,041 controls. The index SNP (rs56372821) is a strong eQTL for CHRNA2 and analyses of the genetic regulated gene expressions identified significant association of CHRNA2 expression in cerebellum with CUD. This indicates a potential therapeutic use in CUD of compounds with agonistic effect on the neuronal acetylcholine receptor alpha-2 subunit encoded by CHRNA2. At the polygenic level analyses revealed a significant decrease in the risk of CUD with increased load of variants associated with cognitive performance.

genomics

Phenome-wide association studies (PheWAS) across large "real-world data" population cohorts support drug target validation

Phenome-wide association studies (PheWAS), which assess whether a genetic variant is associated with multiple phenotypes across a phenotypic spectrum, have been proposed as a possible aid to drug development through elucidating mechanisms of action, identifying alternative indications, or predicting adverse drug events (ADEs). Here, we evaluate whether PheWAS can inform target validation during drug development. We selected 25 single nucleotide polymorphisms (SNPs) linked through genome-wide association studies (GWAS) to 19 candidate drug targets for common disease therapeutic indications. We independently interrogated these SNPs through PheWAS in four large \"real-world data\" cohorts (23andMe, UK Biobank, FINRISK, CHOP) for association with a total of 1,892 binary endpoints. We then conducted meta-analyses for 145 harmonized disease endpoints in up to 697,815 individuals and joined results with summary statistics from 57 published GWAS. Our analyses replicate 70% of known GWAS associations and identify 10 novel associations with study-wide significance after multiple test correction (P<1.8x10-6; out of 72 novel associations with FDR<0.1). By leveraging directionality and point estimate of the effect sizes, we describe new associations that may predict ADEs, e.g., acne, high cholesterol, gout and gallstones for rs738409 (p.I148M) in PNPLA3; or asthma for rs1990760 (p.T946A) in IFIH1. We further propose how quantitative estimates of genetic safety/efficacy profiles can be used to help prioritize candidate targets for a specific indication. Our results demonstrate PheWAS as a powerful addition to the toolkit for drug discovery.\n\nOne Sentence SummaryMatching genetics with phenotypes in 800,000 individuals predicts efficacy and on-target safety of future drugs.

genetics

Whole Genome Sequencing in Psychiatric Disorders: the WGSPD Consortium

As technology advances, whole genome sequencing (WGS) is likely to supersede other genotyping technologies. The rate of this change depends on its relative cost and utility. Variants identified uniquely through WGS may reveal novel biological pathways underlying complex disorders and provide high-resolution insight into when, where, and in which cell type these pathways are affected. Alternatively, cheaper and less computationally intensive approaches may yield equivalent insights. Understanding the role of rare variants in the noncoding gene-regulating genome, through pilot WGS projects, will be critical to determine which of these two extremes best represents reality. With large cohorts, well-defined risk loci, and a compelling need to understand the underlying biology, psychiatric disorders have a role to play in this preliminary WGS assessment. The WGSPD consortium will integrate data for 18,000 individuals with psychiatric disorders, beginning with autism spectrum disorder, schizophrenia, bipolar disorder, and major depressive disorder, along with over 150,000 controls.

genomics

Quantifying the impact of rare and ultra-rare coding variation across the phenotypic spectrum

There is a limited understanding about the impact of rare protein truncating variants across multiple phenotypes. We explore the impact of this class of variants on 13 quantitative traits and 10 diseases using whole-exome sequencing data from 100,296 individuals. Protein truncating variants in genes intolerant to this class of mutations increased risk of autism, schizophrenia, bipolar disorder, intellectual disability, ADHD. In individuals without these disorders, there was an association with shorter height, lower education, increased hospitalization and reduced age. Gene sets implicated from GWAS did not show a significant protein truncating variants-burden beyond what captured by established Mendelian genes. In conclusion, we provide the most thorough investigation to date of the impact of rare deleterious coding variants on complex traits, suggesting widespread pleiotropic risk.\n\nMain abbreviations

genetics

The iPSYCH2012 case-cohort sample: New directions for unravelling genetic and environmental architectures of severe mental disorders

The iPSYCH consortium has established a large Danish population-based Case-Cohort sample (iPSYCH2012) aimed at unravelling the genetic and environmental architecture of severe mental disorders. The iPSYCH2012 sample is nested within the entire Danish population born 1981-2005 including 1,472,762 persons. This paper introduces the iPSYCH2012 sample and outlines key future research directions. Cases were identified as persons with schizophrenia (N=3,540), autism (N=16,146), ADHD (N=18,726), and affective disorder (N=26,380), of which 1928 had bipolar affective disorder. Controls were randomly sampled individuals (N=30,000). Within the sample of 86,189 individuals, a total of 57,377 individuals had at least one major mental disorder. DNA was extracted from the neonatal dried blood spot samples obtained from the Danish Neonatal Screening Biobank and genotyped using the Illumina PsychChip. Genotyping was successful for 90% of the sample. The assessments of exome sequencing, methylation profiling, metabolome profiling, vitamin-D, inflammatory and neurotrophic factors are in progress. For each individual, the iPSYCH2012 sample also includes longitudinal information on health, prescribed medicine, social and socioeconomic information and analogous information among relatives. To the best of our knowledge, the iPSYCH2012 sample is the largest and most comprehensive data source for the combined study of genetic and environmental aetiologies of severe mental disorders.

genetics

Reassessment Of Lesion-Associated Gene And Variant Pathogenicity In Focal Human Epilepsies

PurposeIncreasing availability of surgically resected brain tissue from Focal Cortical Dysplasia and low-grade epilepsy-associated tumor patients fostered large-scale genetic examination. However, assessment of germline and somatic variant pathogenicity remains difficult.\n\nMethodsHere, we critically reevaluated the pathogenicity for all neuropathology-associated variants reported to date in the PubMed and ClinVar databases, including 12 disease-related genes and 88 neuropathology-associated missense variants. We (1) assessed evolutionary gene constraint using the pLI and missense z scores, (2) applied guidelines by the American College of Medical Genetics and Genomics (ACMG), and (3) predicted pathogenicity by using PolyPhen-2, CADD, and GERP.\n\nResultsConstraint analysis classified only seven out of 12 genes to be likely disease-associated, while 35 (40%) of those 88 variants were classified as being variants of unknown significance (VUS) and 53 (60%) as being likely pathogenic (LPII). Pathogenicity prediction yielded discrimination between neuropathology-associated variants (LPII and VUS) and rare variant scores obtained from individuals present in the Genome Aggregation Database (gnomAD).\n\nConclusionWe conclude that several VUS are likely disease-associated and will be reclassified by future molecular evidence. In summary, interpretation of lesion-associated gene variants remains complex while the application of current ACMG guidelines including bioinformatic pathogenicity prediction will help improving interpretation and prediction.

genetics

Fine-mapping of genetic loci driving spontaneous clearance of hepatitis C virus infection

Approximately three quarters of acute HCV infections evolve to a chronic state, while one quarter are spontaneously cleared. Genetic predispositions strongly contribute to the development of chronicity. We have conducted a genome-wide association study to identify genomic variants underlying HCV spontaneous clearance using Immunochip in European and African ancestries. We confirmed two previously reported significant associations, in the IL28B/IFNL41,2 and MHC regions, with spontaneous clearance in the European population. We further fine-mapped the MHC association to a region of about 50 kilo base pairs, down from 1 mega base pairs in the previous study. Additional analyses suggested that the association in the MHC locus might be significantly stronger for virus subtype 1a than 1b, suggesting that viral subtype may have influenced the genetic mechanism underlying the clearance of HCV.

immunology

New mutations, old statistical challenges

Based on targeted sequencing of 208 genes in 11,730 neurodevelopmental disorder cases, Stessman et al. report the identification of 91 genes associated (at a False Discovery Rate [FDR] of 0.1) with autism spectrum disorders (ASD), intellectual disability (ID), and developmental delay (DD)--including what they characterize as 38 novel genes, not previously reported as connected with these diseases1.\n\nIf true, this would represent a substantial step forward. Unfortunately, each of the two discovery analyses (1. De novo mutation analysis and, 2. a comparison of private mutations with public control data) contain critical statistical flaws. When one accounts for these problems, fewer than half of the genes--and very few, if any, of the novel findings--survive. These errors have implications for how future analyses should be conducted, for understanding the genetic basis of these disorders, and for genomic medicine.\n\nWe discuss the two main ana ...

genetics

The rate of false polymorphisms introduced when imputing genotypes from global imputation panels

Previous studies1,2 have shown that large multi-population imputation reference panels increases the number of well-imputed variants. However, to our knowledge, no previous studies have evaluated the rate of introduced variation in monomorphic sites of the study population when using imputation panels with admixed populations. In this study we evaluate the rate of false positive variants introduced by the imputation of Finnish genotype data using global reference panels (Haplotype Reference Consortium1; HRC, and the 1000Genomes project Phase I3; 1000G) and compare the results to a Finnish population-specific reference panel combining whole genome and exome sequenced samples. In sites that were monomorphic in our test set, we observed high false positive rates for the global reference panels (4.0% for 1000G and 2.6% for HRC) compared to the Finnish panel (0.26%). This rate was even higher (7.4%) when using a combination panel of 1000G and Finnish whole genome sequences with cross-panel imputation.

genetics