bioRxiv ScienceSearch

Biology subjects

Carlborg, O.

Publications and source records attributed to Carlborg, O..

6 recordsLinked to original sources

Genotyping by low-coverage whole-genome sequencing in intercross pedigrees from outbred founders: a cost efficient approach

BackgroundExperimental intercrosses between outbred founder populations are powerful resources for mapping loci contributing to complex traits (Quantitative Trait Loci or QTL). Here, we present an approach and accompanying software for high-resolution genotype imputation in such populations using whole-genome high coverage sequence data on founder individuals ([~]30x) and low coverage sequence data on intercross individuals ([~]0.4x). The method is illustrated in a large F2 pedigree between lines of chickens that have been divergently selected for 40 generations for the same trait (body weight at 8 weeks of age).\n\nResultsDescribed is how hundreds of individuals were whole-genome sequenced in a cost- and time-efficient manner using a Tn5-based library preparation protocol optimized for this application. In total, 7.6M markers segregated in this pedigree and 10.0 to 13.7% were informative for imputing the founder line genotypes within the F0-F2 families. The genotypes imputed from low coverage sequence data were consistent with the founder line genotypes estimated using SNP and microsatellite markers both at individual imputed sites (92%) and across the genome of individual chickens (93%). The resolution of the recombination breakpoints was high with 50% being resolved within <10kb.\n\nConclusionsA method for genotype imputation from low-coverage whole-genome sequencing in outbred intercrosses is described and evaluated. By applying it to an outbred chicken F2 cross it is illustrated that it provides high quality, high-resolution genotypes in a time and cost efficient manner.

genetics

On the relationship between high-order linkage disequilibrium and epistasis

A plausible explanation for statistical epistasis revealed in genome wide association analyses is the presence of high order linkage disequilibrium (LD) between the genotyped markers tested for interactions and unobserved functional polymorphisms. Based on findings in experimental data, it has been suggested that high order LD might be a common explanation for statistical epistasis inferred between local polymorphisms in the same genomic region. Here, we empirically evaluate how prevalent high order LD is between local, as well as distal, polymorphisms in the genome. This could provide insights into whether we should account for this when interpreting results from genome wide scans for statistical epistasis. An extensive and strong genome wide high order LD was revealed between pairs of markers on the high density 250k SNP-chip and individual markers revealed by whole genome sequencing in the A. thaliana 1001-genomes collection. The high order LD was found to be more prevalent in smaller populations, but present also in samples including several hundred individuals. An empirical example illustrates that high order LD might be an even greater challenge in cases when the genetic architecture is more complex than the common assumption of bi-allelic loci. The example shows how significant statistical epistasis is detected for a pair of markers in high order LD with a complex multi allelic locus. Overall, our study illustrates the importance of considering also other explanations than functional genetic interactions when genome wide statistical epistasis is detected, in particular when the results are obtained in small populations of inbred individuals.

genetics

Explorations of the polygenic genetic architecture of flowering time in the worldwide Arabidopsis thaliana population

As a locally adapted complex trait, flowering time in Arabidopsis thaliana has attracted much attention in genetics. Most studies have, however, focused on contributions by individual loci rather than the joint contributions by the large number of loci in the genetic architecture of flowering time to local and global adaptation. In an earlier study, we reported 46 loci associated with flowering time variation during growth at 10{degrees}C, 16{degrees}C or both in the 1,001-genomes collection of Arabidopsis thaliana accessions. Here, we explore how these loci together contribute to differences among genetically defined, and geographically divided, subpopulations across the native range of this species. Our approach was to define flowering time as a trait, and the measurements at 10 and 16 {degrees}C as two independent measurements of it. This facilitated explorations of the dynamics in the genetic architecture -which loci contribute and their effects- of flowering time across growth temperatures and their potential roles in local and global adaptation. The overall flowering time differences between populations could be explained by subtle changes in allele-frequencies and gradual changes in phenotype due to globally present alleles. More extreme local adaptations were on several occasions due to contributions by regional alleles with relatively large effects. About 2/3 of the 48 evaluated flowering time loci had similar effects on flowering time at 10{degrees}C and 16{degrees}C, while the remaining 1/3 had different effects in the two temperatures, suggesting an important contribution of gene by temperature interactions to this trait. There are also indications that co-evolution of functionally connected alleles in local populations has been important for local adaptation. Overall, this study provides deeper insights to the polygenic genetic basis of flowering time variation in Arabidopsis thaliana across a wide range of ecological habitats.\n\nAuthor SummaryMany genes can affect flowering time in Arabidopsis thaliana, but their contribution to natural flowering time variation in the worldwide population is largely unknown. We explored how 48 loci associated with flowering time, measured at 10{degrees}C and 16{degrees}C, or their difference, for the same wild collected 1,001-genomes Arabidopsis thaliana accessions together contribute to differences among the genetically defined and geographically divided subpopulations from the native species range. The overall flowering time differences among these subpopulations could be explained by the joint small effects of globally present alleles, suggesting an important contribution by polygenic adaptation for this trait. Most alleles with large effects on flowering were present only in some populations, facilitating more extreme local adaptations. Long-range LD was observed between genes in several biological pathways, indicating possible local adaptation via co-evolution of functionally connected polymorphisms. The genetic architecture of flowering time was also found to depend on the growth temperature. Most flowering time loci had similar effects on flowering time measured at 10{degrees}C and 16{degrees}C, but the effects of about 1/3 of them had effects that varied with temperature. Overall, new insights are provided to how the polygenic architecture of flowering time has facilitated its colonisation of a wide range of ecological habitats.

genetics

A multi-locus association analysis method integrating phenotype and expression data reveals multiple novel associations to flowering time variation in wild-collected Arabidopsis thaliana

AO_SCPLOWBSTRACTC_SCPLOWWhen a species adapts to a new habitat, selection for the fitness traits often result in a confounding between genome-wide genotype and adaptive alleles. It is a major statistical challenge to detect such adaptive polymorphisms if the confounding is strong, or the effects of the adaptive alleles are weak. Here, we describe a novel approach to dissect polygenic traits in natural populations. First, candidate adaptive loci are identified by screening for loci that are directly associated to the trait or control the expression of genes known to affect it. Then, the multi-locus genetic architecture is inferred using a backward elimination association analysis across all the candidate loci using an adaptive false-discovery rate based threshold. Effects of population stratification are controlled by corrections for population structure in the pre-screening step and by simultaneously testing all candidate loci in the multi-locus model. We illustrate the method by exploring the polygenic basis of an important adaptive trait, flowering time in Arabidopsis thaliana, using public data from the 1,001 genomes project. Our method revealed associations between 33 (29) loci and flowering time at 10 (16){degrees}C in this collection of natural accessions, where standard genome wide association analysis methods detected 5 (3) loci. The 33 (29) loci explained approximately 55 (48)% of the total phenotypic variance of the respective traits. Our work illustrates how the genetic basis of highly polygenic adaptive traits in natural populations can be explored in much greater detail by using new multi-locus mapping approaches taking advantage of prior biological information as well as genome and transcriptome data.

genetics

On The Relationship Between Epistasis And Genetic Variance-Heterogeneity

Epistasis and genetic variance heterogeneity are two non-additive genetic inheritance patterns that are often, but not always, related. Here we use theoretical examples and empirical results from analyses of experimental data to illustrate the connection between the two. This includes an introduction to the relationship between epistatic gene-action, statistical epistasis and genetic variance heterogeneity and a brief discussion about how other genetic processes than epistasis can also give rise to genetic variance heterogeneity.\n\nHighlightGenetic effects on the trait variance, rather than the mean, have been found in several studies. Here we discuss how this sometimes, but not always, can be caused by epistasis.

genetics

A complex multi-locus, multi-allelic genetic architecture underlying the long-term selection-response in the Virginia body weight line of chickens

The ability of a population to adapt to changes in their living conditions, whether in nature or captivity, often depends on polymorphisms in multiple genes across the genome. In-depth studies of such polygenic adaptations are difficult in natural populations, but can be approached using the resources provided by artificial selection experiments. Here, we dissect the genetic mechanisms involved in long-term selection responses of the Virginia chicken lines, populations that after 40 generations of divergent selection for 56-day body weight display a nine-fold difference in the selected trait. In the F15 generation of an intercross between the divergent lines, 20 loci explained more than 60% of the additive genetic variance for the selected trait. We focused particularly on seven major QTL and found that only two fine-mapped to single, bi-allelic loci; the other five contained linked loci, multiple alleles or were epistatic. This detailed dissection of the polygenic adaptations in the Virginia lines provides a deeper understanding of genome-wide mechanisms involved in the long-term selection responses. The results illustrate that long-term selection responses, even from populations with a limited genetic diversity, can be polygenic and influenced by a range of genetic mechanisms.

genetics