bioRxiv ScienceSearch

Biology subjects

Reyer, H.

Publications and source records attributed to Reyer, H..

2 recordsLinked to original sources

rePROBE: Workflow for Revised Probe Assignment and Updated Probe-Set Annotation in Microarrays

Commercial and customized microarrays are valuable tools for the analysis of holistic expression patterns, but require the integration of the latest genomic information. This study provides a comprehensive workflow implemented in an R package (rePROBE) to assign the entire probes and to annotate the probe sets based on up-to-date genomic and transcriptomic information. The rePROBE R package is freely available at https://github.com/friederhadlich/rePROBE. It can be applied to available gene expression microarray platforms and addresses both public and custom databases. The revised probe assignment and updated probe-set annotation were applied to commercial microarrays available for different livestock species, i.e. ChiGene-1_0-st (Gallus gallus, 443,579 probes; 18,530 probe sets), PorGene-1_1-st (Sus scrofa, 592,005; 25,779) and BovGene-1_0-st (Bos taurus, 530,717; 24,759) as well as human (Homo sapiens, HuGene-1_0-st) and mouse (Mus musculus, HT_MG-430_PM) microarrays. Using current specie-specific transcriptomic information (RefSeq, Ensembl and partially non-redundant nucleotide sequences) and genomic information, the applied workflow revealed 297,574 probes for chickens (pig: 384,715; cattle: 363,077; human: 481,168; mouse: 324,942) assigned to 15,689 probe sets (pig: 21,673; cattle: 21,238; human: 23,495; mouse: 32,494). These are representative of 12,641 unique genes that were both annotated and positioned (pig: 15,758; cattle: 18,046; human: 20,167; mouse: 16,335). Additionally, the workflow collects information on the number of single nucleotide polymorphisms (SNPs) within respective targeted genomic regions and thus provides a detailed basis for comprehensive analyses such as quantitative trait locus (eQTL) expression studies to identify quantitative and functional traits.

bioinformatics

Design of Experiments for Fine-Mapping Quantitative Trait Loci in Livestock Populations

Single nucleotide polymorphisms (SNPs) which capture a significant impact on a trait can be identified with genome-wide association studies. High linkage disequilibrium (LD) among SNPs makes it difficult to identify causative variants correctly. Thus, often target regions instead of single SNPs are reported. Sample size has not only a crucial impact on the precision of parameter estimates, it also ensures that a desired level of statistical power can be reached. We study the design of experiments for fine-mapping of signals of a quantitative trait locus in such a target region. A multi-locus model allows to identify causative variants simultaneously, to state their positions more precisely and to account for existing dependencies. Based on the commonly applied SNP-BLUP approach, we determine the z-score statistic for locally testing non-zero SNP effects and investigate its distribution under the alternative hypothesis. This quantity employs the theoretical instead of observed dependence between SNPs; it can be set up as a function of paternal and maternal LD for any given population structure. We simulated multiple paternal half-sib families and considered a target region of 1 Mbp. A bimodal distribution of estimated sample size was observed, particularly if more than two causative variants were assumed. The median of estimates constituted the final proposal of optimal sample size; it was consistently less than sample size estimated from single-SNP investigations which was used as a baseline approach. The second mode pointed to inflated sample sizes and could be explained by blocks of varying linkage phases leading to negative correlations between SNPs. Optimal sample size increased almost linearly with number of signals to be identified but depended much stronger on the assumption on heritability. For instance, three times as many samples were required if heritability was 0.1 compared to 0.3. These results enable the resource-saving design of future experiments for fine-mapping of candidate variants in structured and unstructured populations.

genetics