bioRxiv ScienceSearch

Biology subjects

Liley, J.

Publications and source records attributed to Liley, J..

2 recordsLinked to original sources

Accurate error control in high dimensional association testing using conditional false discovery rates

High-dimensional hypothesis testing is ubiquitous in the biomedical sciences, and informative covariates may be employed to improve power. The conditional false discovery rate (cFDR) is widely-used approach suited to the setting where the covariate is a set of p-values for the equivalent hypotheses for a second trait. Although related to the Benjamini-Hochberg procedure, it does not permit any easy control of type-1 error rate, and existing methods are over-conservative. We propose a new method for type-1 error rate control based on identifying mappings from the unit square to the unit interval defined by the estimated cFDR, and splitting observations so that each map is independent of the observations it is used to test. We also propose an adjustment to the existing cFDR estimator which further improves power. We show by simulation that the new method more than doubles potential improvement in power over unconditional analyses compared to existing methods. We demonstrate our method on transcriptome-wide association studies, and show that the method can be used in an iterative way, enabling the use of multiple covariates successively. Our methods substantially improve the power and applicability of cFDR analysis.

genomics

Improved consistency in estimates of conditional false discovery rates increases power relative to both existing methods and parametric estimators

A common aim in high-dimensional association studies is the identification of the subset of investigated variables associated with a trait of interest. Using association statistics on the same variables for a second related trait can improve power. An important quantity in such analyses is the conditional false-discovery rate (cFDR), the probability of non-association with the trait of interest given p-value thresholds for both traits. The cFDR can be used for hypothesis testing and as a posterior probability in its own right. In this paper, we propose new estimators for the cFDR based on kernel density estimates and mixture-Gaussian models of effect sizes, the latter also allowing estimation of a local form of cFDR (cfdr). We also propose a general non-parametric improvement to existing estimators based on estimating a posterior probability previously estimated at 1. We find that new estimators have the desirable property of smooth rejection regions, but, unexpectedly, do not improve the power of the method, even when distributional assumptions are true. Furthermore, we find that although the local cfdr represents a theoretically optimal decision boundary, noisiness in its estimation means it is less powerful than corresponding cFDR estimates. We find, however, that the non-parametric adjustment increases power for every estimator. We demonstrate the best method on transcriptome-wide association study datasets for breast and ovarian cancers. The findings from this analysis are of both theoretical and pragmatic interest, giving insight into the nature of cFDR and the behaviour of false-discovery rates in a two-dimensional setting. Our methods allow improved control over the behaviour of the cFDR estimator and improved power in high-dimensional hypothesis testing.

genomics