bioRxiv ScienceSearch

Biology subjects

Hong Gao

Publications and source records attributed to Hong Gao.

2 recordsLinked to original sources

Increasing the Efficiency of Genome-wide Association Mapping via Hidden Markov Models

With the rapid production of high dimensional genetic data, one major challenge in genome-wide association studies is to develop effective and efficient statistical tools to resolve the low power problem of detecting causal SNPs with low to moderate susceptibility, whose effects are often obscured by substantial background noises. Here we present a novel method that serves as an optimal technique for reducing background noises and improving detection power in genome-wide association studies. The approach uses hidden Markov model and its derivate Markov hidden Markov model to estimate the posterior probabilities of a markers being in an associated state. We conducted extensive simulations based on the human whole genome genotype data from the GlaxoSmithKline-POPRES project to calibrate the sensitivity and specificity of our method and compared with many popular approaches for detecting positive signals including the{chi} 2 test for association and the Cochran-Armitage trend test. Our simulation results suggested that at very low false positive rates (< 10-6), our method reaches the power of 0.9, and is more powerful than any other approaches, when the allelic effect of the causal variant is non-additive or unknown. Application of our method to the data set generated by Welcome Trust Case Control Consortium using 14,000 cases and 3,000 controls confirmed its powerfulness and efficiency under the context of the large-scale genome-wide association studies.

Genetics

Statistics of Cellular Evolution in Leukemia: Allelic Variations in Patient Trajectories Based on Immune Repertoire Sequencing

The evolution of a cancer system consisting of cancer clones and normal cells is a complex dynamic process with multiple interacting factors including clonal expansion, somatic mutation, and sequential selection. As a typical example, in patients with chronic lymphocytic leukemia (CLL), a monoclonal population of transformed B cells expands to dominate the B cell population in the peripheral blood and bone marrow. This expansion of transformed B cells suggests that they might evolve through processes distinct from those of normal B cells. Recent advances in next generation sequencing enable the high-throughput identification and tracking of individual B cell clones through sequencing of the V-D-J junction segments of the immunoglobulin heavy chain (IGH). Here we developed a statistical approach to modeling cellular evolution of the immune repertoire. Adapting the infinitely many alleles model from population genetics, we studied abnormalities occurring in the immune repertoire of patients as substantial deviations from the null model. The Ewens sampling test (EST) distinguished the immune repertoires of CLL patients with imminent relapse from healthy controls and patients in sustained remission. Extensive simulations based on sequencing data showed that EST is sensitive in detecting cancer-related derangements of the IGH repertoire. In addition, we suggest two potentially useful parameters: the rate at which donors B cell clones enter the circulation and the average time to regenerate a transplanted immune repertoire, both of which help to distinguish relapsing CLL patients from those in sustained remission and provide additional information about the dynamics of immune reconstitution in the latter patients. We anticipate that our models and statistics will be useful in diagnosis and prognosis of leukemia, and may be adapted for application to other diseases related to adaptive immunity.

Evolutionary Biology