bioRxiv ScienceSearch

Biology subjects

Bo Huang

Publications and source records attributed to Bo Huang.

2 recordsLinked to original sources

A fast multi-locus random-SNP-effect EMMA for genome-wide association studies

Although the mixed linear model (MLM) such as efficient mixed model association (EMMA), has been widely used in genome-wide association studies (GWAS), relatively little is known about fast and efficient algorithms to implement multi-locus GWAS. To address this issue, we report a fast multi-locus random-SNP-effect EMMA (FASTmrEMMA). In this method, a new matrix transformation was constructed to obtain a new genetic model that includes only quantitative trait nucleotide (QTN) variation and normal residual error; letting the number of nonzero eigenvalues be one and fixing the polygenic-to-residual variance ratio was used to increase computing speed. All the putative QTNs with the [≤]0.005 P-values in the first step of the new method were included in one multi-locus model for true QTN detection. Owing to the multi-locus feature, the Bonferroni correction is replaced by a less stringent selection criterion. Results from analyses of both simulated and real data showed that FASTmrEMMA is more powerful in QTN detection, model fit and robustness, has less bias in QTN effect estimation, and requires less running time than the current single- and multi-locus methodologies for GWAS, such as E-BAYES, SUPER, EMMA, CMLM and ECMLM. Therefore, FASTmrEMMA provides an alternative for multi-locus GWAS.

Genetics

A scalable strategy for high-throughput GFP tagging of endogenous human proteins

A central challenge of the post-genomic era is to comprehensively characterize the cellular role of the [~]20,000 proteins encoded in the human genome. To systematically study protein function in a native cellular background, libraries of human cell lines expressing proteins tagged with a functional sequence at their endogenous loci would be very valuable. Here, using electroporation of Cas9/sgRNA ribonucleoproteins and taking advantage of a split-GFP system, we describe a scalable method for the robust, scarless and specific tagging of endogenous human genes with GFP. Our approach requires no molecular cloning and allows a large number of cell lines to be processed in parallel. We demonstrate the scalability of our method by targeting 48 human genes and show that the resulting GFP fluorescence correlates with protein expression levels. We next present how our protocols can be easily adapted for the tagging of a given target with GFP repeats, critically enabling the study of low-abundance proteins. Finally, we show that our GFP tagging approach allows the biochemical isolation of native protein complexes for proteomic studies. Together, our results pave the way for the large-scale generation of endogenously tagged human cell lines for the proteome-wide analysis of protein localization and interaction networks in a native cellular context.\n\nSIGNIFICANCE STATEMENTThe function of a large fraction of the human proteome still remains poorly characterized. Tagging proteins with a functional sequence is a powerful way to access function, and inserting tags at endogenous genomic loci allows the preservation of a near-native cellular background. To characterize the cellular role of human proteins in a systematic manner and in a native context, we developed a method for tagging endogenous human proteins with GFP that is both rapid and readily applicable at a genome-wide scale. Our approach allows studying both localization and interaction partners of the protein target. Our results pave the way for the large-scale generation of endogenously tagged human cell lines for a systematic functional interrogation of the human proteome.

Cell Biology