bioRxiv ScienceSearch

Biology subjects

Ruczinski, I.

Publications and source records attributed to Ruczinski, I..

4 recordsLinked to original sources

Improved Analysis of Phage ImmunoPrecipitation Sequencing (PhIP-Seq) Data Using a Z-score Algorithm

Phage ImmunoPrecipitation Sequencing (PhIP-Seq) is a massively multiplexed, phage-display based methodology for analyzing antibody binding specificities, with several advantages over existing techniques, including the uniformity and completeness of proteomic libraries, as well as high sample throughput and low cost. Data generated by the PhIP-Seq assay are unique in many ways. The only published analytical approach for these data suffers from important limitations. Here, we propose a new statistical framework with several improvements. Using a set of replicate mock immunoprecipitations (negative controls lacking antibody input) to generate background binding distributions, we establish a statistical model to quantify antibody-dependent changes in phage clone abundance. Our approach incorporates robust regression of experimental samples against the mock IPs as a means to calculate the expected phage clone abundance, and provides a generalized model for calculating each clones expected abundance-associated standard deviation. In terms of bias removal and detection sensitivity, we demonstrate that this z-score algorithm outperforms the previous approach. Further, in a large cohort of autoantibody-defined Sjogrens Syndrome (SS) patient sera, PhIP-Seq robustly identified Ro52, Ro60, and SSB/La as known autoantigens associated with SS. In an effort to identify novel SS-specific binding specificities, SS z-scores were compared with z-scores obtained by screening Ropositive sera from patients with systemic lupus erythematosus (SLE). This analysis did not yield any commonly targeted SS-specific autoantigens, suggesting that if they exist at all, their epitopes are likely to be discontinuous or post-translationally modified. In summary, we have developed an improved algorithm for PhIP-Seq data analysis, which was validated using a large set of sera with clinically characterized autoantibodies. This z-score approach will substantially improve the ability of PhIP-Seq to detect and interpret antibody binding specificities. The associated Python code is freely available for download here: https://github.com/LarmanLab/PhIP-Seq-Analyzer.

bioinformatics

Inferring Disease Risk Genes from Sequencing Data in Multiplex Pedigrees Through Sharing of Rare Variants

We previously demonstrated how sharing of rare variants (RVs) in distant affected relatives can be used to identify variants causing a complex and heterogeneous disease. This approach tested whether single RVs were shared by all sequenced affected family members. However, as with other study designs, joint analysis of several RVs (e.g. within genes) is sometimes required to obtain sufficient statistical power. Further, phenocopies can lead to false negatives for some causal RVs if complete sharing among affecteds is required. Here we extend our methodology (Rare Variant Sharing, RVS) to address these issues. Specifically, we introduce gene-based analyses, refine RV definition based on haplotypes, and introduce a partial sharing test based on RV sharing probabilities for subsets of affected family members. RVS also has the desirable features of not requiring external estimates of variant frequency or control samples, provides functionality to assess and address violations of key assumptions, and is available as open source software for genome-wide analysis. Simulations including phenocopies, based on the families of an oral cleft study, revealed the partial and complete sharing versions of RVS achieved similar statistical power compared to alternative methods (RareIBD and the Gene-Based Segregation Test), and had superior power compared to the pedigree Variant Annotation, Analysis and Search Tool (pVAAST) linkage statistic. In studies of multiplex cleft families, analysis of rare single nucleotide variants in the exome of 151 affected relatives from 54 families revealed no significant excess sharing in any one gene, but highlighted different patterns of sharing revealed by the complete and partial sharing tests.

genetics

Detection of de novo copy number deletions from targeted sequencing of trios

De novo copy number deletions have been implicated in many diseases, but there is no formal method to date however that identifies de novo deletions in parent-offspring trios from capture-based sequencing platforms. We developed Minimum Distance for Targeted Sequencing (MDTS) to fill this void. MDTS has similar sensitivity (recall), but a much lower false positive rate compared to less specific CNV callers, resulting in a much higher positive predictive value (precision). MDTS also exhibited much better scalability, and is available as open source software at github.com/JMF47/MDTS.

bioinformatics

Evolution of Hominin Polyunsaturated Fatty Acid Metabolism: From Africa to the New World

BackgroundThe metabolic conversion of dietary omega-3 and omega-6 18 carbon (18C) to long chain (> 20 carbon) polyunsaturated fatty acids (LC-PUFAs) is vital for human life. Fatty acid desaturase (FADS) 1 and 2 catalyze the rate-limiting steps in the biosynthesis of LC-PUFAs. The FADS region contains two haplotypes; ancestral and derived, where the derived haplotypes are associated with more efficient LC-PUFA biosynthesis and is nearly fixed in Africa. In addition, Native American populations appear to be nearly fixed for the lesser efficient ancestral haplotype, which could be a public health problem due to associated low LC-PUFA levels, while Eurasia is polymorphic. This haplotype frequency distribution is suggestive of archaic re-introduction of the ancestral haplotype to non-African populations or ancient polymorphism with differential selection patterns across the globe. Therefore, we tested the FADS region for archaic introgression or ancient polymorphism. We specifically addressed the genetic architecture of the FADS region in Native American populations to better understand this potential public health impact.\n\nResultsWe confirmed Native American ancestry is nearly fixed for the ancestral haplotype and is under positive selection. The ancestral haplotype frequency is also correlated to Siberian populations geographic location further suggesting the ancestral haplotype s role in cold weather adaptation and leading to the high haplotype frequency within Native American populations. We also find that the Neanderthal is more closely related to the derived haplotypes while the Denisovan clusters closer to the ancestral haplotypes. In addition, the derived haplotypes have a time to the most recent common ancestor of 688,474 years ago which is within the range of the modern-archaic hominin divergence.\n\nConclusionsThese results support an ancient polymorphism forming in the FADS gene region with differential selection pressures acting on the derived and ancestral haplotypes due to the old age of the derived haplotypes and the ancestral haplotype being under positive selection in Native American ancestry populations. Further, the near fixation of the less efficient ancestral haplotype in Native American ancestry suggests the need for future studies to explore the potential health risk of associated low LC-PUFA levels in Native American ancestry populations.

evolutionary biology